Data Quality & Cleaning
Tools for detecting and fixing label errors, outliers, and other data issues in ML datasets to improve training data quality.
11 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 1 |
cleanlab
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels. |
11.6K | 921 | 01/13/26 | Apache-2.0 |
| 2 |
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines. |
7.8K | 1.1K | 08/18/26 | Apache-2.0 |
| 3 |
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷 |
6.9K | 412 | 08/28/26 | Apache-2.0 |
| 4 |
qsv
Blazing-fast Data-Wrangling toolkit |
3.8K | 107 | 08/29/26 | Other |
| 5 |
MNBVC
MNBVC is a massive Chinese corpus benchmarked against ChatGPT's 40T data, covering mainstream and niche cultures, with diverse text forms including news, essays, novels, and more. |
4.3K | 296 | 08/28/26 | MIT |
| 6 |
Curator
Scalable data pre processing and curation toolkit for LLMs |
1.7K | 320 | 08/28/26 | Apache-2.0 |
| 7 |
DataProfiler
What's in your data? Extract schema, statistics and entities from datasets |
1.6K | 187 | 08/26/26 | Apache-2.0 |
| 8 |
fastdup
fastdup is a powerful, free tool designed to rapidly generate valuable insights from image and video datasets. It helps enhance the quality of both images and labels, while significantly reducing data operation costs, all with unmatched scalability. |
1.9K | 93 | 08/23/26 | Other |
| 9 |
dolma
Data and tools for generating and inspecting OLMo pre-training data. |
1.5K | 203 | 08/24/26 | Apache-2.0 |
| 10 |
spotlight
Interactively explore unstructured datasets from your dataframe. |
1.3K | 92 | 08/19/26 | MIT |
| 11 |
lazynlp
Library to scrape and clean web pages to create massive datasets. |
2.3K | 324 | 11/11/20 | Other |