#1 · Primary category: Data Quality & Cleaning
cleanlab
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
Project last updated:01/13/26
GitHub Stars
11.6K
Forks
921
Contributors
55
License
Apache-2.0
Why we included this project
cleanlab is for teams who suspect their training data is messier than it looks. You fit your own classifier, then feed cleanlab the predicted probabilities and feature embeddings, and it flags label errors, outliers, near-duplicates, and other problems without making you hand-check every sample. For anyone building production ML on medical records, user-generated content, or sensor logs, that turns a slow manual audit into a scripted check you can run on each new dataset. The Datalab interface compiles everything into one report, and it handles text, audio, image, and tabular data. When clean labels are scarce or annotation is expensive, it helps because it points human reviewers at the samples where the model's own uncertainty already suggests something is off.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
qsv
Blazing-fast Data-Wrangling toolkit
MNBVC
MNBVC is a massive Chinese corpus benchmarked against ChatGPT's 40T data, covering mainstream and niche cultures, with diverse text forms including news, essays, novels, and more.
Curator
Scalable data pre processing and curation toolkit for LLMs