#1 · Primary category: Data Quality & Cleaning

cleanlab

active-learning annotation anomaly-detection data-annotation data-centric-ai data-cleaning data-curation data-labeling data-profiling data-quality data-science data-validation datasets exploratory-data-analysis labeling machine-learning noisy-labels out-of-distribution-detection outlier-detection weak-supervision

Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.

Project last updated:01/13/26

GitHub Stars

11.6K

Forks

921

Contributors

55

License

Apache-2.0

Why we included this project

cleanlab is for teams who suspect their training data is messier than it looks. You fit your own classifier, then feed cleanlab the predicted probabilities and feature embeddings, and it flags label errors, outliers, near-duplicates, and other problems without making you hand-check every sample. For anyone building production ML on medical records, user-generated content, or sensor logs, that turns a slow manual audit into a scripted check you can run on each new dataset. The Datalab interface compiles everything into one report, and it handles text, audio, image, and tabular data. When clean labels are scarce or annotation is expensive, it helps because it points human reviewers at the samples where the model's own uncertainty already suggests something is off.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category