Data Quality & Cleaning

Tools for detecting and fixing label errors, outliers, and other data issues in ML datasets to improve training data quality.

11 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

1–11 of 11

Rank Project Stars Forks
1 cleanlab

Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.

11.6K 921
2 DataFlow

Easy Data Preparation with latest LLMs-based Operators and Pipelines.

7.8K 1.1K
3 data-juicer

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

6.9K 412
4 qsv

Blazing-fast Data-Wrangling toolkit

3.8K 107
5 MNBVC

MNBVC is a massive Chinese corpus benchmarked against ChatGPT's 40T data, covering mainstream and niche cultures, with diverse text forms including news, essays, novels, and more.

4.3K 296
6 Curator

Scalable data pre processing and curation toolkit for LLMs

1.7K 320
7 DataProfiler

What's in your data? Extract schema, statistics and entities from datasets

1.6K 187
8 fastdup

fastdup is a powerful, free tool designed to rapidly generate valuable insights from image and video datasets. It helps enhance the quality of both images and labels, while significantly reducing data operation costs, all with unmatched scalability.

1.9K 93
9 dolma

Data and tools for generating and inspecting OLMo pre-training data.

1.5K 203
10 spotlight

Interactively explore unstructured datasets from your dataframe.

1.3K 92
11 lazynlp

Library to scrape and clean web pages to create massive datasets.

2.3K 324