#8 · Primary category: Data Quality & Cleaning
fastdup
fastdup is a powerful, free tool designed to rapidly generate valuable insights from image and video datasets. It helps enhance the quality of both images and labels, while significantly reducing data operation costs, all with unmatched scalability.
Project last updated:08/23/26
GitHub Stars
1.9K
Forks
93
Contributors
23
License
Other
Why we included this project
Getting a clean training set is usually the slowest part of any vision project, and fastdup attacks that bottleneck directly. Point it at a folder of images or video frames and it flags duplicates and near-duplicates, outliers, corrupted files, and clusters of visually similar shots, so you can prune a dataset before it ever reaches a training run. It does this on a single machine for tens of millions of images by working from compact similarity embeddings rather than raw pixels, and its gallery-style reports let you eyeball flagged groups and catch label errors without writing your own inspection scripts. For computer vision engineers and small data teams, it is a fast, scriptable first pass over raw visual data that would otherwise eat hours of manual review.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
cleanlab
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
DataFlow
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
qsv
Blazing-fast Data-Wrangling toolkit
MNBVC
MNBVC is a massive Chinese corpus benchmarked against ChatGPT's 40T data, covering mainstream and niche cultures, with diverse text forms including news, essays, novels, and more.