#8 · Primary category: Data Quality & Cleaning

fastdup

data-augmentation data-curation dataset deep-learning image image-analysis image-classfication image-classification image-duplicate-detection image-processing image-similarity machine-learning novelty-detection object-detection outlier-detection python visual-search visualization visualization-tools

fastdup is a powerful, free tool designed to rapidly generate valuable insights from image and video datasets. It helps enhance the quality of both images and labels, while significantly reducing data operation costs, all with unmatched scalability.

Project last updated:08/23/26

GitHub Stars

1.9K

Forks

93

Contributors

23

License

Other

Why we included this project

Getting a clean training set is usually the slowest part of any vision project, and fastdup attacks that bottleneck directly. Point it at a folder of images or video frames and it flags duplicates and near-duplicates, outliers, corrupted files, and clusters of visually similar shots, so you can prune a dataset before it ever reaches a training run. It does this on a single machine for tens of millions of images by working from compact similarity embeddings rather than raw pixels, and its gallery-style reports let you eyeball flagged groups and catch label errors without writing your own inspection scripts. For computer vision engineers and small data teams, it is a fast, scriptable first pass over raw visual data that would otherwise eat hours of manual review.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category