#16 · Primary category: Synthetic Data Generation
augraphy
Augmentation pipeline for rendering synthetic paper printing, faxing, scanning and copy machine processes
Project last updated:07/20/25
GitHub Stars
573
Forks
63
Contributors
20
License
MIT
Why we included this project
Paired clean-and-noisy document images are scarce, and that scarcity is the problem Augraphy was built around. It takes a clean document image and runs it through a configurable pipeline that reproduces the wear real office equipment leaves on paper, including dirty laser or inkjet printers, low-resolution fax machines, bleed-through, folds, shadows, and scanner noise. Every degraded output comes with a perfect ground-truth copy because the pipeline starts from a known original, which makes it practical for training document restoration, OCR, and deskewing models. It also propagates masks, keypoints, and bounding boxes through spatial transforms, so it fits object-detection and segmentation workflows, not just image-to-image tasks. Teams that need realistic synthetic training data without collecting and labeling thousands of physical pages get the noisy copies and the clean originals in one package.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
SDV
Synthetic data generation for tabular data
distilabel
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
kubric
A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
UltraChat
Large-scale, Informative, and Diverse Multi-round Chat Data (and Models)
synthetic-data-generator
SDG is a specialized framework designed to generate high-quality structured tabular data.