#16 · Primary category: Synthetic Data Generation

augraphy

augmentation-pipeline computer-vision crappification data-augmentation data-pipeline deep-neural-networks image-processing machine-learning synthetic-data synthetic-dataset-generation training-data

Augmentation pipeline for rendering synthetic paper printing, faxing, scanning and copy machine processes

Project last updated:07/20/25

GitHub Stars

573

Forks

63

Contributors

20

License

MIT

Why we included this project

Paired clean-and-noisy document images are scarce, and that scarcity is the problem Augraphy was built around. It takes a clean document image and runs it through a configurable pipeline that reproduces the wear real office equipment leaves on paper, including dirty laser or inkjet printers, low-resolution fax machines, bleed-through, folds, shadows, and scanner noise. Every degraded output comes with a perfect ground-truth copy because the pipeline starts from a known original, which makes it practical for training document restoration, OCR, and deskewing models. It also propagates masks, keypoints, and bounding boxes through spatial transforms, so it fits object-detection and segmentation workflows, not just image-to-image tasks. Teams that need realistic synthetic training data without collecting and labeling thousands of physical pages get the noisy copies and the clean originals in one package.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category