#11 · Primary category: Synthetic Data Generation
mostlyai
Synthetic Data SDK ✨
Project last updated:05/08/26
GitHub Stars
793
Forks
67
Contributors
19
License
Apache-2.0
Why we included this project
MostlyAI is a Python SDK that generates synthetic copies of your own data. You train a model on tabular, text, or time-series records, then produce as many synthetic samples as you need, with controls for differential privacy and conditional generation, plus rebalancing for underrepresented segments. It runs locally on your own compute by default, or you can point it at a remote endpoint, and it includes quality reports that score fidelity and privacy, so you can tell whether the output is trustworthy. Teams that need realistic test data without exposing real records will find it useful, whether they are data engineers, ML practitioners, or just building demos and staging environments from sensitive production data.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
SDV
Synthetic data generation for tabular data
distilabel
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
kubric
A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
synthetic-data-generator
SDG is a specialized framework designed to generate high-quality structured tabular data.
unrealcv
UnrealCV: Connecting Computer Vision to Unreal Engine