#2 · Primary category: Synthetic Data Generation
distilabel
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
Project last updated:08/24/26
GitHub Stars
3.4K
Forks
256
Contributors
41
License
Apache-2.0
Why we included this project
Distilabel is for teams that need training data for fine-tuning but do not want to hand-label thousands of examples. You assemble data-generation and AI-judging steps into Python pipelines and synthesize instruction pairs, preference data, and other training material at scale, following published research methods for tasks like preference labeling and reward-model data. A single API reaches across many model providers, so you are not tied to one vendor when generating or critiquing data. The project has shipped real results, including preference datasets with millions of examples and fine-tuned models that improved after AI-feedback filtering. If your bottleneck is clean, varied training data rather than model code, this framework is worth a close look.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
SDV
Synthetic data generation for tabular data
kubric
A data generation pipeline for creating semi-realistic synthetic multi-object videos with rich annotations such as instance segmentation masks, depth maps, and optical flow.
synthetic-data-generator
SDG is a specialized framework designed to generate high-quality structured tabular data.
UltraChat
Large-scale, Informative, and Diverse Multi-round Chat Data (and Models)
unrealcv
UnrealCV: Connecting Computer Vision to Unreal Engine