#3 · Primary category: Data Quality & Cleaning

data-juicer

data data-analysis data-pipeline data-processing data-science data-visualization foundation-models instruction-tuning large-language-models llm llms multi-modal pre-training synthetic-data

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

Project last updated:08/28/26

GitHub Stars

7.0K

Forks

412

Contributors

54

License

Apache-2.0

Why we included this project

Data-Juicer is for anyone whose time goes to cleaning data instead of training models. It includes more than 200 composable operators for cleaning, filtering, deduplicating, and restructuring text, image, audio, and video datasets headed into foundation-model work. Pipelines are defined as versioned YAML recipes, so a whole processing run can be committed and shared like code, which makes reproducible data prep practical for a team. The same operators run on a laptop or across a Ray cluster, so a quick prototype and a multi-terabyte pre-training corpus use the same building blocks. People preparing data for pre-training, instruction tuning, RL, RAG, or even embodied-agent traces get working pieces rather than writing glue code.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category