#89 · Primary category: MLOps & Evaluation
llm-datasets
Curated list of datasets and tools for post-training.
Project last updated:04/29/26
GitHub Stars
4.8K
Forks
394
Contributors
8
License
Other
Why we included this project
If you are getting ready to fine-tune or align your own model, this is the reference you want open in a tab. It is a living directory of SFT and RL/preference datasets grouped by domain, and each entry notes sample counts, whether the data includes reasoning traces, and the license. The same repository pulls together post-training tools and guides, so you can go from picking a starting mixture for a chat assistant or a reasoning model to actually assembling the pipeline. The opening section on dataset quality, covering accuracy, diversity, and complexity, gives a quick sanity check before you commit to any collection. It is a curated resource rather than deployable software, but it will save you real time when you are building your own training data.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models