#4 · Primary category: Data Catalog & Metadata Management
LLMDataHub
A quick guide (especially) for trending instruction finetuning datasets
Project last updated:11/28/23
GitHub Stars
3.4K
Forks
234
Contributors
1
License
MIT
Why we included this project
For anyone training or fine-tuning their own instruction-following model, finding good training data is often the hardest part of the job. This repository gathers a large set of openly available datasets in one place, grouped by what the data is for: alignment, domain-specific, pretraining, and multimodal work. Each entry notes the dataset size, language, type, and which published model trained on it, which is exactly the detail you need when picking a corpus for an Alpaca-style recipe. It isn't software you run; it's a curated reference map that saves hours of digging through scattered Hugging Face collections and research repositories. Because the entries track when each dataset appeared, you can also tell what was available at any given point.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.