#138 · Primary category: MLOps & Evaluation
evaluation-guidebook
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
Project last updated:12/03/25
GitHub Stars
2.1K
Forks
126
Contributors
13
License
Other
Why we included this project
Deciding how to test an LLM is often harder than building it, and this guidebook collects the lessons the authors learned running the Open LLM Leaderboard and building lighteval. It covers automatic benchmarks, human evaluation, and LLM-as-a-judge setups, then goes on to troubleshooting and reproducibility notes that read like they were written after real deployment struggles. Beginners can start with the basics in each chapter, while teams already doing evals can jump straight to the design and tuning sections, and the starred links point to deeper reading the authors actually recommend. It is a solid reference for anyone writing custom evals for a production model or trying to pick sensible metrics before launch. One caveat: the repository is no longer maintained and redirects to a newer hosted version, so treat it as the stable core of the material.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models