#260 · Primary category: AI Tool Directories & Curated Lists

Awesome-LLM-Eval

awsome-list awsome-lists benchmark bert chatglm chatgpt dataset evaluation gpt3 large-language-model leaderboard llama llm llm-evaluation machine-learning nlp openai qwen rag

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.

Project last updated:11/24/25

GitHub Stars

656

Forks

84

Contributors

5

License

MIT

Why we included this project

Choosing how to evaluate a large language model is often harder than picking the model itself, and this curated index is a good place to start. It gathers evaluation tools, benchmark datasets, leaderboards, and papers in one spot, sorted by what you actually want to test: general knowledge, coding, RAG pipelines, agent capabilities, long-context handling, multimodal behavior, and inference speed. The list also covers the models themselves and the frameworks used to train them, so you can see the full range of evaluation work in a single pass. Because the project is maintained alongside an academic survey on LLM evaluation, the selection reflects a considered view of the field rather than a random scrape of links. Teams planning an evaluation strategy can use it to find out which benchmarks and harnesses exist before committing to a specific toolchain.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category