#358 · Primary category: Education & Research

LLM-eval-survey

benchmark evaluation large-language-models llm llms model-assessment

The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".

Project last updated:08/01/26

GitHub Stars

1.6K

Forks

101

Contributors

19

License

Other

Why we included this project

This repository is the maintained companion to a survey paper on evaluating large language models, and for keeping current it beats the paper itself, since the arXiv version cannot be updated in real time. Rather than shipping code, it groups hundreds of papers and benchmarks by what is being tested, with sections on natural language understanding, reasoning, robustness, ethics, social and natural science tasks, and medical and agent applications, and by where the evaluation happens. That structure lets you trace how a capability such as reasoning has been measured over time and which datasets keep appearing as reference points. The maintainers accept pull requests, so the collection keeps pace with new benchmarks. Researchers and engineers designing evaluation studies of their own can use it to find established baselines and to spot gaps before investing in new protocols.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category