#48 · Primary category: MLOps & Evaluation

opencompass

benchmark chatgpt evaluation large-language-model llama2 llama3 llm openai

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

Project last updated:08/27/26

GitHub Stars

7.4K

Forks

854

Contributors

185

License

Apache-2.0

Why we included this project

Choosing between large language models usually means trusting vendor claims unless you have comparable numbers of your own. OpenCompass addresses that by running models through more than a hundred datasets, spanning general knowledge, math reasoning, long-context tasks, and coding benchmarks, so you can see how candidates genuinely differ on the work you care about. It covers a broad range of models, from open weights like Llama, Mistral, Qwen, and GLM to hosted APIs such as GPT-4 and Claude, which makes cross-family comparison straightforward. For evaluations where a single accuracy score is not enough, it adds LLM-as-a-judge evaluation and cascading evaluators for more complex assessment pipelines. The configuration-driven setup suits teams that want reproducible, standardized results before picking a model or shipping a release.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category