#48 · Primary category: MLOps & Evaluation
opencompass
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Project last updated:08/27/26
GitHub Stars
7.4K
Forks
854
Contributors
185
License
Apache-2.0
Why we included this project
Choosing between large language models usually means trusting vendor claims unless you have comparable numbers of your own. OpenCompass addresses that by running models through more than a hundred datasets, spanning general knowledge, math reasoning, long-context tasks, and coding benchmarks, so you can see how candidates genuinely differ on the work you care about. It covers a broad range of models, from open weights like Llama, Mistral, Qwen, and GLM to hosted APIs such as GPT-4 and Claude, which makes cross-family comparison straightforward. For evaluations where a single accuracy score is not enough, it adds LLM-as-a-judge evaluation and cascading evaluators for more complex assessment pipelines. The configuration-driven setup suits teams that want reproducible, standardized results before picking a model or shipping a release.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models