#54 · Primary category: MLOps & Evaluation
chinese-llm-benchmark
A live Chinese LLM benchmark evaluating 395+ models across 300 dimensions, offering leaderboards and a 2M+ defect library for research.
Project last updated:08/29/26
GitHub Stars
6.4K
Forks
262
Contributors
2
License
Other
Why we included this project
Choosing a Chinese-language model for a real product usually means comparing vague marketing claims, and ReLE tries to replace that with a fine-grained look at what models actually do. The project runs hundreds of commercial and open-weight models through roughly seven domains and about 300 sub-areas, covering school subjects, medical licensing, finance, law, and agent tool use, then publishes the scores as live leaderboards. The more unusual part is the defect library with over two million recorded failures, which lets researchers and engineers hunt down specific weaknesses instead of trusting an aggregate score. If you're picking a model for a Chinese-facing product or sanity-checking a private model before deployment, the per-dimension results give you something more actionable than a single rank. And because the benchmark keeps adding newly released models, the comparisons stay relevant as the field moves.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models