#134 · Primary category: MLOps & Evaluation
ClawProBench
ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.
Project last updated:08/25/26
GitHub Stars
823
Forks
54
Contributors
3
License
Apache-2.0
Why we included this project
ClawProBench is built for a specific problem: comparing agent models inside the OpenClaw runtime on tasks that actually run, not on static question-answer sets. It executes agents live against more than a hundred scenarios with deterministic grading, so the same task gives the same pass/fail result across repeated trials rather than a score that wobbles between runs. That consistency is what makes a benchmark result worth trusting before you commit a workflow to a particular model. The project also publishes a public leaderboard and a Hugging Face dataset, so you can reproduce results or add your own scenarios to the suite. Teams picking agent models for production, and researchers who want a transparent, runtime-grounded evaluation setup, will find the harness and its structured reports directly useful.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models