#134 · Primary category: MLOps & Evaluation

ClawProBench

agent benchmark evaluation harness leaderboard llm openclaw

ClawProBench is a live-first benchmark harness for evaluating LLM agents in the OpenClaw runtime with deterministic grading and repeated-trial reliability.

Project last updated:08/25/26

GitHub Stars

823

Forks

54

Contributors

3

License

Apache-2.0

Why we included this project

ClawProBench is built for a specific problem: comparing agent models inside the OpenClaw runtime on tasks that actually run, not on static question-answer sets. It executes agents live against more than a hundred scenarios with deterministic grading, so the same task gives the same pass/fail result across repeated trials rather than a score that wobbles between runs. That consistency is what makes a benchmark result worth trusting before you commit a workflow to a particular model. The project also publishes a public leaderboard and a Hugging Face dataset, so you can reproduce results or add your own scenarios to the suite. Teams picking agent models for production, and researchers who want a transparent, runtime-grounded evaluation setup, will find the harness and its structured reports directly useful.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category