#142 · Primary category: MLOps & Evaluation

deep_research_bench

agent benchmark deepresearch nlp

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Project last updated:05/11/26

GitHub Stars

817

Forks

84

Contributors

2

License

Apache-2.0

Why we included this project

Deep research agents are hard to judge because the work is open-ended and there is no single right answer. This benchmark targets that problem directly: 100 expert-written, PhD-level tasks across 22 domains, split between Chinese and English, and grounded in the kind of web-search queries people actually make. The repo ships the evaluation pipeline and a rubric-based scoring framework, and a public leaderboard lets you compare your agent against other teams' results. For anyone shipping a research agent, that is a more reliable way to measure report quality, factual accuracy, and presentation than reading outputs by hand.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category