#142 · Primary category: MLOps & Evaluation
deep_research_bench
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Project last updated:05/11/26
GitHub Stars
817
Forks
84
Contributors
2
License
Apache-2.0
Why we included this project
Deep research agents are hard to judge because the work is open-ended and there is no single right answer. This benchmark targets that problem directly: 100 expert-written, PhD-level tasks across 22 domains, split between Chinese and English, and grounded in the kind of web-search queries people actually make. The repo ships the evaluation pipeline and a rubric-based scoring framework, and a public leaderboard lets you compare your agent against other teams' results. For anyone shipping a research agent, that is a more reliable way to measure report quality, factual accuracy, and presentation than reading outputs by hand.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models