#108 · Primary category: MLOps & Evaluation
tau2-bench
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Project last updated:08/27/26
GitHub Stars
1.9K
Forks
480
Contributors
13
License
MIT
Why we included this project
Most agent benchmarks hand a model a fixed list of questions and grade the answers. τ-Bench targets the messier reality of customer service, staging full conversations between an agent, the tools it calls, and the human it serves across retail, airline, and banking domains. You can see not just whether the answer was right but whether the agent asked the right follow-ups and actually closed the issue. The τ³ release extends beyond text to voice and knowledge-retrieval scenarios, and the repo ships a CLI for running and scoring evaluations plus a live leaderboard for comparing models against a common reference point. That makes it a practical harness for teams building conversational agents or tool-calling products that will talk to real users.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models