#108 · Primary category: MLOps & Evaluation

tau2-bench

ai benchmark conversational-agents language-model-agent llm

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Project last updated:08/27/26

GitHub Stars

1.9K

Forks

480

Contributors

13

License

MIT

Why we included this project

Most agent benchmarks hand a model a fixed list of questions and grade the answers. τ-Bench targets the messier reality of customer service, staging full conversations between an agent, the tools it calls, and the human it serves across retail, airline, and banking domains. You can see not just whether the answer was right but whether the agent asked the right follow-ups and actually closed the issue. The τ³ release extends beyond text to voice and knowledge-retrieval scenarios, and the repo ships a CLI for running and scoring evaluations plus a live leaderboard for comparing models against a common reference point. That makes it a practical harness for teams building conversational agents or tool-calling products that will talk to real users.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category