Model or dataset
sierra-research/tau2-bench avatar
sierra-research/tau2-bench

tau2-bench: evaluating customer service agents that talk to users and call tools

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

2,153 stars549 forksPythonMIT

At a glance

What is it?
tau2-bench (also written τ-bench) is a simulation framework for scoring tool-using agents against a policy, a user simulator and a task set. It installs with uv, runs text or voice domains, and its scoring is version-sensitive.
Who is it for?
Adopt tau2-bench if you have an agent that must follow a written policy while calling tools, and you want a repeatable score instead of a demo. Skip it if you need a generic chatbot benchmark or cannot supply paid model and voice API keys, since the framework drives real providers through LiteLLM and realtime audio APIs.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap tau2-bench fills: agents that must talk and act at once

Most agent evaluations pick one axis. Tool-calling benchmarks check whether the right function was invoked with the right arguments. Chat benchmarks check whether a reply reads well. Customer service agents have to do both in the same conversation, under a written policy, with a user who changes their mind mid-task. tau2-bench is built for that combination.

The README describes it as a simulation framework for evaluating customer service agents across multiple domains, supporting text-based half-duplex (turn-based) evaluation and voice full-duplex (simultaneous) evaluation. The audience is narrow and clear: people building or procuring support agents who need a number they can compare across models, not a vibe check. The repository ships five domains: mock, airline, retail, telecom and banking_knowledge. Each domain defines a policy the agent must follow, a set of tools it can call, a task set, and optionally user tools for the simulator. That four-part structure is the whole design. If your problem does not fit it, the framework will not bend.

How a run works: policy, tools, tasks and a simulated user

A tau2-bench run is a conversation between three parties. The agent under test sees the domain policy and the tool schemas. The user simulator plays a customer with a goal drawn from the task set. The orchestrator moves messages between them and records the trajectory. The README's orchestrator documentation covers two communication modes: half-duplex, where turns alternate, and full-duplex, where the voice path streams audio in both directions at once.

Scoring is not a single pass or fail. The evaluation documentation explains that task files carry evaluation_criteria.actions and a reward_basis field, and that reward_basis gates which parts of the reward are counted. So a task can require specific actions to have been taken, and the reward only reflects the dimensions the task declares. Grading therefore depends on the task data as much as on the agent, which is exactly why the July 2026 v1.0.1 release re-graded leaderboard submissions after fixing banking_knowledge task errors.

Voice mode swaps the text user simulator for an end-to-end audio path through realtime providers. The pyproject optional dependencies list elevenlabs, deepgram-sdk, pyaudio, livekit-agents and Google and AWS SDKs under the voice extra, and the README names OpenAI, Gemini and xAI as realtime providers. That is a heavier stack than most benchmarks, and it is why voice lives behind an extra rather than in the core install.

Installing tau2-bench and running a first airline evaluation

The project now installs with uv rather than pip install -e ., and requires Python >=3.12,<3.14. Clone the repository and sync the core environment, which covers text mode for airline, retail, telecom and mock.

bash
git clone https://github.com/sierra-research/tau2-bench
cd tau2-bench
uv sync

Optional features are separate syncs. Voice pulls the audio and realtime packages; knowledge pulls the retrieval pipeline used by banking_knowledge. On macOS the README notes voice also needs system dependencies installed with brew install portaudio ffmpeg.

bash
uv sync --extra voice
uv sync --extra knowledge
uv sync --all-extras

Keys come from a .env file copied from the template. The framework routes model calls through LiteLLM, so any provider LiteLLM supports can be configured, and the template lists ANTHROPIC_API_KEY, OPENAI_API_KEY, ELEVENLABS_API_KEY, DEEPGRAM_API_KEY, PINE_API_KEY, PINE_REALTIME_BASE_URL and OPENROUTER_API_KEY.

bash
cp .env.example .env

A first run names the domain and both models. This example is the one the README gives, with one trial and five tasks, and it writes results into data/simulations/.

bash
tau2 run --domain airline --agent-llm gpt-4.1 --user-llm gpt-4.1 \
  --num-trials 1 --num-tasks 5

Afterwards, tau2 view browses the saved simulations, and tau2 intro prints an overview of domains, commands and examples. One practical warning from the README: if you are evaluating rather than training, use the base task split, which is the default and matches the original τ-bench structure. The voice and knowledge additions live in their own splits.

The version trap: why two tau2-bench scores may not be comparable

The most important operational fact in the release notes is a compatibility break. The v1.0.1 release fixed task errors in banking_knowledge, and the README states plainly that results produced with tau2-bench below 1.0.1 are not comparable with 1.0.1 or later, and that affected leaderboard submissions were re-graded. Other domains are unaffected.

This is not a bug report, it is a property of the design. When the benchmark's own task data is the ground truth, correcting the data changes the score. Anyone quoting a banking_knowledge number from an earlier run is quoting a different benchmark. The release notes offer two escape hatches: re-score existing result files with tau2 evaluate-trajs --fresh-tasks, or pin the pre-v1.0.1 tag to reproduce the old behaviour. Both are documented, which is better than most benchmarks manage, but the burden is on you to record which one you used.

The v1.0.0 release went further, folding in 75 or more task fixes across airline, retail and banking, drawn from analysis in the SABER paper, plus voice and knowledge domains. A benchmark that revises its tasks this often is being maintained carefully, but it also means a score table without a version column is close to meaningless. The last push to the repository was on 2026-09-10, so the codebase is current.

Where tau2-bench is the wrong tool

The framework assumes a policy-following agent with a fixed tool surface and a scripted customer. Three consequences follow.

First, it is not a general assistant benchmark. There is no open-ended question answering or coding track here; the domains are customer service scenarios with policies and tools. If your agent is a research assistant, the mock domain will tell you nothing you care about.

Second, the user simulator is itself a model call, and it is configured separately from the agent with --user-llm. Variance from that model is part of your measurement. The README's example uses --num-trials 1, which is fine for a smoke test and far too thin for a claim; the flag exists because runs are stochastic.

Third, voice evaluation depends on external realtime services and, for the shipped personas, on voice IDs that external users do not have. The .env.example notes that the default voice IDs are Sierra-internal and will not work for external users, pointing to docs/voice-personas.md for creating your own in ElevenLabs. So the voice track is reproducible in structure but not in the exact audio conditions Sierra used, unless you supply equivalent voices. The README does not document rollback for a partially completed run, so plan around whole runs rather than resuming mid-way.

Alternatives and the difference in approach

The natural comparison is the original τ-bench, which this repository supersedes. The README's backward compatibility note says that evaluating an agent should use the base task split, described as matching the original τ-bench structure. In other words, the original task set is still inside this project as a split, and the newer voice and knowledge splits sit alongside it. Choosing between them is not choosing between two tools; it is choosing which split you run, and whether you want multimodal and retrieval coverage that the original never had.

Against tool-calling benchmarks that grade a single function call, the difference is the user. Here the agent must elicit information from a simulated customer before it can act, and the reward depends on whether the right actions happened across a multi-turn trajectory rather than at one step. That is a harder and noisier measurement, and it is the reason the task-fix releases matter so much.

Against building your own harness, the trade-off is control. A custom harness lets you define exactly the failure modes you care about. tau2-bench gives you a shared policy-and-tools format, a task schema with evaluation_criteria, a viewer, a leaderboard submission path, and an agent developer guide for plugging in your own agent. You accept its task data and its grading model in exchange for comparability.

Licence, upgrade cost and what to check before you publish

The project is MIT licensed, and pyproject.toml declares license = "MIT". That is permissive and places few obligations on internal or commercial use, but the licence covers the framework code, not the model providers you call through it or the voice services you connect to. Those carry their own terms, and running evaluations costs money per token and per audio minute. Nothing here is legal advice; check the terms of each provider you enable in .env.

Upgrade cost is real but bounded. Moving from τ²-bench to τ³-bench changed the installer to uv and raised the Python floor from >=3.10 to >=3.12,<3.14, and the README warns that some internal APIs were refactored. The pyproject pins litellm>=1.80.15,<1.82.7, so provider behaviour is constrained to a known range. Extras are additive, which keeps the core text install small.

Before publishing a number, verify three things: the tau2-bench version your results were produced with, the task split you ran, and whether the domain is banking_knowledge, where the 1.0.1 grading fix invalidated earlier scores. The leaderboard submission documentation and the JSON schema generator in the Makefile exist precisely so submissions carry that metadata.

Editorial conclusion

Adopt tau2-bench if you have an agent that must follow a written policy while calling tools, and you want a repeatable score instead of a demo. Skip it if you need a generic chatbot benchmark or cannot supply paid model and voice API keys, since the framework drives real providers through LiteLLM and realtime audio APIs. Before publishing numbers, pin the version: results from tau2-bench below 1.0.1 are not comparable with 1.0.1 or later on banking_knowledge, and check whether your task split is base or the newer voice and knowledge sets.

Frequently asked questions

What is tau2-bench?

It is a simulation framework from Sierra Research for evaluating customer service agents across multiple domains, each defining a policy, a set of tools, a task set and optionally user tools. It supports text-based half-duplex evaluation and voice full-duplex evaluation through realtime audio providers.

Which domains does tau2-bench include?

The README lists mock, airline, retail, telecom and banking_knowledge. The first four run in text mode with the core install; banking_knowledge requires the knowledge extra for its retrieval pipeline.

How do I install tau2-bench?

Clone the repository and run uv sync, which installs the core text-mode environment. Optional features come from extras such as uv sync --extra voice, --extra knowledge, --extra gym or --all-extras, and Python >=3.12,<3.14 is required.

How do I run a tau2-bench evaluation?

Set up .env with your provider keys, then run tau2 run with a domain and both models, for example tau2 run --domain airline --agent-llm gpt-4.1 --user-llm gpt-4.1 --num-trials 1 --num-tasks 5. Results are saved to data/simulations/ and can be browsed with tau2 view.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. sierra-research/tau2-bench on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sierra-research-tau2-bench.svg)](https://hysenlabs.com/projects/sierra-research-tau2-bench)