τ²-Bench Airline tests whether a model can do an airline support agent's job: follow a policy manual, talk to a simulated customer, and call the right tools to search flights, change bookings, and issue refunds. The model doesn't need any domain knowledge; it's scored entirely on how well it executes tool calls across multi-step task trajectories. We run it continuously against the same provider endpoints that serve OpenRouter traffic, so a score reflects both the model and the provider running it. We use this benchmark because it has a high floor, so we can assess provider variance and not model capability. Our routing algorithm for tool call requests uses these same signals to send traffic to the best performing endpoints.
Last benchmark run
Google: Gemini 3.7 Flash leads τ²-Bench Airline at 80.6% as of Sep 16, 2026, 9:00 AM UTC. Airline tasks, policy, simulator and customer setup, step budget, and grading are held fixed across production provider endpoints, so the displayed results compare models and the providers serving them under the same evaluation conditions.
3.9% → 5.0%
Top-level rows use default routing where available; click a row to expand provider-pinned results.
| # | Model | Std dev | ||||
|---|---|---|---|---|---|---|
| 1 | Google: Gemini 3.7 Flash Pareto | 80.6% | -- | $0.077 | 2.0m | 14.7k |
| 2 | 80.2% | ±3.0pp | $0.96 | 2.2m | 5.79k | |
| 3 | 79.6% | ±2.3pp | $0.50 | 2.2m | 7.26k | |
| 4 | 78.7% | ±2.0pp | $1.28 | 1.7m | 5.69k | |
| 5 | 78.5% | ±3.6pp | $0.099 | 6.1m | 15.8k | |
| 6 | 78.3% | ±1.7pp | $0.098 | 3.1m | 9.09k | |
| 7 | 78.0% | -- | $0.69 | 2.0m | 6.91k | |
| 8 | 77.7% | ±1.1pp | $0.11 | 5.5m | 19.1k | |
| 9 | 77.3% | ±2.0pp | $0.12 | 2.5m | 20.5k | |
| 10 | StepFun: Step 3.7 Flash Pareto | 77.3% | -- | $0.020 | 3.8m | 10.9k |
| 11 | 77.1% | ±3.5pp | $0.15 | 9.1m | 35k | |
| 12 | 77.0% | ±2.6pp | $0.53 | 2.9m | 10.4k | |
| 13 | 76.9% | ±0.8pp | $0.10 | 2.6m | 8.57k | |
| 14 | 76.9% | ±3.0pp | $0.080 | 4.7m | 12.2k | |
| 15 | 76.9% | ±1.5pp | $0.41 | 1.7m | 5.68k | |
| 16 | 76.7% | -- | $0.028 | 3.7m | 15.4k | |
| 17 | 76.7% | -- | $1.86 | 5.6m | 22.9k | |
| 18 | 76.7% | ±0.0pp | $0.24 | 7.4m | 15.2k | |
| 19 | 76.7% | ±2.0pp | $0.20 | 2.3m | 7.91k | |
| 20 | 76.7% | ±0.7pp | $0.21 | 2.1m | 5.14k | |
| 21 | 76.2% | ±2.0pp | $0.50 | 2.2m | 8.14k | |
| 22 | Google: Gemma 4 31B Pareto | 76.1% | ±4.0pp | $0.016 | 5.5m | 8.28k |
| 23 | 76.0% | ±2.8pp | $0.042 | 2.7m | 7.59k | |
| 24 | 76.0% | -- | $0.35 | 4.1m | 11.6k | |
| 25 | 75.7% | ±2.5pp | $0.48 | 3.3m | 10k | |
| 26 | 75.7% | ±1.7pp | $0.11 | 1.7m | 25.1k | |
| 27 | 75.7% | ±4.3pp | $0.43 | 8.7m | 37.6k | |
| 28 | 75.7% | ±0.3pp | $0.30 | 2.5m | 10.8k | |
| 29 | 75.3% | ±1.7pp | $0.32 | 3.3m | 11.4k | |
| 30 | 75.3% | ±2.7pp | $0.50 | 2.6m | 8.28k | |
| 31 | 75.3% | ±2.0pp | $0.36 | 2.3m | 12.8k | |
| 32 | Z.ai: GLM 5.3 Flash Pareto | 75.3% | ±2.0pp | $0.006 | 2.2m | 4.18k |
| 33 | 75.3% | ±4.2pp | $0.036 | 2.0m | 5.54k | |
| 34 | 75.1% | ±3.3pp | $0.009 | 2.4m | 8.29k | |
| 35 | 74.9% | ±2.7pp | $0.038 | 2.8m | 7.71k | |
| 36 | 74.7% | ±0.7pp | $0.42 | 2.8m | 28.3k | |
| 37 | 74.6% | ±4.7pp | $0.10 | 10.7m | 23.8k | |
| 38 | 74.1% | ±3.6pp | $0.057 | 2.4m | 5.73k | |
| 39 | 74.1% | ±4.0pp | $0.028 | 6.6m | 12.7k | |
| 40 | 74.0% | ±0.0pp | $0.032 | 3.0m | 17.6k | |
| 41 | 73.7% | ±2.7pp | $0.061 | 3.8m | 9.41k | |
| 42 | 73.3% | ±1.3pp | $0.100 | 1.6m | 4.17k | |
| 43 | 73.3% | ±3.5pp | $0.040 | 3.5m | 13.4k | |
| 44 | 73.3% | -- | $0.24 | 3.5m | 12.5k | |
| 45 | 73.2% | ±8.2pp | $0.055 | 7.6m | 35k | |
| 46 | 73.0% | ±4.4pp | $0.037 | 2.6m | 5.62k | |
| 47 | 73.0% | ±3.3pp | $0.33 | 3.4m | 10.3k | |
| 48 | 73.0% | ±0.3pp | $0.24 | 2.4m | 22.9k | |
| 49 | 72.8% | ±5.9pp | $0.008 | 5.7m | 16.3k | |
| 50 | 72.3% | ±1.0pp | $0.097 | 2.4m | 51.2k | |
| 51 | 72.1% | ±2.6pp | $0.023 | 4.6m | 8.07k | |
| 52 | 71.8% | ±3.3pp | $0.067 | 4.2m | 7.51k | |
| 53 | 71.4% | ±7.6pp | $0.043 | 3.8m | 23.9k | |
| 54 | 71.3% | ±1.3pp | $0.064 | 5.8m | 34.2k | |
| 55 | 71.2% | ±2.7pp | $0.029 | 3.2m | 6k | |
| 56 | 71.1% | ±4.3pp | $0.018 | 2.3m | 5.87k | |
| 57 | 71.0% | ±6.2pp | $0.023 | 2.2m | 6.2k | |
| 58 | 70.9% | ±2.0pp | $0.34 | 8.1m | 9.11k | |
| 59 | 70.9% | ±3.6pp | $0.014 | 2.4m | 4.53k | |
| 60 | 70.7% | -- | $0.24 | 2.5m | 17.9k | |
| 61 | 70.7% | ±0.0pp | $0.009 | 1.7m | 5.4k | |
| 62 | 70.7% | ±2.4pp | $0.008 | 4.8m | 14.3k | |
| 63 | 69.7% | ±3.0pp | $0.13 | 11.6m | 58.4k | |
| 64 | 69.1% | ±4.1pp | $0.039 | 6.7m | 11.9k | |
| 65 | 68.7% | ±10.8pp | $0.045 | 1.8m | 3.93k | |
| 66 | Xiaomi: MiMo-V2-Flash Pareto | 68.5% | ±6.1pp | $0.006 | 43s | 3.58k |
| 67 | 68.4% | ±5.5pp | $0.025 | 11.9m | 33.7k | |
| 68 | 68.4% | ±3.8pp | $0.68 | 13.4m | 61.2k | |
| 69 | 68.3% | ±3.9pp | $0.017 | 4.8m | 13k | |
| 70 | 68.0% | ±2.0pp | $0.33 | 2.9m | 8.95k | |
| 71 | 67.8% | ±4.3pp | $0.056 | 7.7m | 8.93k | |
| 72 | 67.8% | ±3.3pp | $0.12 | 2.0m | 12.7k | |
| 73 | 67.3% | ±3.6pp | $0.080 | 62s | 3.68k | |
| 74 | 67.3% | -- | $0.11 | 80s | 4.28k | |
| 75 | 67.0% | ±0.3pp | $0.033 | 2.5m | 17.5k | |
| 76 | 66.6% | ±3.6pp | $0.015 | 82s | 5.49k | |
| 77 | 65.6% | ±11.1pp | $0.037 | 3.6m | 6.29k | |
| 78 | 64.1% | ±3.4pp | $0.013 | 3.5m | 13.3k | |
| 79 | 64.1% | ±4.8pp | $0.016 | 2.8m | 6.61k | |
| 80 | 63.0% | ±5.0pp | $0.092 | 2.2m | 12.5k | |
| 81 | 62.9% | -- | $0.020 | 2.7m | 22.2k | |
| 82 | 61.9% | ±6.3pp | $0.008 | 2.6m | 7.86k | |
| 83 | 61.3% | ±0.7pp | $0.22 | 3.2m | 15.1k | |
| 84 | 61.2% | ±4.2pp | $0.059 | 6.9m | 21.1k | |
| 85 | 60.9% | ±3.1pp | $0.047 | 3.8m | 3.8k | |
| 86 | 60.2% | ±10.5pp | $0.059 | 5.9m | 9.02k | |
| 87 | 58.4% | ±6.5pp | $0.024 | 61s | 2.79k | |
| 88 | 57.7% | ±0.3pp | $0.029 | 79s | 8.17k | |
| 89 | 57.3% | -- | $0.012 | 2.0m | 16.8k | |
| 90 | 55.3% | ±5.2pp | $0.11 | 2.7m | 3.28k | |
| 91 | 54.5% | ±5.9pp | $0.086 | 13.7m | 17.6k | |
| 92 | 51.8% | ±2.8pp | $0.025 | 7.3m | 51k | |
| 93 | 51.4% | ±2.5pp | $0.023 | 18.7m | 96.5k | |
| 94 | 48.7% | ±0.7pp | $0.12 | 51s | 2.66k | |
| 95 | 47.6% | ±4.5pp | $0.043 | 11.6m | 99.7k | |
| 96 | 47.3% | -- | $0.024 | 2.9m | 45.1k | |
| 97 | 47.0% | ±3.6pp | $0.015 | 62s | 2.3k | |
| 98 | 46.7% | ±5.2pp | $0.048 | 65s | 1.89k | |
| 99 | 46.3% | ±1.7pp | $0.20 | 52s | 2.08k | |
| 100 | 46.0% | -- | $0.099 | 65s | 2.75k | |
| 101 | 45.9% | ±4.1pp | $0.018 | 1.7m | 2.28k | |
| 102 | 44.7% | ±1.3pp | $0.21 | 47s | 2.22k | |
| 103 | 44.6% | ±2.3pp | $0.041 | 1.8m | 1.94k | |
| 104 | 44.0% | ±1.3pp | $0.029 | 74s | 2.45k | |
| 105 | 43.7% | ±6.1pp | $0.015 | 3.1m | 10.3k | |
| 106 | 43.7% | ±2.3pp | $0.014 | 1.6m | 10.2k | |
| 107 | 43.1% | ±0.6pp | $0.024 | 3.6m | 3.15k | |
| 108 | 43.0% | ±3.8pp | $0.042 | 5.7m | 8.59k | |
| 109 | 42.3% | ±1.0pp | $0.030 | 8.7m | 20.1k | |
| 110 | 42.1% | ±9.9pp | $0.015 | 7.7m | 13.2k | |
| 111 | 40.7% | ±4.4pp | $0.035 | 2.3m | 3.44k | |
| 112 | 40.0% | -- | $1.13 | 72s | 3.25k | |
| 113 | 38.7% | ±3.6pp | $0.012 | 43s | 550 | |
| 114 | 38.3% | ±0.3pp | $0.071 | 4.7m | 7.51k | |
| 115 | 37.7% | ±3.4pp | $0.016 | 1.9m | 2.64k | |
| 116 | 34.7% | ±3.2pp | $0.073 | 3.6m | 6.68k | |
| 117 | 32.8% | ±4.5pp | $0.020 | 67s | 2.67k | |
| 118 | 31.4% | ±6.5pp | $0.006 | 65s | 1.52k | |
| 119 | 27.0% | ±3.0pp | $0.021 | 66s | 3.95k | |
| 120 | 26.0% | ±1.3pp | $0.079 | 9.8m | 5.67k | |
| 121 | 18.7% | ±6.3pp | $0.020 | 1.8m | 2.7k | |
| 122 | 16.7% | -- | $0.017 | 58s | 2.49k | |
| 123 | 10.3% | ±0.3pp | $0.007 | 61s | 2.32k |
It's a tool-calling benchmark that is hard to game. Grading depends on live tool-call trajectories rather than memorized answers, so it resists training-data leakage better than Q&A-style evals. It exercises every tool-calling failure mode (wrong arguments, skipped policy checks, giving up, hallucinated confirmations) at a relatively low cost per run. The relative scores also carry more signal than the absolute ones. The same model can score differently across providers, and those deltas are what Exacto routing uses to pick higher-accuracy endpoints.
Each task is a simulated airline support conversation with a scripted user, a toolbox (flight search, booking changes, refunds, loyalty policies), and a gold reference solution. A task passes only if the final database state and the messages to the user match the reference; partial credit is not awarded.
There is still headroom. Top models fail roughly one in five tasks, and the airline domain is the hardest τ²-Bench split. Accuracy differences here separate models that follow multi-step policies from ones that merely chat well.
The floor is high, though. Many tasks reward inaction. A refusal task with an empty gold action list passes for any agent that changes nothing. Even weak models score well above zero, so the meaningful spread sits at the top of the range.
The benchmark is public, so tasks may appear in training corpora. Contamination inflates scores less here than in Q&A-style evals, though, since a leaked task still has to be executed correctly, step by step, against a live database.
Scoring fidelity has limits. The checker verifies two things: the final database hash and exact substring matches in the agent's messages. Each task's natural-language assertions ("agent should refuse the cancellation") are metadata, and no judge model reads the transcript. So a savings calculation fails if the agent says "$23,552.50" when the checker greps for "23553".
The user simulator matters too. We pin it to gemini-2.5-flash so agent scores stay comparable, but the sim is itself an LLM with failure modes of its own. It can stop the conversation before the agent finishes, leak its hidden task instructions, or keep a stuck agent looping until the 200-step ceiling kills the run. Swapping the sim model shifts absolute scores, which is why cross-paper τ²-Bench numbers rarely line up exactly.
Every task ships a gold solution: a list of tool calls, strings the agent must say, and natural-language assertions. After the conversation ends, the checker replays the gold tool calls against a fresh database and compares hashes with the agent's final database. It then greps the agent's messages for each required string. The reward is the product of those two checks:
reward = db_match × communicate_met // each ∈ {0, 1}
db_match = hash(agent DB) == hash(gold DB)
communicate_met = every required string appears in an agent message
any run that hits MAX_STEPS instead of a clean stop scores 0 outrightThe rollouts below are from real runs, with gemini-2.5-flash as the user simulator throughout.
Task 17 agent: openai/gpt-5.1
For reservation FQ8APE: add 3 checked bags, swap the passenger to Omar Rossi, and upgrade basic economy to economy, paying with a gift card.
The agent looked up the user, found the right reservation among several, confirmed the changes and payment method, then made all three writes: passenger swap, cabin upgrade, and bags. The final database hashes match the gold state and the run ends on USER_STOP, so reward is 1. This is what the eval is designed to measure: multi-step tool use under policy constraints, done correctly.
Task 0 agent: openai/gpt-4o-mini
Cancel reservation EHGLP3. The booking is more than 24 hours old, basic economy, no travel insurance, so policy says no cancellation.
The agent pulled the reservation, saw a basic economy fare booked more than 24 hours ago with no insurance, and cancelled it anyway. It even promised a $208 refund. The checker compares the final database against the gold database (unmodified), the hashes differ, reward is 0. The same model refused this exact cancellation in a different epoch; sampling variance flips the outcome.
Task 18 agent: openai/gpt-4o-mini
Downgrade all five business-class reservations to economy, then report the total amount saved. The correct total is $23,553.
The agent executed all five downgrades correctly, and the database check passed. Then it computed the savings from a partial list of fares and told the user $14,965 instead of $23,553. The communicate check greps every agent message for the literal string "23553", finds nothing, and zeroes the whole task. One wrong arithmetic answer erased five correct database writes.
Task 11 agent: meta-llama/llama-3.1-8b-instruct
Remove passenger Sophia from reservation GV1N64. Policy forbids changing the passenger count; the correct move is to downgrade both passengers to basic economy and refund $5,244.
The agent’s very first move is a write with fully invented arguments: a reservation ID the user never gave and a passenger "John Doe" born 1990-01-01 who exists nowhere in the data. After finding the real reservation it sends update_reservation_passengers with a one-passenger array against a two-passenger booking, gets "number of passengers does not match", and retries the identical call four more times, at one point hallucinating "Sophia Smith, 1992-01-01" as the second passenger. It never consults the policy (removing a passenger is forbidden; the correct move is a cabin downgrade), quotes a fabricated $200 refund, and both checks fail. The arguments are well-formed JSON that the tool schema accepts; they are just wrong about the world, which is why small models can look fine on schema-level tool-call metrics and still fail tasks like this.
Task 14 agent: openai/gpt-4o-mini
Rebook the cheapest business round trip and split payment across gift cards, one certificate, and a credit card. The correct split puts $44 on the card.
The agent built a booking where the payment amounts did not sum to the ticket price. The tool rejected it with the same error every time, and the agent retried the identical call dozens of times until the run hit its 200-step ceiling. Any run that terminates on MAX_STEPS scores 0 before the database is even compared. This transcript is 202 messages long; the excerpt below is the loop.
Task 0 agent: openai/gpt-4o-mini
Same task as the policy-break failure: the user wants to cancel EHGLP3, and policy says no.
This run earned a legitimate pass: the agent checked the reservation, cited the 24-hour rule, and refused. But look at what the checker actually verified: an untouched database and an empty communicate list. An agent that stonewalled every request, or transferred to a human immediately, would score identically. Refusal tasks measure "did nothing break", so they inflate scores for overly cautious models.
Scores aggregate all successful runs, weighted by task count, with a minimum of 45 graded tasks per model-provider pair. A model's headline score uses its default routing (not pinned to a provider) when one exists; otherwise it falls back to the median provider. The standard deviation is measured across runs for that representative result. Cost, time, and token figures are per-task averages from the same runs. Best value is the cheapest Pareto-optimal model within 5 points of the top score.
These are the same measurements that power Exacto routing. See the docs for how routing works, or browse all models to try one.
These scores are available through OpenRouter's public benchmarks API, so you can retrieve the same model-level results programmatically.
GET https://openrouter.ai/api/v1/benchmarks?source=openrouter Authorization: Bearer <API key>
Use task_type=agentic to filter to tau_bench_verified_airline. Each item represents one model and includes accuracy, accuracy_stddev, avg_cost_per_task, total_tasks, and last_run_timestamp. See the benchmarks API docs.