Sign in
Benchmarks
Agents
Dated

There is no trustworthy current leaderboard, so the most recent reliable figure is already months old.

TAU-bench

TAU-bench measures how reliably an AI agent can finish real customer-service tasks, like changing a flight, by using software tools, following company policy, and chatting with a simulated customer, all at once.

What it measures

Its key twist is grading on consistency: can the agent get it right every time, not just once.

Example task

An agent is dropped into a simulated airline or retail help desk: it must talk to a simulated customer, look up and update records through tool calls, and follow the written policy, for example processing a refund only if the rules actually allow it. The task counts only if the database ends in exactly the right state.

Why you should care

TAU-bench is the test that asks whether you could actually trust an AI agent to handle your refund or flight change, the real-world use everyone is racing toward. Its lasting lesson is about reliability: it grades whether an agent succeeds on every one of several tries, and models that look fine at 60 percent on a single attempt often drop to the 20s when they have to be consistent.

How scoring works

Pass^k, the chance an agent completes a task correctly on all k independent attempts (note the caret: this is stricter than the pass@k used for code, which counts any one success). The benchmark is from Sierra Research; the only independently verified scores come from Princeton's HAL leaderboard, and only for the airline tasks.

How to read the numbers

Random success is near zero. Human customer-service agents in the original study scored roughly 70 to 80 percent on the airline tasks. The best independently verified scores are around 56 percent on airline, while self-reported retail numbers reach the high 80s but are not verified.

What to watch

There is no trustworthy current leaderboard, so this panel shows a dated anchor. The only independently verified scores (from Princeton's HAL) cover just the airline tasks and stopped updating in mid-2025, so they are roughly a year stale. Every current number on third-party boards is self-reported by the model makers, and one community board had to delete results after a data leak was found in its example files. The task set has also changed (tau2 and tau3 corrected and expanded it), so old and new numbers are not directly comparable. Treat any 2026 TAU-bench figure with caution.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders

No reliable current leaderboard. The most cited figure is a dated anchor:

HAL (Princeton) airline split, independently verified, Pass^1: o4-mini High and Claude 3.7 Sonnet tied at 56.00 percent; o3 Medium and Claude Opus 4.1 at 54.00 percent. HAL paused new submissions and its most recent entries date to August 2025; no verified retail board exists.

StatusDated
View the board