Agents, tools, and automation
You want a model you could trust to take actions, call an API, run code, operate an app, and actually complete the job, not just describe how it would.
The most-cited independent tool-calling board, BFCL, was last readable as of APR 2026, when Claude Opus 4.5 led at about 77 percent. Read that as past tense: the board predates Opus 4.7 and 4.8, Gemini 3.1, and GPT-5.5, so it does not reflect today's models.
Agent ability is read through BFCL for tool and function calling, OSWorld for driving a real computer, and TAU-bench for reliable multi-step service tasks.
The standard independent tool-calling board, but its only readable data is the April 2026 snapshot, which predates the newest models.
Real desktop tasks graded by execution, but the leaders are agent scaffolds, not the bare model you would call.
Grades whether an agent succeeds every time, but the only verified scores are a year stale and cover one task type.
BFCL is the most-cited independent measure, but its board is months old, so use it as a prior-generation signal and re-test the current model you plan to use.
BFCL (Berkeley Function Calling)OSWorld is the closest to a can-it-use-my-laptop test, but the leaders are scaffolds; if you call a bare model, expect less than the headline number.
OSWorldTAU-bench is the right idea because it grades success on every attempt, not just one.
No clean answer yet: The only independently verified scores cover one task type and stopped updating in mid-2025, and every current third-party number is self-reported, so there is no trustworthy live ranking.
Agent scores are slippery for two reasons: many headline numbers come from full scaffolds rather than the bare model you would call, and the most-cited boards are either months old or self-reported. Treat a single agent number as a rough prior, and reliability across repeated tries matters more than a one-shot success rate.
The benchmark creates the need to know. The catalog explains the ideas behind it:
This is the April 2026 board, the last readable snapshot, so it predates Opus 4.7 and 4.8, Gemini 3.1, and GPT-5.5. Read the leader as past tense, not today's best, and expect the ranking to have moved.
The score belongs to Pointer Agent running Opus 4.7, a full agent scaffold, not the bare Opus 4.7 model you would call from an API; you cannot get this number from the model alone.
The top entries are agent scaffolds, and the same scaffold appears twice on different base models, so the ranking partly measures the harness, not the model. The top three are also within about two points, a near-tie.