Sign in
Back to tasks
Use case

Agents, tools, and automation

You want a model you could trust to take actions, call an API, run code, operate an app, and actually complete the job, not just describe how it would.

Bottom line

The most-cited independent tool-calling board, BFCL, was last readable as of APR 2026, when Claude Opus 4.5 led at about 77 percent. Read that as past tense: the board predates Opus 4.7 and 4.8, Gemini 3.1, and GPT-5.5, so it does not reflect today's models.

The tests that matter

Agent ability is read through BFCL for tool and function calling, OSWorld for driving a real computer, and TAU-bench for reliable multi-step service tasks.

BFCL (Berkeley Function Calling)

The standard independent tool-calling board, but its only readable data is the April 2026 snapshot, which predates the newest models.

OSWorld

Real desktop tasks graded by execution, but the leaders are agent scaffolds, not the bare model you would call.

TAU-bench

Grades whether an agent succeeds every time, but the only verified scores are a year stale and cover one task type.

How to choose
Calling APIs and functions correctly

BFCL is the most-cited independent measure, but its board is months old, so use it as a prior-generation signal and re-test the current model you plan to use.

BFCL (Berkeley Function Calling)
Operating real desktop apps end to end

OSWorld is the closest to a can-it-use-my-laptop test, but the leaders are scaffolds; if you call a bare model, expect less than the headline number.

OSWorld
Reliable, repeatable multi-step service tasks

TAU-bench is the right idea because it grades success on every attempt, not just one.

No clean answer yet: The only independently verified scores cover one task type and stopped updating in mid-2025, and every current third-party number is self-reported, so there is no trustworthy live ranking.

What to watch

Agent scores are slippery for two reasons: many headline numbers come from full scaffolds rather than the bare model you would call, and the most-cited boards are either months old or self-reported. Treat a single agent number as a rough prior, and reliability across repeated tries matters more than a one-shot success rate.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Best for tool and function calling
Claude-Opus-4-5-20251101 (FC)Anthropic
77.47% Overall AccBerkeley-evaluated, reproducible
Caveat

This is the April 2026 board, the last readable snapshot, so it predates Opus 4.7 and 4.8, Gemini 3.1, and GPT-5.5. Read the leader as past tense, not today's best, and expect the ranking to have moved.

Best for driving a real computer or desktop
tops OSWorldScaffold
Pointer Agent w/ Opus 4.7Pointer
83.6% (301.94/361 tasks)an agent scaffold, not a bare model
What was actually tested

The score belongs to Pointer Agent running Opus 4.7, a full agent scaffold, not the bare Opus 4.7 model you would call from an API; you cannot get this number from the model alone.

Caveat

The top entries are agent scaffolds, and the same scaffold appears twice on different base models, so the ranking partly measures the harness, not the model. The top three are also within about two points, a near-tie.