Sign in
Back to tasks
Use case

Coding and building apps

You want a model that does not just produce plausible-looking code, but code that compiles, passes the tests, and resolves the issue you actually filed.

Bottom line

The most job-like coding board, SWE-bench Verified, is contested: OpenAI stopped reporting it in early 2026 and its current number one is under an access dispute. As of JUN 2026 the top verified score on it is about 95 percent, but read that as a disputed claim, not a settled winner.

The tests that matter

Coding is read mainly through SWE-bench Verified for real bug fixing, with LiveCodeBench and HumanEval as the contamination-aware and the saturated reference points.

SWE-bench Verified

The closest test to the real job, but contested: OpenAI dropped it and the top score is disputed.

LiveCodeBench

Built to defeat memorization with fresh problems, but no trustworthy current board exists.

HumanEval

The historic default, now saturated: nearly every model clears it, so a high score says little.

Design2Code

Screenshot-to-code similarity, but no trustworthy current board exists, only self-reported aggregator figures.

How to choose
Fixing a real bug in a real repository

Start from SWE-bench Verified, the test closest to a maintainer's pull request, but read its top as contested and lean on a model you can actually access today.

SWE-bench Verified
Competition-style problems on fresh, unseen prompts

LiveCodeBench is the right idea because it only grades problems published after a model's training cut-off.

No clean answer yet: Its official board is frozen at mid-2025 and three aggregators each report a different current top three, so there is no trustworthy live ranking to quote.

Short, self-contained functions

HumanEval is effectively retired: nearly every frontier model clears it, so use it only as a floor. Below about 90 percent is a red flag; above that the number does not separate anyone.

HumanEval
Turning a design or screenshot into front-end code

Design2Code is the named benchmark for screenshot-to-code.

No clean answer yet: Its only public figures are self-reported through an aggregator with no independent referee, so treat any leader here as an unverified claim and judge the output yourself.

What to watch

Coding scores are the most contaminated on this whole surface: popular problems and their solutions sit in nearly every training set, so a sky-high number can reflect memorization rather than skill. The job-like test, SWE-bench Verified, is the one labs argue about most. Prefer a model you can actually use today over whoever tops a frozen or disputed board.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

SWE-bench Verified
As of JUN 2026
Best for fixing a real bug in an existing codebase
Claude Fable 5Anthropic
Disputed
95.00%vals.ai verified; access since suspended
JUN 2026ยทvals.ai
Caveat

SWE-bench Verified is contested: OpenAI stopped reporting it in February 2026, saying tasks had leaked into training and many tests reject correct fixes, and the number one shown was suspended from public access in June 2026. Read it as a disputed claim, not today's settled best.