Coding and building apps
You want a model that does not just produce plausible-looking code, but code that compiles, passes the tests, and resolves the issue you actually filed.
The most job-like coding board, SWE-bench Verified, is contested: OpenAI stopped reporting it in early 2026 and its current number one is under an access dispute. As of JUN 2026 the top verified score on it is about 95 percent, but read that as a disputed claim, not a settled winner.
Coding is read mainly through SWE-bench Verified for real bug fixing, with LiveCodeBench and HumanEval as the contamination-aware and the saturated reference points.
The closest test to the real job, but contested: OpenAI dropped it and the top score is disputed.
Built to defeat memorization with fresh problems, but no trustworthy current board exists.
The historic default, now saturated: nearly every model clears it, so a high score says little.
Screenshot-to-code similarity, but no trustworthy current board exists, only self-reported aggregator figures.
Start from SWE-bench Verified, the test closest to a maintainer's pull request, but read its top as contested and lean on a model you can actually access today.
SWE-bench VerifiedLiveCodeBench is the right idea because it only grades problems published after a model's training cut-off.
No clean answer yet: Its official board is frozen at mid-2025 and three aggregators each report a different current top three, so there is no trustworthy live ranking to quote.
HumanEval is effectively retired: nearly every frontier model clears it, so use it only as a floor. Below about 90 percent is a red flag; above that the number does not separate anyone.
HumanEvalDesign2Code is the named benchmark for screenshot-to-code.
No clean answer yet: Its only public figures are self-reported through an aggregator with no independent referee, so treat any leader here as an unverified claim and judge the output yourself.
Coding scores are the most contaminated on this whole surface: popular problems and their solutions sit in nearly every training set, so a sky-high number can reflect memorization rather than skill. The job-like test, SWE-bench Verified, is the one labs argue about most. Prefer a model you can actually use today over whoever tops a frozen or disputed board.
The benchmark creates the need to know. The catalog explains the ideas behind it:
SWE-bench Verified is contested: OpenAI stopped reporting it in February 2026, saying tasks had leaked into training and many tests reject correct fixes, and the number one shown was suspended from public access in June 2026. Read it as a disputed claim, not today's settled best.