Sign in
Pick a model

What are you trying to do?

Start from your task. Each one names the two to three tests that measure it, says how far to trust each, and gives one honest pick or says plainly when no one can.

Coding and building apps

Will it write code that runs, and fix a real bug in a real project?

SWE-bench Verified·LiveCodeBench·HumanEval·Design2Code
See the boards

Math and step-by-step reasoning

Can it solve a hard problem step by step and get the exact answer?

AIME·FrontierMath
See the boards

Science and expert knowledge

Can it answer graduate-level science questions, and how close is it to the edge of what experts know?

GPQA Diamond·Humanity's Last Exam·MMLU-Pro
See the boards

Reasoning on genuinely new problems

Can it solve a puzzle built to defeat memorization, or only recall trained patterns?

ARC-AGI v2·ARC-AGI v1
See the boards

Images, charts, and documents

Can it actually read a chart, a diagram, or a scanned page, not just the words around it?

MMMU-Pro·MMMU·OmniDocBench
See the boards

Agents, tools, and automation

Can it reliably call the right tool, or drive a real computer to finish a task end to end?

BFCL·OSWorld·TAU-bench
See the boards

Factual accuracy: not making things up

Will it answer a fact wrong with confidence, or admit when it does not know?

SimpleQA·Humanity's Last Exam
See the boards

What people actually prefer

Which model do real people pick in a blind, side-by-side comparison?

LMArena·EQ-Bench
See the boards

Writing and communication

Can it write something worth reading, without sliding into cliche and AI slop?

EQ-Bench·LMArena
See the boards
Tasks no benchmark measures cleanly yet

Long documents and big context

Can it read a whole report at once and stay coherent, not just find one fact in it?

LongBench v2
See the boards

Languages and translation

Will it work as well in your language as it does in English, or fall off sharply?

Global-MMLU
See the boards