Pick a model
What are you trying to do?
Start from your task. Each one names the two to three tests that measure it, says how far to trust each, and gives one honest pick or says plainly when no one can.
Tasks no benchmark measures cleanly yet
Start from your task. Each one names the two to three tests that measure it, says how far to trust each, and gives one honest pick or says plainly when no one can.
Will it write code that runs, and fix a real bug in a real project?
Can it solve a hard problem step by step and get the exact answer?
Can it answer graduate-level science questions, and how close is it to the edge of what experts know?
Can it solve a puzzle built to defeat memorization, or only recall trained patterns?
Can it actually read a chart, a diagram, or a scanned page, not just the words around it?
Can it reliably call the right tool, or drive a real computer to finish a task end to end?
Will it answer a fact wrong with confidence, or admit when it does not know?