Lexicon · Measuring it

The benchmark alphabet

Plain English. The names in every launch deck, decoded. SWE-bench — real GitHub issues; the Verified variant was retired on contamination, succeeded by Pro. ARC-AGI — novel abstract puzzles designed to resist memorisation, scored against semi-private sets. LMArena — human preference votes converted to Elo ratings. Artificial Analysis Intelligence Index — a composite of roughly nine evaluations, including Terminal-Bench (agent tasks in a terminal), GPQA Diamond (graduate-level science questions) and Humanity's Last Exam (frontier academic questions). FrontierMath — research-grade mathematics, run by Epoch AI.

Why it moves money. Each name is a different instrument measuring a different thing — preference, coding, abstraction, recall — and deck-writers pick the flattering one. Composites bring their own trap: the index version changes between readings, so score movement partly reflects the ruler, not the model. And every instrument on the list decays (see benchmark saturation).

What to watch. Which benchmark a launch cites and which it omits; whether the quoted number is comparable across the version change.

From the signals. Astra debuts at 61 on a nine-evaluation composite. Index v4.3: a tie at 53, on a changed ruler. SWE-bench Verified retired.

Further reading. swebench.com · arcprize.org.

← All terms