Lexicon · Measuring it

Benchmark-maxing and Goodharting

Plain English. Goodhart's law, from 1975: when a measure becomes a target, it ceases to be a good measure. Benchmark-maxing is the AI-industry instance — optimising models for the leaderboard rather than the capability the leaderboard was meant to proxy.

Why it moves money. Benchmarks now allocate capital: they decide which labs raise, which models get bought, which narratives run. That weight makes them targets, and targeted measures degrade — the ACM has run the argument formally against AI leaderboards. But the distinction investors should keep is between spurious gaming and real-but-narrow optimisation: when a harness change moved an ARC-AGI-3 score 36 points, the optimisation was frequently genuine — publishing a benchmark simply recruits the industry to solve that task, which is not the same as general capability rising.

What to watch. The gap between leaderboard rank and field performance; scores quoted from a vendor's own evaluation suite; models delisted or vanishing from leaderboards without explanation.

From the signals. Goodhart comes for the leaderboards, and a model vanishes from one. Tencent claims a narrow lead — on its own eval. The ARC harness gap, and why Goodhart is the wrong reading.

Further reading. Goodhart's law.

← All terms