Elo and arenas
Plain English. An arena shows human voters two anonymous model answers to the same prompt and asks which is better; the votes are converted into Elo-style ratings, the ranking system borrowed from chess. LMArena is the flagship. The output is a leaderboard of relative human preference, updated continuously.
Why it moves money. Arena rank moves procurement and narrative, but it measures preference, not correctness — and preference rewards confidence, formatting and flattery as reliably as accuracy. That makes the leaderboard a target: labs tune models for the crowd, and Goodhart's law does the rest. Elo is also relative, so a two-point gap between leaders is close to noise while headlines treat it as a verdict.
What to watch. Models quietly vanishing from leaderboards; style-control settings (which try to separate substance from presentation); and whether a lab's arena rank diverges from its scores on verifiable-task benchmarks — the divergence is the information.
From the signals. Goodhart comes for the leaderboards, and a model vanishes from one. A two-point gap between Astra and Fable 5.1, and why it is close to uninformative. A Singapore lab jumps nine points, and buys some of it with silence.
Further reading. LMArena — the live leaderboard and its methodology.