Lexicon · Capability & training

Test-time compute

Plain English. Spending more computation when the model answers, not when it is trained — letting it think longer, try multiple approaches, or check its own work. The mechanism behind "reasoning" settings, and the third axis of scaling after model size and data.

Why it moves money. It moves cost from training capex to inference opex. A model that thinks for thousands of tokens before answering multiplies the compute — and the bill — per query, which is bullish for compute demand and awkward for anyone selling flat-rate access. It also makes capability a dial rather than a fixed property: the same model at different thinking budgets is effectively different products at different prices, and benchmark scores quietly depend on how much thinking the tester paid for.

What to watch. Tokens consumed per task, and whether extra thinking keeps buying accuracy or hits diminishing returns per domain. Score claims that don't disclose the reasoning budget are marketing.

From the signals. Harness choice moved Astra's ARC-AGI-3 score by 36 points at max reasoning. GLM-5.3 scored 60 on Artificial Analysis and paid for it in tokens. Kimi K3 thinks in code — and that is why token demand goes up.

Further reading. Snell et al., Scaling LLM Test-Time Compute Optimally (2024)

← All terms