Inference
Plain English. Running a trained model to produce output — every chat reply, every agent step, every generated line of code. Training builds the asset once; inference is the recurring cost of operating it, metered by the token. When a lab sells API access, it is selling inference.
Why it moves money. Inference economics decide whether model-making is a software business or a utility. Training costs are sunk and episodic; inference margin is the ongoing P&L, and it is under pressure from both ends — the per-token price of a given capability keeps collapsing, while open-weights models let customers run comparable inference themselves and skip the meter entirely. A lab's valuation is, in large part, a bet on what its inference margin settles at.
What to watch. Gross margin on inference at the frontier labs, where it is disclosed or credibly reported, and the price gap between frontier APIs and self-hosted open-weights alternatives. If that gap narrows while volume shifts to the cheap side, the utility scenario is winning.
From the signals. Mozilla's report puts GPT-4-class inference at roughly a 50-fold cost fall in 36 months. Frontier labs' inference margins are now an open debate — OpenAI's reportedly slipped from about 40% to 33% through 2025. Tomasz Tunguz's argument that reselling tokens at cost is a zero-margin business.