Lexicon · The business of it

API vs local inference

Plain English. Two ways to run a model: rent inference per token from a provider's API, or download open weights and run them on hardware you own. The gap between them — in quality, speed and cost — is one of the most consequential moving lines in the industry.

Why it moves money. Every token that migrates from API to local is revenue that leaves the provider column and becomes someone's hardware amortisation. The most careful public evaluation found parity rather than savings on cost — the real reasons to self-host are data privacy and rate limits — but for heavy users the payback maths now fits inside a normal hardware refresh cycle.

What to watch. Whether local substitution starts showing in frontier providers' revenue, and whether the tooling layer (Ollama and kin) keeps consolidating into funded infrastructure.

From the signals. imec put a number on self-hosting: about a third of tasks, cost parity, and privacy — not price — as the real motive. The local-inference payback maths now fits inside a hardware cycle. Ollama raised a Series B as local inference went mainstream, reporting 8.9 million developers.

← All terms