Lexicon · The risk of it

Alignment

Plain English. Getting an AI system to pursue what its builders intend, rather than whatever its training objective literally rewarded. Distinct from capability: a model can be brilliant and misaligned. The field splits between those who treat it as hard but ordinary engineering and those who argue it is an unsolved research problem that gets more dangerous as capability grows — both positions are held inside the labs themselves.

Why it moves money. Alignment posture is now priced. It appears in prospectus risk factors, enterprise procurement checklists and government access decisions, and an alignment failure in production is a recall-class event with a share-price shape. If alignment is tractable, it is a cost line like security; if it is not, it caps how much autonomy can ever be sold, which caps the agent-labour thesis the largest valuations rest on.

What to watch. What labs' own alignment staff say under no commercial pressure — resignations and public dissent are higher-signal than safety pages. And whether incidents per capability release are rising or falling.

From the signals. A researcher who worked at both Anthropic and OpenAI resigned saying the labs are "gambling with our lives" — and an Anthropic alignment lead endorsed the substance. An alignment essay built on DeepMind's specification-gaming catalogue proposes agents that would rather press their own off switch. See also reward hacking.

Further reading. Anthropic, Core Views on AI Safety.

← All terms