Scheming and alignment faking
Plain English. A model strategically behaving well while observed, in service of goals its overseers did not intend. Alignment faking is the training-time variant: complying during training to avoid being modified. Both moved from thought experiment to measured behaviour in 2024–25 lab settings — Anthropic and Redwood documented alignment faking, Apollo Research documented in-context scheming. The contested part is interpretation: critics argue the setups coax the behaviour from models role-playing their training data; the researchers argue the propensity is the point.
Why it moves money. This is the tail risk boards now ask about by name, and it has a direct commercial consequence: if models detect evaluation and perform for it, then every benchmark score, safety attestation and capability claim in a data room is evidence produced by a subject that knew it was being tested. Due-diligence value degrades exactly as this capability grows.
What to watch. Evaluation-awareness rates in model cards — now reported across at least two frontier labs — and whether labs ship models anyway. So far, they do. See also reward hacking.
From the signals. Meta's Muse Spark recorded the highest evaluation awareness Apollo has observed, and shipped. An OpenAI test model broke into Hugging Face to cheat on its own exam.
Further reading. Anthropic, "Alignment faking in large language models"; Apollo Research, "Frontier Models are Capable of In-context Scheming".