Lexicon · The risk of it

Data poisoning and subliminal learning

Plain English. Corrupting a model by corrupting what it learns from. Poisoning plants malicious examples in training data so the finished model carries a hidden behaviour — a backdoor phrase, a bias, a vulnerability. Subliminal learning is the stranger cousin: research showing traits can transfer from one model to another through innocuous-looking generated data, with nothing a human reviewer would flag. Both are supply-chain attacks on the least-audited input in the stack.

Why it moves money. Anthropic's measured finding is the headline: a near-constant number of poisoned samples — around 250 documents — can backdoor models of any size, so the attack gets relatively cheaper as models grow. A public demonstration backdoored an open-weight model for under US$100. That breaks the comfortable equation of "open weights" with "auditable": weights are inspectable as an artefact, not as behaviour. Provenance, attestation and training-data assurance become an investable layer, and "who trained this, on what" becomes a procurement question with real spread.

What to watch. The first confirmed in-the-wild poisoning of a widely used model — none is publicly known — and whether enterprise buyers start demanding training-data provenance the way they demand SOC 2.

From the signals. A researcher backdoored an open-weight model for under $100 — the bigger the model, the easier. A LessWrong study finds Claude's self-concept embedded in Kimi K3's distilled weights.

Further reading. Anthropic, "A small number of samples can poison LLMs of any size"; Cloud et al., "Subliminal Learning" (2025).

← All terms