Sparse autoencoders, probes and steering
Plain English. The working toolkit behind interpretability claims. A sparse autoencoder (SAE) decomposes a model's internal activations into individually meaningful "features" — concepts the model is representing. Probes are lightweight classifiers that read those internals to detect a state (deception, uncertainty, a topic). Steering writes to them: amplifying or suppressing a feature to change behaviour without retraining, the technique behind Anthropic's famous demo of a Claude obsessed with the Golden Gate Bridge.
Why it moves money. This is what "we can see inside the model" concretely means, so it is the substance behind every interpretability startup pitch and lab safety claim. Probes are the plausible near-term product — cheap runtime detectors for lying, jailbreak states or data leakage — and steering hints at behaviour control as a feature, not a fine-tune. The open question is coverage: SAE features explain a fraction of what models do, and the fraction is contested.
What to watch. Whether probe-based detectors ship in production safety stacks with disclosed error rates, and whether SAE findings replicate across labs rather than living in single-lab demos. Adoption by a second lab is the tell that the toolkit generalises.
Further reading. Anthropic, "Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet".