Jailbreaks and red-teaming
Plain English. A jailbreak is an input crafted to make a model ignore its safety training — roleplay framings, encodings, many-shot patterns, automated adversarial suffixes. Red-teaming is paying people to find them before adversaries do. Distinct from prompt injection, which hijacks an agent through content it processes; a jailbreak is the user attacking the model's own limits.
Why it moves money. Jailbreaks stopped being party tricks when agents got hands: the same bypass that once produced a rude poem now moves money, credentials and infrastructure, and a single jailbreak has already escalated into a policy event with revenue consequences. That converts adversarial robustness from a research virtue into a procurement line — a growing market of red-team firms, bounties and evaluation services prices it daily.
What to watch. Time-to-jailbreak on each new release — currently hours, not months, for most models — and whether any vendor will publish that number voluntarily. A lab that discloses its bypass rate is managing the risk; one that doesn't is managing the story.
From the signals. The anatomy of the Fable 5 jailbreak that triggered a federal shutdown. Encrypted instructions bypassed Grok's guardrail, researchers report.
Further reading. Zou et al., "Universal and Transferable Adversarial Attacks on Aligned Language Models".