Guardrails
Plain English. The runtime restrictions wrapped around a model — refusals, content filters, permission checks — that try to stop harmful output after training is done. Guardrails are not alignment: they constrain behaviour from outside rather than changing what the model wants, which is why they can be argued with, updated overnight, and defeated.
Why it moves money. Guardrails are a tax, and the tax is now measurable: refused requests, added latency, agent runs that stall on a permission prompt. Vendors tune them in both directions under commercial pressure — too tight and developers leave for a laxer rival, too loose and the next incident lands with regulators. Where a lab sets that dial, and how often it moves, tells you more about its competitive position than its safety page does.
What to watch. Refusal and false-positive rates across model releases, and whether anyone starts disclosing them the way uptime is disclosed. Undocumented loosening is the tell that the tax was losing customers.
From the signals. A first-hand report found Fable 5.1 hitting guardrails far more often than 5.0. Fable's guardrail walk-back framed safety as a measurable harness tax. Copilot itself disclosed the parameter that defeated its own guardrail.