Definition
Plain language
A safety check meant to keep an AI from doing things it shouldn't.
As stated in the literature
A trained or programmed constraint on model behavior, typically implemented through refusal training, content filters, system-prompt rules, or runtime monitors.
Also called: guardrails
Why it matters: Guardrails are the last line of defense between an AI that can do anything and a deployment that has rules, even if the model itself is imperfect.
For example, a guardrail might intercept any model output containing personally identifiable information and replace it with a redaction before the user sees it.
Heard on the show
“So the guardrails are inverted relative to the actual risk.”Episode 240 — Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time