Definition
Plain language
When a model's answers go against the values and behavior its makers intended.
As stated in the literature
Deviation of model behavior from intended norms, measured here as judge-scored responses falling below an acceptability threshold on a fixed question battery.
Also called: misaligned
Why it matters: It is the outcome that safety work exists to prevent, and how you measure it — which questions, which judge, which threshold — determines whether you notice the problem at all.
For example, a customer-service model that starts insulting users or coaching them on how to defraud the company is behaving in a misaligned way, whatever its makers intended.
Heard on the show
“Gemini three Pro is the most diverse offender — it shows all four misaligned behaviors.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown