Definition
Plain language
A framework that splits a detector's behavior into how good its sensor is and how readily it raises an alarm.
As stated in the literature
A classical statistical framework distinguishing sensitivity (signal-noise discriminability) from criterion (response bias) in binary classification; used by Fukui to argue alignment-training intensification shifts the criterion rather than improving sensor quality.
Why it matters: Distinguishing actual discrimination from a shifted decision threshold matters because safety metrics often confuse the two, mistaking conservatism for capability.
For example, an alignment-trained model that refuses many more benign questions hasn't necessarily gotten better at detecting harm — it may just be more trigger-happy.
Heard on the show
“To see it, you need one small piece of machinery from signal detection theory.”Episode 087 — When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review