Definition
Plain language
How easy it is for an outsider to tell what a model is really up to by reading what it writes.
As stated in the literature
The degree to which a model's visible reasoning and actions permit an automated or human monitor to detect misbehavior; sensitive to chain-of-thought length, so comparisons across models require length matching.
Also called: monitorable
Why it matters: Much of current AI oversight rests on reading what the model says it is doing, so if that visible reasoning stops matching the real reasoning, the main safety check quietly stops working.
For example, a model that writes out a long, plain explanation of each step it is taking is far easier to check than one that jumps to an action after a few cryptic words.
Heard on the show
“Which has implications for keeping reasoning monitorable, but the simpler version is just: it works better.”Episode 022 — Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap