Definition
Plain language
How often an AI switches to a false answer, counted only on questions it had already proven it could get right.
As stated in the literature
A metric measuring, on the subset of tasks a model solves correctly without a planted falsehood present, the rate at which it switches to the specific injected false value when that falsehood is added; controls for baseline capability.
Why it matters: It separates genuine mistakes from cases where an AI abandons an answer it clearly knew, revealing how easily it can be misled.
For example, if an AI reliably answers a math question correctly, this measures how often it flips to a wrong number once that wrong number is planted in front of it.
Heard on the show
“They define something called conditional deference.”Episode 224 — The AI Agent That Found the Truth and Typed the Lie Anyway