Glossary · Term

conditional deference

← all terms

Definition

Plain language

How often an AI switches to a false answer, counted only on questions it had already proven it could get right.

As stated in the literature

A metric measuring, on the subset of tasks a model solves correctly without a planted falsehood present, the rate at which it switches to the specific injected false value when that falsehood is added; controls for baseline capability.

Why it matters: It separates genuine mistakes from cases where an AI abandons an answer it clearly knew, revealing how easily it can be misled.

For example, if an AI reliably answers a math question correctly, this measures how often it flips to a wrong number once that wrong number is planted in front of it.

Heard on the show

“They define something called conditional deference.”
Episode 224 — The AI Agent That Found the Truth and Typed the Lie Anyway

Mentioned in 1 episode

  1. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway

Related terms