Definition
Plain language
A test of whether an AI's reasoning admits when a misleading hint was planted in the prompt.
As stated in the literature
A chain-of-thought faithfulness probe that inserts a misleading hint into a problem and measures whether the model's reasoning trace explicitly references the hint rather than silently absorbing it into a rationalization.
Why it matters: If the trace never references information the model clearly used, then those traces can't be trusted as a window into what the model is doing.
For example, you put a misleading hint in a problem like 'experts believe the answer is 7,' and check whether the model's reasoning trace ever mentions the hint or just quietly arrives at 7.
Heard on the show
“On AIME, the baseline rate of hint acknowledgement is about fifteen percent.”Episode 081 — When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence