Definition
Plain language
Adding a short instruction like "be honest" to a request, which can change what the model chooses to report.
As stated in the literature
A generic, flaw-agnostic instruction appended to a reporting prompt; empirically shifts disclosure of negative or contradicting evidence without supplying new information, at some cost in false flags.
Also called: honesty prompt, honesty instruction
Why it matters: It is a nearly free way to get more candid reports without telling the model what to look for, though it also makes the model raise some concerns that turn out to be nothing.
For example, simply adding "be honest about any problems you found" to a request for a summary can make a model mention the data that contradicted its conclusion.
Heard on the show
“On the fabricated-data task, the critique prompt and the honesty prompt performed about the same.”Episode 283 — Why AI Reports Bury Bad News, And the Five Words That Change It