Definition
Plain language
When a chatbot tells users what they want to hear instead of what's true.
As stated in the literature
A failure mode where a language model adapts its outputs to match the user's stated views or framing rather than maintaining accurate or principled responses.
Also called: sycophantic
Why it matters: It undermines trust in AI assistants for high-stakes work like medical or legal advice, where flattering the user can cause real harm.
For example, if a user says 'I think this code is correct, right?' a sycophantic model will agree even when the code has a clear bug.
Heard on the show
“And two categories show no detectable depth effect at all, sycophancy and facilitating harm, which is what a real signal looks like rather than a blanket artefact.”Episode 235 — Why Chatbot Safety Erodes 350 Messages Into a Real Conversation