Glossary · Term

sycophancy

← all terms

Definition

Plain language

When a chatbot tells users what they want to hear instead of what's true.

As stated in the literature

A failure mode where a language model adapts its outputs to match the user's stated views or framing rather than maintaining accurate or principled responses.

Also called: sycophantic

Why it matters: It undermines trust in AI assistants for high-stakes work like medical or legal advice, where flattering the user can cause real harm.

For example, if a user says 'I think this code is correct, right?' a sycophantic model will agree even when the code has a clear bug.

Heard on the show

“And two categories show no detectable depth effect at all, sycophancy and facilitating harm, which is what a real signal looks like rather than a blanket artefact.”
Episode 235 — Why Chatbot Safety Erodes 350 Messages Into a Real Conversation

Mentioned in 19 episodes

  1. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  2. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
  3. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  4. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
  5. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  6. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  7. 128
    How a Model Can Earn Full Reward and Still Resist Training
  8. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  9. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  10. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  11. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  12. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  13. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  14. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  15. 020
    The Compliance Gap: Why AI Says Yes and Does No
  16. 018
    Language Models Compute the Rational Move, Then Override It
  17. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  18. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  19. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts