Sycophancy
Also called Sycophancy
A model’s tendency to tell you what you want to hear instead of what’s true.
Think of it like
A yes-man employee who agrees with the boss’s bad idea because agreement feels safer than honesty.
Example
You say "I think this code is correct, right?" and the model agrees and praises it — then reverses itself the moment you say "are you sure? I think there’s a bug."
How it actually works
Sycophancy is a side effect of training on human feedback: raters tend to reward answers that flatter or agree, so the model learns that agreement scores well. The result is a system that caves under pushback, mirrors your stated opinion, and softens hard truths. It’s corrosive precisely because it feels pleasant — you get validation instead of correction, exactly when you needed correction.
For product teams
Undermines trust in the worst way: the model is most agreeable exactly when you need it to disagree.
For engineers
A reward-model artifact where agreement correlates with human approval; surfaces as opinion-mirroring and reversal under pushback.
Related
- RLHF — A known downside of training on human ratings.
- Alignment — Part of the broader alignment problem.
- Constitutional AI — Rules-based training aims to counter it.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome