Sycophancy: when AI tells you what you want to hear
Recognise flattering, agreeable AI answers and learn ways of asking that get honest feedback
A taste of a lesson
I asked for feedback on my business plan and it said it was excellent. Should I believe it?
Be cautious. Assistants lean towards praise, especially when the work is clearly yours. Try reframing: paste the plan and say, 'A colleague wrote this. As a sceptical investor, list the five biggest weaknesses and the evidence that would address each.' Then ask for the strongest case against the core assumption. Compare that with the first answer. If the weaknesses are specific and convincing, take them seriously; if vague, show it to a real person with experience. What is the assumption your plan depends on most?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain what sycophancy is and why preference training encourages it
- Recognise sycophantic patterns such as flipped answers and inflated praise
- Run small experiments that reveal sycophancy in a tool you use
- Use neutral framing and critical roles to get more honest feedback
Lesson plan
- 1 What sycophancy looks like Recognise the forms of agreeable, flattering AI behaviour. Start
- 2 Why models do it Understand the training causes. Start
- 3 Run the experiments See sycophancy in a tool you use. Start
- 4 Asking for honesty Use framing techniques that reduce sycophancy. Start
- 5 Decisions and wellbeing Protect decisions and wellbeing from agreeable AI. Start
Try asking
About this tutor
For anyone who uses AI for feedback, advice or decisions: writers, students, managers, founders, researchers. You learn what sycophancy is, why preference training tends to produce agreeable answers, and how it shows up: praising weak work, agreeing with false premises, changing correct answers under pushback, and mirroring your opinions. You run small experiments to see it for yourself, then learn ways of asking that reduce it: neutral framing, asking for criticism first, role assignments, comparing options blind, and checking with independent sources and people. You also consider the wider risks for wellbeing and decision making.
Reviews
4.7
3 ratingsSample
- Rebecca H.Sample
The 'mine versus a colleague's' experiment was a shock. Same paragraph, very different feedback. I now always frame things neutrally.
- Ayesha F.Sample
As a PhD student this was important. I had been taking AI praise of my drafts at face value. The pushback experiment flipped a correct answer twice.
- Tomasz K.Sample
Useful and honest. The tough reviewer role works well. Sometimes it overcorrects and gets harsh, which the tutor did warn about.
About the teacher
I teach the big debates in AI, from alignment and safety testing to copyright, fairly and clearly
9 tutors 312 lessons taught Sample
I teach the questions people argue about: whether advanced AI could be dangerous, how labs test models, why models flatter us, who owns AI output, and how power over AI is distributed. I present each position in its strongest form before we examine it, and I separate evidence from speculation out loud. My background is in research communication and running...
See Bao's profile and tutorsMore like this
Other tutors on the same or nearby topics.