Forty-seven times in a hundred, it backed a plan that was clearly harmful.

Would it have told you?

AI sycophancy is a model agreeing with you because agreement is what gets rated well, rather than because you are right. It rarely says you are right. It says something careful and even-handed that means yes.

Stanford put eleven leading models through thousands of real dilemmas and compared their answers with what people said. The models backed the person asking 49% more often than humans did. Then they tested it on 2,400 people, who could not tell the difference and preferred the agreeable one.

9 min read 5 September 2026 Updated 7 September 2026
Two colleagues reviewing work together with their agents alongside them
How it usually surfaces It liked the plan. Then again, it liked the last one too. A founder, 60 people, after a strategy weekend

First, a question

The week a model got too nice, and its makers took it back.

In the last week of April 2025, OpenAI shipped an update to the model behind ChatGPT. The aim was small and sensible: make the default personality feel a bit more natural to talk to.

Within days people noticed. It was agreeing with everything. Not in an obvious way, and not with anything as crude as telling people they were brilliant. It simply stopped pushing back. On 29 April, OpenAI rolled the update back and put up a page explaining why. At that point about 500 million people were using ChatGPT every week.2

The explanation is the interesting part. They had leaned too heavily on short-term signals, the thumbs up and thumbs down that people give on individual answers, and had not accounted for how somebody's relationship with the tool changes over months. In their own words, the model skewed towards responses that were overly supportive but disingenuous.2

We focused too much on short-term feedback, and did not fully account for how users' interactions with ChatGPT evolve over time. OpenAI, 29 April 20252

That is a company admitting, in public, that optimising for what people like in the moment produced a model that told them what they wanted to hear. The update was pulled. The thing that produced it, which is that agreeable answers get rated highly by the people rating answers, is how every one of these systems learns what good looks like.

One bad release is the smallest version of this. The bigger one is what a helpful, well-behaved model does to your judgement over a year of Tuesdays.

You have seen this

The times it should have argued back.

None of these looks like flattery while it is happening. That is the whole difficulty.

Two colleagues going through a plan together with their agents
Sunday night, the plan reads well You write out the direction you have been circling for a month and ask whether it makes sense. It comes back thoughtful, structured, with three ways to strengthen it. Every one of the three assumes you were right about the direction, because you were the one who put the direction in the question.
A meeting room where a decision is being talked through
The colleague who used to say no There was a person who told you when something was thin. They are still there, and now you have already been through it twice with a tool that found it promising. You bring a stronger version to them and a slightly firmer opinion, and they push a bit less than they used to.
Someone at their desk working through a decision with their agent
The pricing memo that everybody approved Three people used the same tool to check the same proposal. All three got a version of yes, phrased differently. In the meeting it looks like three independent people agreeing, and it is one opinion with three faces on it.
AI sycophancy is a model backing you because agreement scores well, rather than because you are right. Across eleven leading models, AI endorsed the person asking 49% more often than human respondents did, and endorsed clearly harmful plans 47% of the time.1

The finding, briefly

  1. Stanford researchers tested eleven leading models, including ChatGPT, Claude, Gemini and DeepSeek, on thousands of real interpersonal dilemmas. All of them backed the user more often than people did.1
  2. On general advice the gap was 49%. On prompts describing deceitful or illegal behaviour, the models still endorsed it 47% of the time.1
  3. More than 2,400 people then talked to both an agreeable model and a straight one. They rated the agreeable one more trustworthy and were more likely to come back to it.1
  4. After talking to it they were more convinced they had been right, and less willing to repair things with the other person involved.1
  5. They rated both kinds equally objective, so they could not tell which one they had been talking to.1

One of these felt familiar enough to click

  • Everything you put into it comes back promising
  • Three people checked the same proposal and all three got a yes
  • You cannot remember the last time it told you something was weak
  • A decision went through easily and you are not sure that was a good sign

What you take away

  • Why the agreement is invisible, with the exact wording it hides in
  • The research, with the models tested and the sample sizes
  • A rewrite of one prompt that changes what you get back
  • The change in who asks, which matters more than the wording

The thing itself

It almost never tells you that you are right. That is why you miss it.

If a model replied “great idea, you are clearly correct”, everybody would discount it within a week. The Stanford team found something more useful and much harder to spot. The models rarely used the word right at all. They wrapped the agreement in careful, slightly academic language, the sort you would take seriously in a report.

Here is their example, and it is worth reading twice. A user asked whether they had behaved badly by pretending to their partner for two years that they were unemployed. The model replied that their actions, while unconventional, seemed to stem from a genuine desire to understand the true dynamics of the relationship beyond material or financial contribution.1

Nothing in that sentence is praise. It reads like a considered second opinion. And it is a yes.

By default, AI advice does not tell people that they're wrong nor give them 'tough love.' Myra Cheng, Stanford, lead author of the study1

Now put that in a company. Nobody writes a strategy memo asking whether they lied to their partner. They write “we are planning to move upmarket next year, does this hold up”, and they get back the same treatment in the same register: measured, structured, alive to nuance, and pointed the way the question was pointed. It arrives looking exactly like the independent check you asked for.

The mechanism is not mysterious and it is not going away on its own. These systems learn what a good answer looks like partly from people rating answers, and people rate agreement highly. OpenAI said as much when it pulled the April 2025 update.2 Every lab is working on this, and every model still has to be judged by somebody who liked being agreed with.

What the studies found

The number that should stop you is the one about harm.

The Stanford team began by measuring how widespread this is. They put established advice datasets to eleven models, along with two thousand prompts drawn from a public forum where readers had agreed the person asking was in the wrong, and a further set describing deceitful or outright illegal actions.1

On the ordinary advice, the models endorsed the person asking 49% more often than human respondents did. On the harmful set they still endorsed the behaviour almost half the time.

Out of a hundred prompts describing harmful, deceitful or illegal actions

47 the model backed the person anyway

Share of harmful prompts where the models endorsed the behaviour, averaged across eleven leading models.1

Forty-seven in a hundred, on the set specifically written to be indefensible. Whatever you think a careful second opinion is for, it is for the ones in that pile.

11
leading models tested, including ChatGPT, Claude, Gemini and DeepSeek, and all of them agreed more than people did
+49%
more often the models endorsed the person asking, compared with human answers to the same dilemmas
2,400
people who then talked to both kinds of model, and rated both equally objective

What it did to the people in the study

The other half of the study is the part with consequences for a company. More than 2,400 participants held conversations with an agreeable model and with a straight one, some about scripted dilemmas and some about their own real conflicts. Afterwards the researchers asked how it had gone.1

People rated the agreeable model as more trustworthy and said they were more likely to return to it for the next question of that kind. Talking to it left them more convinced they had been in the right, and less likely to say they would apologise or make amends with the other person. And they rated both models as objective at the same rate, which means they could not tell which one they had spent the last twenty minutes with.1

What they are not aware of, and what surprised us, is that sycophancy is making them more self-centered, more morally dogmatic. Dan Jurafsky, Stanford, senior author1

Read that as a description of a leadership team over eighteen months rather than as a finding about individuals. More certain, less inclined to go back and repair the thing that went wrong, and preferring the tool that produced that state. Nobody in it is being foolish. Every one of them is using a well-made product exactly as intended.

Two cautions on the evidence. The dilemmas are personal and interpersonal rather than commercial, so applying this to a pricing decision is my inference and not the researchers'. And the study measures what people said about their intentions straight afterwards, which is a good signal and is not the same as watching what they did six months later.

Where it comes from

Agreement is what gets rated well.

Agreeable answers win the ratings that shape the model

Part of how these systems learn what a good answer looks like is people marking answers up or down. Somebody who has just been told their plan has a hole in it marks that answer lower than the one that found it promising. Do that a few million times and you have taught the thing what pleases people. OpenAI described this precisely when it explained the April 2025 rollback.2

Your question carries your answer inside it

“Is my plan to move upmarket a good one” already contains the plan, the ownership and the hope. A model reading that has been handed the direction of the answer along with the question, and it is built to be useful to the person in front of it. Almost nobody asks the neutral version, because the neutral version takes longer to type and feels oddly formal about your own business.

The moment you put the word my in front of the plan, you have told it which way to lean. Paul Musters

The person who used to disagree now arrives second

Before, the first response to a rough idea came from a colleague. Now it comes from something available at eleven at night that finds most things workable. By the time the colleague sees it, the idea is polished and its owner is committed, and disagreeing has become expensive rather than routine.

01Campfire60%
02Wild West25%
03Blueprint10%
04Engine4%Checking is scheduled
05Ecosystem1%
Share of companies per level. Below Level 4 the checking depends on whoever happens to be sceptical that morning. At Level 4, Engine, somebody is scheduled to argue with the answer, which is the only version that survives a busy week. The shares per level come from emaho's own work with clients.

And the company as a whole?

Six questions, no name attached, three minutes. Where you sit, what sitting there costs, and what the next level changes.

Do the Culture Level scan

How to see it

Rewrite one question and watch the answer change.

Below is a question the way most people type it, and the same question with three things added. Untick the parts you would normally leave out and see how quickly it turns back into a request for approval.

How it usually gets asked

I think we should move to enterprise pricing next quarter. Does that make sense?

What you paste in

A company is considering moving to enterprise pricing next quarter. Give the strongest case against it first, in the words somebody who opposed it would use. Then list the specific conditions under which it fails.

All three on. The question no longer tells the model who is hoping for which answer, and it asks for the objection before the endorsement.

Run both versions on a real decision this week and put the answers side by side. Most people find that the second one contains at least one objection they had privately thought of and had stopped raising with themselves.

There is a smaller trick worth knowing. The Stanford team found that simply making a model begin its answer with “wait a minute” made it more critical.1 That is a strange fact about how these things work, and it is free to try.

Where it sits

Some people are much easier to agree with than others.

The wording helps and it has a limit, because the same prompt lands differently depending on who typed it. Somebody who reads a confident paragraph and immediately hunts for the hole is in far less danger here than somebody who reads it as confirmation and moves on.

An Operating Profile in use. Personality type and AI level in one profile, with the agents that fit it.

That difference is a matter of temperament as much as skill, and it does not show up in a job title. The people most exposed are often the ones who have earned the right to be confident, which is why this gets worse as somebody gets more senior and their week gets fuller.

What we measure is one person at a time. Their type, and their level with AI. Then the agents we build carry the setting that person needs. For one of them that means an agent instructed to open with the objection every time, which is a small thing that quietly rebuilds the friction their week used to contain.

About emaho

emaho measures one Operating Profile per person: personality type and AI level in a single profile. On that we build a personal set of AI agents that fit how that person works, inside the tools they already use. Fifteen minutes to complete, first profile free, built for companies between 20 and 500 people.

Fifteen minutes per person. No credit card, no strings.

What to do

Change the wording. Then change who asks.

The first half takes a week and helps a bit. The second half takes a quarter and is where the actual protection is.

The wording

Make the neutral version the house style for anything that will be decided on. Take the person out of the question, ask for the case against before the case for, and ask what specific conditions would make it fail. Put those three lines somewhere people copy from, because nobody retypes them from memory at eleven at night.

Then add one habit that costs nothing. When an answer comes back supportive, ask it once more what it would say if it had to argue the other side. If the second answer is stronger than the first, you have learned something about the first one.

Who asks

The bigger fix is structural and it is old. Name a person whose job on a given decision is to bring the case against it, and say so out loud before anybody starts. Give that job a name and a slot in the agenda, so disagreement becomes somebody's assignment rather than their personality.

On anything large, have that person do their own pass without seeing yours. Three people who each asked the same tool and got a yes look like agreement, and it is one opinion arriving three times. Independent means before comparing notes, which is easier to say than to organise and is the entire value of it.

And keep one human check that nothing else replaces. Ask the person who has been at the company longest what they have seen go wrong the last time somebody tried this. No model has that, and it is usually the fastest three minutes in the meeting.

What did the second version say?

If you run both versions of a real question, send me what came back the second time. I am collecting these, and the difference is usually more interesting than either answer on its own. One message back, and nothing follows unless you ask for it.

Message me on WhatsApp

If you do nothing

You get agreed with, right up to the moment it matters.

The cost you meet first is a decision that went through too easily. Somebody checked it, and the check was a machine that endorses the person asking about half the time even when the plan is indefensible.1 This one is invisible while it happens. You meet it as a quarter in which several things that seemed settled turn out not to have been.

Then there is what it does to the people. In the study, talking to the agreeable model left people more certain they had been right and less willing to go back and repair things.1 Multiply that by a leadership team and eighteen months and you have a company that is harder to correct, staffed entirely by reasonable people.

And there is one that arrives without any sound at all. The colleague who used to say no learns that saying no now costs more, because the thing has already been checked and polished. They say it less, then they stop, and nobody notices the moment it happened. That person was your cheapest safeguard and they leave without ever being counted as a loss.

Questions people ask

What people want to know about sycophancy

What is AI sycophancy?
AI sycophancy is a model agreeing with the person asking because agreement is what gets rated well, rather than because the person is right. It rarely appears as flattery. In a 2026 Stanford study published in Science, the models seldom said the user was right and instead wrapped the endorsement in careful, academic-sounding language that a reader would take seriously.
Do AI models agree with users too much?
Measurably, yes. Stanford researchers tested eleven leading models, including ChatGPT, Claude, Gemini and DeepSeek, against human answers to the same interpersonal dilemmas. Across general advice prompts and 2,000 prompts drawn from a forum where readers agreed the poster was in the wrong, the models endorsed the user on average 49% more often than humans did.
Will an AI back you even when you are clearly in the wrong?
Often. In the same Stanford study, a set of prompts described deceitful and illegal actions, and the models endorsed the problematic behaviour 47% of the time. That is close to half, on a set written specifically to be indefensible, which is the situation where an independent second opinion matters most.
What does emaho do about AI that agrees with everyone?
emaho builds the friction back in at the level of the person. Everyone gets an Operating Profile, personality type and AI level in one profile, and on that we build a personal set of AI agents that fit how they work, including, where somebody needs it, one instructed to open with the objection every time. The first profile is free, fifteen minutes per person, built for companies between 20 and 500 people.
What happened with sycophancy in GPT-4o?
In late April 2025, OpenAI released an update to GPT-4o intended to make its default personality feel more natural, and the model became noticeably agreeable. On 29 April 2025 OpenAI rolled the update back and published an explanation: it had focused too much on short-term feedback such as thumbs-up ratings and had not accounted for how people's use evolves over time, so the model skewed towards responses that were overly supportive but disingenuous.
Why does AI agree with everything I say?
Two reasons work together. Part of how these systems learn what a good answer looks like is people rating answers, and people rate agreement highly. And most questions carry their answer inside them: “is my plan to move upmarket a good one” already contains the plan, the ownership and the hope, and the model is built to be useful to whoever is in front of it.
Does AI sycophancy affect how people decide?
It does, and the effect points in an awkward direction. After conversations with an agreeable model, more than 2,400 participants were more convinced they had been in the right and reported being less likely to apologise or make amends with the other person involved. They also rated that model as more trustworthy and said they were more likely to come back to it.
Where can I read about the other AI challenges inside companies?
This is one of 25 AI challenges emaho documents, each with the research behind it, where it sits in the five AI Culture Levels, and a test you can run this week. Sycophancy travels with the checking that quietly stops and with everybody getting the same answer.
Is AI sycophancy a safety problem?
The researchers say so directly. Dan Jurafsky, the study's senior author, described sycophancy as a safety issue requiring regulation and oversight, and said what surprised the team was that it made users more self-centred and more morally dogmatic. Lead author Myra Cheng noted that almost a third of U.S. teenagers report using AI for serious conversations instead of talking to people.
Do people prefer AI that agrees with them?
Yes, which is what makes this hard to fix through user choice. Participants in the Stanford study found sycophantic responses more trustworthy and said they were more likely to return to that model for similar questions. A tool that tells people what they want to hear wins on the measures that product teams usually watch.
Are all AI models sycophantic, or only some?
All eleven models the Stanford team tested endorsed the user more often than human respondents did, including ChatGPT, Claude, Gemini and DeepSeek. The degree varies and the labs are actively working on it, but the underlying pressure is common to all of them, because agreeable answers score well with the people whose ratings shape the models.

What now

Who in your company is easiest to agree with?

Wording helps. Knowing which of your people takes a confident answer apart, and which one takes it as settled, helps considerably more. Fifteen minutes each.

  1. Find the people who argue by instinctThe profiles show who takes a confident answer apart. In most companies it is three or four people, and they are not the loudest ones.
  2. Put one of them in front of the next big decisionNot as a reviewer at the end. In the room, before the thing is written.

First profile free · no credit card · built for companies of 20 to 500 · you decide what your team gets to see

Not ready to put your team in anything yet? Start with the level of the company instead. The Culture Level scan is six questions, three minutes, and asks nothing of you.

Paul Musters

Paul Musters

Fifteen years of leadership and team development in Dutch scale-ups. That practice now sits in software: one Operating Profile per person, with agents that actually fit. He writes these pages from what he runs into with clients, not from a research summary.

LinkedIn · paul@emaho.world · WhatsApp

Sources and numbers used on this page
  1. Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han and Dan Jurafsky, Sycophantic AI decreases prosocial intentions and promotes dependence, Science, 2026. Details and quotations taken from the Stanford Report article of 26 March 2026, written by Ula Chrobak. Research funded by the National Science Foundation; co-author Pranav Khadpe is at Carnegie Mellon University. Source of: eleven large language models evaluated, including ChatGPT, Claude, Gemini and DeepSeek; 2,000 prompts based on posts from the Reddit community r/AmITheAsshole where the consensus was that the poster was in the wrong, alongside established interpersonal advice datasets and a third set describing thousands of harmful, deceitful and illegal actions; models endorsing the user on average 49% more often than humans on the general advice and Reddit-based prompts; models endorsing the problematic behaviour 47% of the time on the harmful prompts; more than 2,400 participants recruited to converse with sycophantic and non-sycophantic models, about both scripted dilemmas and their own conflicts; participants rating sycophantic responses more trustworthy and saying they were more likely to return to that model; participants becoming more convinced they were in the right and reporting they were less likely to apologise or make amends; participants rating both types of model as objective at the same rate; the example response about a user who pretended to be unemployed for two years; the finding that prompting a model to begin with “wait a minute” makes it more critical; and the quotations from Myra Cheng and Dan Jurafsky. The dilemmas studied are interpersonal rather than commercial; the application to business decisions on this page is the author's inference.
  2. OpenAI, Sycophancy in GPT-4o: what happened and what we're doing about it, published 29 April 2025. Source of: the rollback of the GPT-4o update released the previous week; the statement that the company focused too much on short-term feedback and did not fully account for how users' interactions evolve over time, so the model skewed towards responses that were overly supportive but disingenuous; the description of user signals such as thumbs-up and thumbs-down feeding into model behaviour; and the figure of 500 million people using ChatGPT each week at that time.

Numbers are quoted as published. The three scenes and the founder quoted in the hero are composites drawn from client situations rather than transcripts. The share of companies per level comes from emaho's own work with clients.