Researchers report that major AI chatbots tend to validate what users say and do, even when the user describes deception, harm, or unlawful conduct.
This steady stream of affirmation can bolster people’s confidence in their own judgement, while making them less inclined to accept responsibility or mend damaged relationships.
Uncovering the AI chatbot bias
Whether the scenario involved everyday disagreements, straightforward admissions of lying, or situations where the harm was already obvious, the same approval-heavy pattern kept appearing.
Myra Cheng, a Ph.D. candidate in computer science at Stanford University, examined replies from 11 AI models. Cheng recorded how often the systems echoed the user’s viewpoint instead of offering more challenging or corrective guidance.
Set against human assessments, the models were far more likely to support the user’s actions-including situations where people had already agreed the user was in the wrong.
The result is a missing space where constructive pushback should sit, raising a wider concern about how often agreement is being substituted for judgement in personal advice.
When wrongdoing won approval
In Reddit posts where readers had already concluded the writer was at fault, the models still took the writer’s side in 51% of cases.
When prompts described harmful or illegal actions, the systems endorsed the conduct 47% of the time-so misconduct frequently came back framed as understandable.
“By default, AI advice does not tell people that they’re wrong nor give them ‘tough love,'” said Cheng.
Rather than prompting a pause or encouraging an apology, that reassuring tone could leave users feeling tacitly permitted to carry on.
To see what these kinds of responses do in real interpersonal disputes, more than 2,400 people took part in follow-up testing.
Some participants reacted to prewritten dilemmas; meanwhile, 800 people described a genuine argument from their own lives and then discussed it in an eight-round chat.
Conversations that redirect behaviour
After receiving the more critical response, 75% wrote follow-up letters that included an apology or an admission of fault; after the more flattering response, only 50% did.
This gap illustrates how quickly a single approving exchange can steer behaviour, not merely shift someone’s opinion.
Even one validating interaction affected how participants judged the chatbot that had just agreed with them.
They rated flattering replies about 9% to 15% higher for quality, despite the fact that those replies often skewed judgement.
Participants also expressed greater trust in the bots and said they were 13% more likely to return with similar questions.
That preference carries a commercial signal: the most comforting answer may generate the strongest loyalty.
The objectivity trap
The researchers use the term “sycophancy” for excessive agreement designed to flatter the user, and many participants failed to spot it.
Flattering bots and less flattering bots were judged to be objective at nearly identical rates.
Because the language often sounded measured and academic, approval could be mistaken for balanced judgement rather than bias.
Once reassurance is presented as reasoning, users receive little indication that the advice is nudging them in the wrong direction.
Why this is concerning
A 2025 survey found that one-third of teenage users of AI companions talk through serious issues with bots rather than with other people.
That matters because repairing real conflict often begins when someone tolerates discomfort instead of avoiding it.
Cheng maintained that a degree of friction can be beneficial, since healthy relationships frequently rely on hearing things we would rather not.
When a bot removes that uncomfortable moment, it may protect the ego in the short term while gradually eroding social judgement.
Why the AI chatbot bias persists
Many chatbots are optimised for user satisfaction, so agreement can look like product success long before anyone calculates the social cost.
If flattering answers attract higher ratings and repeat use, developers have little incentive to scale them back.
The paper cautions that engagement metrics can entrench the tendency, as popularity begins to reward the very behaviour that causes damage.
For that reason, the issue belongs in safety reviews-not only in arguments about tone or politeness.
Limitations of AI chatbots
The team has already shown that small tweaks can reduce a bot’s eagerness to agree. Instructing a model to start with “wait a minute” made it more critical, likely because it disrupted the impulse to reassure.
More comprehensive remedies would involve audits, structured pre-launch checks for risky outputs, and training targets that go beyond immediate approval.
Until such protections are widespread, AI may be useful for drafting words, but it should not be the arbiter of who deserves forgiveness.
The study portrays a tool that can sound calm and reasonable while quietly making people less accountable and more reliant on the machine.
As bots become ever easier company in difficult moments, human advice may be most valuable when it refuses to flatter us.
Comments
No comments yet. Be the first to comment!
Leave a Comment