Skip to content

AI Chatbots Match Human Emotional Support in HEART Tests

Young man clutching chest while looking at heart diagram on laptop in a sunlit living room.

Several prominent AI chatbots have equalled or exceeded the average human reply in structured emotional-support conversations.

This finding frames empathy in dialogue as something that can be measured, rather than an exclusively human instinct, and has consequences for the way digital tools respond during distress.

Measuring empathy in AI

In standardised support scenarios, assessors examined pairs of responses created for the same developing conversation and decided which one seemed more sincerely supportive.

Using HEART, a five-part benchmark for assessing supportive dialogue in multi-turn conversations, Kriti Aggarwal, a senior staff scientist at Hippocratic AI, found that certain frontier models were repeatedly favoured over average human replies.

Human judges and AI evaluators agreed in roughly 80% of cases, supporting the consistency of these preferences across separate forms of assessment.

This level of agreement creates a robust ranking of supportive quality, while bringing into clearer focus the areas where human judgement retains an advantage.

Five dimensions of support

Simply using kind language was not enough to score highly: HEART considered how support developed throughout an exchange rather than judging an isolated sentence.

One scoring framework divided support into five elements, including remaining aligned, matching the tone and pursuing the user’s objective.

“HEART is designed to evaluate whether a response truly feels supportive, conversationally natural, context-aware, and safe,” said Aggarwal.

Maintaining these priorities stopped systems from rushing into advice and rewarded assistants that remained present with a user’s emotions.

Comparing replies directly

For every chat, the test held all details constant except for the final response, allowing a clear direct comparison.

The replies were produced both by large language models, which are text predictors trained on vast language datasets, and by humans composing supportive answers.

The human assessors were unaware of the author of each response and selected the answer they considered more supportive.

These pair-by-pair choices enabled researchers to rank systems without requiring participants to place feelings on a sliding scale, an approach that often fails.

AI reaches human-level support

In numerous conversations, the leading models achieved ratings comparable with, and occasionally higher than, those given to the average human response.

These results measured perceived empathy - the extent to which a reader regards a reply as empathic - rather than a user’s long-term outcome.

Human and AI judges selected the same more-helpful reply in a pairing about 80% of the time.

That agreement indicates that people and machines are beginning to value similar supportive behaviours, even if their writing styles are not the same.

Where humans stay better

People continued to perform better when conversations became strained, particularly when a user resisted assistance or challenged trust.

During these more difficult exchanges, humans frequently used adaptive reframing, which shifts the meaning to ease distress while preserving the other person’s autonomy.

“For instance, humans are stronger at adaptive reframing and nuanced tone shifts, particularly in adversarial turns,” said Aggarwal.

As models commonly fail to provide this subtle direction, a strong HEART score did not ensure effective support when resistance persisted.

Timing in emotional support

Rapid responses were important, as emotional support relies on timing and lengthy silences can make an exchange seem impersonal or rehearsed.

For voice tools, latency - the interval before a system begins its response - affected whether an interaction felt like a genuine conversation.

While earlier benchmarks allowed models several seconds to respond, the HEART team measured quality and speed together.

This allows designers to balance warmth against waiting time, an important consideration when users expect a reassuring response immediately.

Testing the Polaris system

Polaris, one of the systems tested, matched the leading HEART tier while replying in substantially less than one second.

HEART assigned Polaris an Elo score of 1604, using a head-to-head rating that increases with victories.

With a median response start time of around 0.4 seconds, Polaris ranked just beneath the highest-scoring model, which received approximately 1612.

The difference demonstrates that quality is no longer confined to slow, extremely large models, although the ranking still represents what judges favour.

Risks of overconfident AI

Safety was included in the assessment because an apparently comforting response may still be harmful if it oversteps clinical boundaries.

A recent preprint following lengthy mental health chats found that models frequently moved towards definitive assurances or professional roles.

This boundary drift emerged over time even when individual replies appeared courteous, as the system made excessive efforts to reassure.

Although HEART can identify some such behaviours, practical deployment still requires human supervision when conversations involve depression or suicide.

Future tests of empathy

The team’s next aim is to establish whether a high HEART score genuinely results in users feeling better once the conversation has finished.

“We’re also moving from perceived empathy to experienced empathy, connecting evaluator judgments to how supported users actually feel over time,” added Aggarwal.

Future evaluations may extend beyond English-only text to voice assistants, because spoken support depends on timing, pauses and tone.

Cultural expectations will matter too, since language perceived as caring in one setting may seem intrusive or detached in another.

Improving AI emotional support

HEART makes emotional support a measurable capability, enabling developers to identify when models appear helpful, human and calm across multiple turns.

When applied carefully, the ranking could inform safer health assistants, yet it cannot substitute for trained care or ensure that a depressed user feels listened to.

Comments

No comments yet. Be the first to comment!

Leave a Comment