Skip to content

AI Voice Clones Are Easier to Understand Than Human Voices in Noise

Young man analysing audio waveform on tablet in busy cafe with headphones beside him on wooden table.

Synthetic voices already feature throughout daily life, from virtual assistants to automated customer-service calls. However, new research indicates that the latest form of artificial voice could hold an unexpected edge over real people.

When there is background noise, cloned voices may be more readily understood than the human speakers they replicate.

Researchers at University College London and the University of Roehampton examined how effectively people could understand human speech and cloned speech against a noisy backdrop.

Even the research team was surprised by the outcome. Rather than sounding poorer or less natural, the cloned voices repeatedly performed better.

How AI copies human voices

Voice clones differ from the older synthetic voices familiar through Siri, Alexa and sat-nav systems. Conventional artificial voices generally require many hours of recordings from a voice actor.

By comparison, voice-cloning technology can produce a close recreation of somebody's voice from just a few seconds of speech. That makes it considerably easier to deploy and far more scalable.

As a result, the pool of potential voices has grown dramatically. Instead of needing a professional speaker to spend hours in a recording studio, a very brief sample can be enough to copy almost anyone's voice.

This creates opportunities for a wide range of applications, including accessibility technology and entertainment. It also enables much more concerning uses, such as fraud and impersonation.

Testing voice clarity in noise

The study concentrated on a more limited issue: basic intelligibility. How easily can everyday listeners understand these cloned voices? The researchers had expected the clones to perform poorly, which would have seemed intuitive.

A machine-generated version of a voice, particularly one created from a brief sample, might be expected to sound less natural and therefore be more difficult to follow. Yet that was not the result.

“I thought initially that voice clones would be less intelligible because they were unfamiliar,” said study co-author Patti Adank, a professor of Speech Perception and Production at UCL.

“I found they were up to 20 percent more intelligible, which was quite shocking. A small part of our paper is talking about that experiment, and then a large part is me and my collaborator frantically trying to find out what it is that makes those voice clones more intelligible.”

Putting AI voices to the test

To investigate further, the team assessed human voices alongside their cloned equivalents in noisy listening environments.

An online experiment involving 80 participants compared 10 human voices with 10 clones at four different signal-to-noise levels.

The findings showed that cloned voices were easier to understand, with their intelligibility advantage reaching as much as 20 percent.

Crucially, the effect remained when the researchers altered both the listeners and the listening circumstances.

Cloned voices are easy to understand

Following the initial experiment, the researchers carried out the work again with older volunteers to establish whether hearing difficulties would change the pattern.

They also repeated the test with American listeners, as the original group had been British, before examining a filter designed to simulate cochlear implants.

Across every test, the cloned voices remained easier to understand. This consistency makes it difficult to dismiss the result as a chance finding.

Had cloned voices only outperformed human speech in one tightly defined setting, the result would have been easier to account for. Instead, the advantage persisted among different groups and in varied conditions.

That points to something systematic in the way these voices are produced.

Where AI voices could help

Initially, the finding may appear to be little more than an interesting technical observation. Yet it has significant implications, as synthetic speech is rapidly becoming a routine part of life.

If cloned voices remain clearer than human voices in noisy settings, they could be appealing for public announcements and assistive technology. They could also enhance navigation systems and communication tools developed for people with hearing difficulties.

However, this same benefit might cause cloned voices to seem more convincing or authoritative than they genuinely are.

A voice that carries through noise more successfully may also seem more polished, more controlled and, in certain situations, more trustworthy.

This presents another question – one the paper does not directly resolve, but leaves quietly in the background. If AI voices are easier to understand than people, how much of the listening world might move towards them?

Why AI voices sound better

The findings also reveal something intriguing about human speech. Natural voices contain countless forms of variation, shaped by pauses, breathing, stress, accents, roughness and thousands of subtle imperfections.

A voice clone could retain sufficient elements of the original speaker's identity while smoothing out some of the more untidy characteristics that make real speech difficult to catch in poor listening conditions.

This has not been demonstrated, but it is one possible explanation for why cloned voices could be outperforming human ones. That final point is an inference from the reported results, rather than something the study has conclusively proven.

The mystery of artificial speech

For now, the researchers do not have a definitive explanation for the effect. That lack of certainty is one reason the study is so compelling.

“I am now going to try to recreate [the effect] by studying how synthesizers work and how they use digital signal processing to generate those voices, just to get a better handle on this,” Adank said.

The unexpected issue, then, is no longer whether cloned voices can equal real voices. In this research, they appear to have surpassed them.

The central mystery is now the reason why. Understanding it could reveal not only more about artificial speech, but also about what makes any voice easy for the human ear to follow and trust in the first place.

Comments

No comments yet. Be the first to comment!

Leave a Comment