
Voice has begun to receive more attention, especially when
artificial intelligence (AI) creates it. Tone and inflections in an ad or video are extremely important when trying to convey a feeling or emotion.
Emotion and tone predict listener
preference more strongly than the way a human voice sounds, and a robotic sounding voice remains the biggest complaint when listening to a voice generated by AI.
Askable Labs, an Askable
company that uses audio for AI-powered conversational data capture, built and ran a study in house that was published July 24, 2026. The first run of the recurring benchmark was published in English,
with Spanish, French and German versions to follow.
The prompts and human audio were sourced from a dataset of more than 40,000 Askable user interviews that lasted about 30 minutes each.
For this first version, the company focused on user interviews conducted in English, but plan to expand to more languages in subsequent versions.
advertisement
advertisement
Samples of at least 10 seconds in
length were curated, with a maximum of three minutes of contiguous speech. From these samples, the company filtered for clips that contained emotional responses, such as stories of frustration or
excitement. and then picked the top 100 of these options from unique participants. Analysts manually reviewed each prompt to ensure quality and accuracy of the transcription.
In the study,
titled VOICE-H, Askable found the robotic sound generated 590 negative mentions, while overly theatrical delivery from the voice also reduced preferences for a specific item.
The question
raters answered was "how well the emotion and tone match the context of the text." The focus was on fit rather than delivery mechanics, and found that flat and overdone delivery was perceived
negatively.
Designed to evaluate what people value in AI-generated speech, the findings compare the synthetic sound against the original human recording from the same conversation and capture
not only the voice people preferred, but why they preferred it.
Across nearly 4,500 comparisons spanning nine text-to-speech models -- including Google, OpenAI and ElevenLabs, with each
judged against the original human recording of the same speech -- participants consistently preferred voices whose emotional delivery matched the content being spoken.
Flat and monotone
delivery was the most common complaint, as well as voices that sounded exaggerated, theatrical or overly performative.
Listeners rewarded voices that struck the appropriate emotional balance
for each scenario.
The study suggests that humans prefer voices that deliver the correct emotion for the moment.
Naturalness alone did not determine preference, the study found.
Participants overlooked minor imperfections when a voice communicated emotion the way they thought it should sound, while even highly natural voices lost favor when their delivery felt mismatched to
the words.
Google's Gemini 3.1 Flash Voice ranked first overall with an Elo rating score of 1,101, ahead of the human recording on 1045 and Cartesia's Sonic 3.5 on 1,041.
Gemini also
led every benchmark in the study, with Gemini 3.1 Flash Voice scoring 1,101 vs. the human recording at 1,045.
For “naturalness,” a rating that suggests a natural voice, Gemini was
preferred -- coming in with a rating score of 1,094 vs. the human recording score of 1,062.
When it came to “accuracy,” Gemini rated 1,117 vs. the human recording at 947.
Analyzing motion and tone, Gemini came in with a rating of 1,146 vs. the human recording of 1,034.
These figures represent Elo rating, a standard chess-style rating system where the
synthetic score indicates that listeners of the AI voice prefer its emotional fit over the human recording's score, or vice versa. The score suggests how effectively the pitch, pace and emotional cues
resonated with the listener.
Askable used Elo ratings combined with an ordinal Bradley-Terry model to weigh the strength of each rater’s stated preference, with 95% bootstrap
confidence intervals.
One important point to mention is that rather than relying on scripted prompts, the benchmark draws on more than 40,000 real interviews conducted through Askable's
research platform.
Researchers selected 100 emotionally rich clips covering moments such as frustration, excitement and uncertainty before generating equivalent recordings from each AI
model.
Although not benchmarked in the study, Askable has worked with brands including HelloFresh, Sony, Visa, Toyota, Canon, Canva, British Airways, Pizza Hut, according to its website.