Skip to content
tarıtas
Build decisions POST-28 5 min read

How to test voice naturalness with AI judges: a repeatable blind A/B method

Compare two clips that differ in one thing, play each pair in both orders, and count a judge's vote only when both orders agree. Taritas runs this with a screened panel of three audio model judges and a same-clip control. One run of 176 comparisons cost about $5 and gave a clear, reliable decision on three voice features in two languages. Before trusting a judge, screen it on pairs where you already know the answer and drop any model with order bias: some picked the first clip 75 to 94% of the time. Match loudness, trim silence, and write the pass rule before the run, with inconclusive meaning do not ship. Decide on order-consistent wins and losses, not on 1-to-5 naturalness scores. In our run the judges picked a winner in only 2 of 30 control pairs, and the plain voice sounded as good as or better than breaths, fillers and longer pauses.

Published · Updated · Supreet Tare

All names, numbers, and identifiers in this post are anonymized. The patterns are real.

A repeatable blind listening test. Each pair differs in one thing and is loudness matched. Every judge hears it in both orders, and a vote counts only if both orders agree. Three screened audio model judges vote, and judges with order bias are screened out first. A control pair of the same sentence voiced twice checks the judges. Results from 176 comparisons for about $5 gave a clear decision on breath, filler and pause features.

Voice quality is hard to measure, and it is harder still in a language your team does not speak. We needed to decide whether three small touches would make a bilingual voice agent sound more human: a soft “Hmm…” before some replies, slightly longer pauses between sentences, and a quiet breath before speaking.

So we built a blind listening test with audio-understanding models as judges. It cost about $5, took one run, and gave a clear answer. This post is the method, step by step, so you can reuse it. (The same agent handles its decisions and turn detection with Jev, covered in our guide to using Jev in a voice agent.)

Step 1: map what your TTS model can do

Start by measuring your text-to-speech (TTS) model’s real capabilities. TTS models accept tags in the text for effects such as a pause or a breath, and support varies by model.

What we found on ours:

What we testedResult
Pause tagsWorked
Written fillers (“Hmm…”)Pronounced as written
Breath tags like [inhales] on the fast modelRead aloud as the word “inhales”
Expressive model that supports breath tagsAbout 0.8 s to first audio, too slow for conversation

That shaped the design: pause tags, written fillers, and breath clips recorded in each voice and inserted before the reply, all on the fast model.

Step 2: screen your judges

Before trusting a judge, test it on pairs where you already know the answer. We screened eight audio models this way.

  • Order bias ruled some out. A few models picked whichever clip played first 75 to 94% of the time.
  • The best panel had three judges: two Gemini models plus MiMo from a second provider. MiMo was the most disciplined judge and never reported a preference that was not there.
  • The panel was strong on clear faults: a spelled-out “H M M”, a doubled word, rushed speech.
  • Subtle differences need the full method in the next step, because that is where a single judge’s vote is least reliable.

Step 3: design each comparison

Play both orders. Each judge hears A then B, and B then A. The vote counts only if both orders pick the same clip.

def order_consistent_vote(judge, clip_a, clip_b):
    first = judge.compare(clip_a, clip_b)   # "first" or "second"
    second = judge.compare(clip_b, clip_a)
    if first == "first" and second == "second":
        return "A"
    if first == "second" and second == "first":
        return "B"
    return None  # inconsistent, so it does not count

Change one thing per pair. For pauses, the sentence audio is identical and only the gaps differ.

Match loudness and trim silence. Judges notice loudness quickly, and a louder clip once got tagged “sounds excited”. Normalize every clip first. For example, with ffmpeg:

ffmpeg -i clip.wav -af "silenceremove=start_periods=1:start_threshold=-50dB,loudnorm" clip_norm.wav

Include a control. Add pairs where both clips are the same sentence voiced twice. This tells you how often the judges hear a difference that is not there.

Write the pass rule before you run. Decide what counts as a win, and treat “inconclusive” as “do not ship”. Writing the rule first, called pre-registration, keeps the decision honest once the numbers arrive. A simple template:

feature: breath_before_reply
compare: with_breath vs plain
votes: order-consistent only
ship_if: feature wins a clear majority of consistent votes
inconclusive: do not ship
control_limit: judges pick a winner in under 10% of control pairs

Step 4: read the right numbers

Use order-consistent wins and losses. Skip the 1-to-5 naturalness scores: they ranged from 2 to 5 on the same clip, and written reasons sometimes changed between the two orders.

What the method told us

One run, 176 comparisons, about $5:

FeatureResultDecision
Breath before speaking0 wins, 7 losses in English. 38% of pairs flagged “robotic breath”.Keep the plain voice
”Hmm” filler, second languageFlagged unnatural in 42% of pairs (“sounds spelled out”).Keep the plain voice
Longer pausesNo audible difference in 56 of 57 pairs.Keep the default pacing
Control, same clip twiceJudges picked a winner in only 2 of 30 pairs.Results confirmed

The control row is what makes the rest trustworthy: the judges rarely invented a preference, so the other results reflect real differences. The decision was clear in one run: the plain voice sounded as good or better. That kept the pipeline simpler and saved the weeks we would otherwise have spent tuning breaths and fillers.

Listening test checklist

  • Measure what your TTS model supports before designing features.
  • Screen judges on known pairs, and drop any with order bias.
  • Use a panel of at least two providers.
  • One difference per pair.
  • Loudness matched and silence trimmed.
  • Every pair played in both orders.
  • Control pairs included in every run.
  • Pass rule written before the run, with “inconclusive” meaning “do not ship”.
  • Decide on order-consistent votes, not on 1-to-5 scores.
  • Remember that a transcript checks the words, not how they sound: speech-to-text heard every filler correctly.

What this means if you are an IT services firm

Clients often ask for a voice agent that “sounds more human”, sometimes in languages your team does not speak. This method turns that request into a measurable decision, in any language the judges support, for a few dollars per run.

Taritas uses this evaluation when building voice agents with engineering teams who are new to conversational AI. If you want that rigor on your next client project, see how we work with partners.

Related questions
Can AI models judge whether synthetic speech sounds natural?
Yes, for clear differences. A screened panel of audio models reliably caught spelled-out fillers, doubled words and rushed speech. For very subtle differences, use the full method: both play orders, a control pair and a fixed pass rule.
What is order bias in audio model judges, and how do you remove it?
Order bias is a judge preferring whichever clip it hears first. Some models picked the first clip 75 to 94% of the time. Play every pair in both orders and count a vote only when both orders pick the same clip.
Why include a control pair in a listening test?
A control pair plays the same sentence voiced twice, so any preference is noise. It measures how often judges hear a difference that is not there. In our run they picked a winner in only 2 of 30 control pairs, which confirmed the other results.
How much does an AI-judged listening test cost?
Our run of 176 comparisons, with three judges and both play orders, cost about $5. That makes it cheap enough to run before every voice change.
Can you test TTS quality in a language your team does not speak?
Yes. The judges listen to the audio directly, so the method works in any language they support. It is how we evaluated a second language that nobody on our team speaks.

Reading this because a client asked for voice AI? That is the conversation we are built for. What taritas does for partners.

More from Build decisions
PROJECT taritas.com/blog
DWG POST-28
REV 1.0
DATE 2026-09-28