Voice AI

How a Voice AI Model Was Put Through 100 Hours of Human Testing
Software can check whether an AI voice said every word. It cannot reliably tell whether the voice sounded human.
This is a company developing an AI voice model, the kind that reads text aloud or speaks for a business on the phone. Automated tests are good at catching missing words. They are poor at hearing whether an apology sounded sorry, or whether a question rose at the end the way a person's would. That still takes a trained human ear.
TUSTRA carried out human evaluation of the model across more than 600 voice prompts and 100 hours of listening.
Two Ways of Testing
Rubric scoring
One clip is heard at a time.
Each quality is scored separately against fixed criteria.
It tells the developer why the voice falls short and exactly where.
A/B preference testing
Two versions of the same line are played in random order and left unlabelled.
The listener judges which version sounds better overall, or on one specific quality.
It tells the developer which version of the model listeners prefer.
The A/B tests were blind. Hiding which version is which stops a listener favouring the one they expect to be better.
What Every Prompt Was Judged On
Emotion
Does the feeling match the words?
A cancelled-flight apology should not sound cheerful.
Tone
Is it right for the setting?
Calm for a bank alert, warmer for a welcome message.
Pitch
Does the voice sit at a natural level without drifting or jumping?
Intonation
Does it rise and fall where a person's would?
This includes rising naturally at the end of a question.
Pronunciation
Are names, numbers and harder words said correctly and clearly?
Rhythm
Are the pauses and emphasis where a real speaker would put them?
Why It Takes 100 Hours
600 prompts in 100 hours is around ten minutes a prompt.
Judging six qualities separately, and comparing two versions of the same line, is slow by design. A quick overall impression misses exactly the faults a developer needs to find, such as one word stressed in the wrong place.
What It Does Not Do
Human evaluation shows where a voice falls short. It does not fix it.
The fixes happen in the developer's own training, and so do the results. It is also slower and more expensive than automated scoring, so it belongs on the qualities software cannot judge, not on everything.
Where This Applies
Any business putting an AI voice in front of customers faces the same question on a smaller scale.
A voice agent that answers with the wrong tone, or mispronounces the company's own name, loses trust on the first call. The same structured, blind testing can be run on any voice before it goes live.
TUSTRA runs human evaluation of AI voices, using the rubric and blind A/B methods that AI developers use to test their own models.