FUURAA AI Knowledge Library · speech-synthesis evaluation

How to evaluate AI speech-synthesis capability claims

Use six evidence gates to turn “human-like speech,” “natural narration” or “consistent multilingual voices” into exact text, language and voice conditions, a reconstructable generation pipeline, intelligibility, pronunciation, prosody, naturalness, speaker similarity, a complete failure denominator, delivery latency, cost and a dated target-listening boundary.

Published31 August 2026Evidence statusMethod synthesis grounded in primary speech-generation, corpus and listening-test researchScopeText-to-speech, multi-speaker, expressive-control and streaming speech synthesis

Naturalness is not correct reading; speaker similarity is not complete speech capability

Speech-synthesis capability must keep text, pronunciation, the generation pipeline, listening design, every output and real delivery in one evidence chain.

Pleasant audio can still omit or repeat words and misread numbers or names; a high MOS may come from samples, listeners and devices unlike the target use. Evaluation must separate intelligibility, pronunciation, prosody, naturalness, speaker similarity, long-form continuity, selection rate and delivery cost.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, voice, language, corpus or leaderboard. It does not guarantee speech quality, pronunciation correctness, identity consistency, accessibility fitness or business outcomes.

Six rejectable evidence gates

Each gate requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the speech task, text, voice and decision

Decision question
What exact text-to-speech, multi-speaker, style-control or streaming task, language, voice condition, listener and acceptance decision does the claim cover?
Minimum evidence
Verbatim claim and date; exact model, checkpoint, text front end, acoustic model, vocoder and serving versions; input text, language and variety; voice and style conditions; listener, channel, use, error costs, thresholds and allowed fallback.
Stop condition
Stop when “human-like speech” or one polished sample replaces a named task, exact pipeline, eligible text population and decision threshold.
02

Version text, pronunciation, audio and listener evidence

Decision question
Which exact texts, pronunciations, speakers, recordings and listening populations were eligible, and what counted as the reference?
Minimum evidence
Text and audio provenance, collection date, permitted use, recording chain and sample rate; normalization and pronunciation references; speaker, language, accent, style, length and domain distribution; listener language proficiency, audio equipment, instructions, randomization, reference anchors, exclusions and disagreement.
Stop condition
Stop when test sentences, reference pronunciations, recording conditions, listener population or selected outputs are unknown.
03

Rebuild text processing, acoustic generation and delivery

Decision question
Can every waveform be traced from raw text through normalization, pronunciation, timing, acoustic representation, vocoder, post-processing and playback?
Minimum evidence
Unicode, numbers, dates, currencies, abbreviations and pronunciation lexicon; grapheme-to-phoneme or token path; language and speaker conditioning; duration, pitch, energy and style controls; sampling, seeds, vocoder, loudness, codec, streaming chunks, retries, caching and serving logs.
Stop condition
Stop when only the final audio remains and the normalized text, pronunciation path, generation settings, selection and post-processing cannot be reconstructed.
04

Separate intelligibility, pronunciation, prosody, naturalness and voice similarity

Decision question
Is the output understandable and correctly pronounced, with appropriate timing and expression, rather than merely sounding pleasant or similar?
Minimum evidence
Human transcription or verified ASR intelligibility; word, phoneme and critical-token errors; numbers, names, abbreviations and code-switching; duration, pauses, stress, pitch and rhythm; naturalness MOS with confidence intervals; speaker-similarity listening and embeddings; paired tests, rater agreement and diagnostic error labels.
Stop condition
Stop when one MOS, an automatic quality predictor or speaker embedding substitutes for pronunciation, intelligibility, prosody, human listening design and critical-error review.
05

Count every utterance, failure, variance and cost

Decision question
What happened across every eligible sentence, speaker, language slice, repeated generation and delivery attempt?
Minimum evidence
Complete text and audio denominator; empty, truncated, repeated, skipped, mispronounced, unstable, noisy and unintelligible outputs; long-form discontinuity; repeated-run variation and selection rate; language, accent, speaker, style, length and text-type slices; synthesis speed, time to first audio, tail latency, memory, compute, human correction and total cost.
Stop condition
Stop when selected samples or average MOS hide failed sentences, pronunciation-critical errors, long-form drift, regeneration, manual editing, slow tails or cost.
06

Transfer to the target listening workflow and expire the claim

Decision question
Does controlled quality remain useful with target texts, listeners, devices, networks, latency limits and changing voices or languages?
Minimum evidence
Shadow or staged tests on target text distributions, languages, devices, codecs, noise and accessibility needs; listener comprehension and task outcomes; fallback and human correction; pronunciation, language, model, voice, codec and serving change triggers; monitoring, owner, rollback and expiry date.
Stop condition
Stop when a studio sample or static benchmark transfers directly to live narration, accessibility, assistants or high-volume delivery without target validation and rollback.

Minimum speech-synthesis claim failure matrix

Check these conditions deliberately before transferring selected audio, one MOS or a similarity score into real speech-synthesis capability.

  • 01
    Numbers, dates, currencies, units, abbreviations or names are spoken incorrectly

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

  • 02
    The audio sounds natural while omitting, repeating or changing words

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

  • 03
    A speaker-similarity score is high while pronunciation or prosody is wrong

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

  • 04
    Automatic MOS rewards audio that target listeners rate differently

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

  • 05
    Code-switching, rare words, punctuation or an unfamiliar script breaks the text front end

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

  • 06
    Long-form audio drifts in voice, pace, loudness, style or continuity

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

  • 07
    Only the best regeneration is presented while failed attempts are omitted

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

  • 08
    A model, lexicon, voice, codec or serving change invalidates prior evidence

    Record the affected text, language, voice, pipeline, listener, device and use; preserve the narrowest conclusion that remains and specify whether to add data, revise the lexicon, regenerate, rerun listening tests, hand off to a human, restrict use or withdraw the claim.

Minimum speech-synthesis capability evaluation record

Let the next reviewer reconstruct the conclusion with the same text, pronunciation references, pipeline, listening design and complete outputs.

  1. 01Claim, date, owner, speech task, text population, language, voice, listener, use and decision
  2. 02Model, checkpoint, text front end, acoustic model, vocoder, codec and serving versions
  3. 03Text and audio provenance, permitted use, collection, recording chain, sample rate and exclusions
  4. 04Normalization, lexicon, pronunciation references, token or phoneme path and language handling
  5. 05Speaker, style, duration, pitch, energy, seed, sampling and post-processing settings
  6. 06Listening design, raters, devices, anchors, randomization, agreement and confidence intervals
  7. 07Intelligibility, pronunciation, critical tokens, prosody, naturalness and similarity results
  8. 08All utterances, failures, repeats, selections, long-form continuity and subgroup results
  9. 09Human correction, synthesis speed, first-audio and tail latency, compute and total cost
  10. 10Target evidence, monitoring, change triggers, fallback, rollback, owner and expiry

Common evidence states

Bind conclusions to the exact pipeline, text and voice population, complete denominator, target-listening workflow and date.

Supported

The exact pipeline meets declared intelligibility, pronunciation, prosody, naturalness, similarity, complete-denominator, latency, cost and target-listener thresholds, with a current review date.

Conditional

Support holds only for named languages, voices, text types, styles, lengths, devices, codecs or operating controls.

Mixed

Intelligibility, pronunciation, prosody, naturalness, similarity, long-form stability, latency or cost differ materially across slices.

Insufficient

Text or audio provenance, pipeline identity, pronunciation references, listening design, complete denominator, target validation or expiry is missing.

FUURAA analysisThe minimum decision unit for an AI speech-synthesis capability claim is exact model, checkpoint, text-front-end, acoustic-model, vocoder, codec and serving version × synthesis task, text population, language, voice, style, listener, device and use × text and audio provenance, normalization, pronunciation references, speaker and listener distribution × token or phoneme path, duration, pitch, energy, sampling, waveform, post-processing and delivery × intelligibility, pronunciation, critical tokens, prosody, naturalness, similarity and long-form continuity × all utterances, failures, repeats, selection, human correction, latency and cost × target-listening workflow boundary and cut-off date. One natural sample is an observation, not transferable proof of speech-synthesis capability.

Primary sources and non-transfer boundaries

These sources constrain waveform generation, sequence-to-sequence TTS, English and Mandarin corpora, end-to-end generation and large-scale in-context speech synthesis.

Sources rechecked 31 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published 12 September 2016WaveNet — A Generative Model for Raw Audio

Introduced an autoregressive model for raw waveforms and reported listener-rated text-to-speech naturalness for English and Mandarin, including multi-speaker conditioning.

BoundaryResults for selected corpora, speakers and listening tests do not establish real-time delivery, long-form stability, every pronunciation, language variety or target acoustic environment.

Open primary source ↗
Published 16 December 2017Tacotron 2 — Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Connects a sequence-to-sequence mel-spectrogram predictor to a WaveNet vocoder and reports human mean-opinion scores close to professional recordings on its evaluation setting.

BoundaryA high MOS on one English, professionally recorded voice does not establish pronunciation coverage, speaker transfer, expressive control, multilingual quality or complete utterance reliability.

Open primary source ↗
Published 5 April 2019LibriTTS — A Corpus Derived from LibriSpeech for Text-to-Speech

Provides about 585 hours of 24 kHz English read speech from 2,456 speakers with corresponding text, designed for multi-speaker TTS research.

BoundaryAudiobook-derived English read speech does not represent spontaneous conversation, every accent, specialist term, noisy input, emotional style or live product workload.

Open primary source ↗
Published 22 October 2020AISHELL-3 — A Multi-Speaker Mandarin TTS Corpus and the Baselines

A Chinese-led project releases about 85 hours of high-fidelity, emotion-neutral Mandarin speech from 218 native speakers with character and pinyin transcripts and multi-speaker baselines.

BoundaryEmotion-neutral Mandarin recordings and embedding-based similarity do not establish every Chinese variety, code-switching, expressive speech, pronunciation correctness or target-listener preference.

Open primary source ↗
Published 11 June 2021VITS — Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Unifies acoustic modelling and waveform generation through variational inference, normalising flows, adversarial learning and a stochastic duration predictor.

BoundaryNaturalness on a single-speaker benchmark does not prove text-normalisation coverage, multilingual transfer, streaming latency, long-form continuity or robustness to unfamiliar inputs.

Open primary source ↗
Published 4 June 2024Seed-TTS — A Family of High-Quality Versatile Speech Generation Models

A Chinese-led technical report evaluates large-scale autoregressive speech generation across speaker similarity, naturalness, in-context synthesis, speaker adaptation and emotion control.

BoundarySelected objective and subjective tests in a technical report do not establish every voice, language, long-form session, pronunciation edge case, operating cost or deployment condition.

Open primary source ↗

Continue checking

Move from speech synthesis into speech recognition, machine translation, video generation or the full library.

Enter the AI Knowledge LibraryEvaluate speech recognitionEvaluate machine translationEvaluate video generationEnter AI Evidence Atlas