Common evidence states
Bind conclusions to the exact pipeline, text and voice population, complete denominator, target-listening workflow and date.
SupportedThe exact pipeline meets declared intelligibility, pronunciation, prosody, naturalness, similarity, complete-denominator, latency, cost and target-listener thresholds, with a current review date.
ConditionalSupport holds only for named languages, voices, text types, styles, lengths, devices, codecs or operating controls.
MixedIntelligibility, pronunciation, prosody, naturalness, similarity, long-form stability, latency or cost differ materially across slices.
InsufficientText or audio provenance, pipeline identity, pronunciation references, listening design, complete denominator, target validation or expiry is missing.
FUURAA analysisThe minimum decision unit for an AI speech-synthesis capability claim is exact model, checkpoint, text-front-end, acoustic-model, vocoder, codec and serving version × synthesis task, text population, language, voice, style, listener, device and use × text and audio provenance, normalization, pronunciation references, speaker and listener distribution × token or phoneme path, duration, pitch, energy, sampling, waveform, post-processing and delivery × intelligibility, pronunciation, critical tokens, prosody, naturalness, similarity and long-form continuity × all utterances, failures, repeats, selection, human correction, latency and cost × target-listening workflow boundary and cut-off date. One natural sample is an observation, not transferable proof of speech-synthesis capability.