FUURAA AI Knowledge Library · Speech capability evaluation

How to evaluate AI speech recognition and spoken-language capability claims

Use six evidence gates to turn “supports multilingual speech” into an exact task, language and population, real audio chain, data provenance, reference transcript, scoring rules, complete failures, streaming latency, human correction, downstream outcome and dated boundary. Apply it to research and due diligence for transcription, meetings, captions, calls, spoken interfaces and voice agents.

Published27 August 2026Evidence statusMethod synthesis grounded in primary speech benchmarks and corpus researchScopeTranscription, captions, meetings, calls, spoken interfaces and voice-agent research

Hearing audio is not the same as understanding speech

A speech claim must keep audio, language, references, scoring, failures and the user task in one evidence chain.

A language count, clean-audio WER or polished demo does not explain dialects, code-switching, field noise, overlapping speech, streaming latency or correction burden. Even a correct transcript does not automatically establish that a downstream summary, dialogue or decision is correct.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, language, accent, dataset or leaderboard. It is not an accuracy guarantee, certification, procurement, audit, investment, legal or compliance opinion.

Six rejectable evidence gates

Each gate answers one decision question, requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the speech task, language and user outcome

Decision question
Is the claim about transcription, translation, language identification, diarisation, keyword spotting or a spoken dialogue outcome; for which language varieties, users and consequences?
Minimum evidence
Verbatim claim and date; system, model and API version; task and output schema; languages, varieties and code-switch policy; user population, downstream decision and acceptance thresholds.
Stop condition
Stop when a language count, demo transcript or generic label such as “human-level speech” replaces an exact task, population and quality floor.
02

Preserve audio acquisition, channel and preprocessing

Decision question
How was audio captured, encoded, segmented and transformed before recognition; was the run batch, streaming or interactive?
Minimum evidence
Original audio and hashes; microphone, distance, device and room; sample rate, codec and packet loss; voice-activity detection, denoising, resampling, channel mixing, segmentation and endpointing; client and network configuration.
Stop condition
Stop when clean close-mic clips or undisclosed enhancement are presented as evidence for field, telephone, meeting-room or real-time performance.
03

Define speaker, acoustic and data provenance distributions

Decision question
Which accents, dialects, ages, speaking styles, domains, noise, reverberation, overlap and vocabulary are represented; could evaluation audio or transcripts overlap training?
Minimum evidence
Sampling frame and subgroup counts; recording and annotation lineage; speaker and site separation; semantic and acoustic duplicate checks; domain, time and held-out splits; missing and excluded audio.
Stop condition
Stop when a curated corpus, selected speakers or a total-hour figure substitutes for representative deployment coverage.
04

Reproduce reference transcripts, normalisation and scoring

Decision question
Who produced the reference, how was ambiguity resolved, and how do tokenisation, casing, punctuation, numbers, segmentation, speakers and proper names enter the metric?
Minimum evidence
Annotation guide and agreement; adjudicated references; normalisation code and version; WER or CER alignment; punctuation, entity, timestamp and diarisation measures; matched scoring scripts and uncertainty.
Stop condition
Stop when WER or CER values produced by different references, tokenisation, normalisation or excluded segments are compared as if equivalent.
05

Count every failure, delay and correction burden

Decision question
Across all assigned audio, how often was output missing, late, truncated, wrongly attributed or unusable; what retries and human correction were required?
Minimum evidence
Complete denominator; empty, partial and timed-out outputs; dropped audio and language switches; hallucinated text and insertion bursts; real-time factor, endpoint and tail latency; retries, edit distance, correction time and task completion.
Stop condition
Stop when failed files, silence, endpoint errors, retries or human cleanup disappear from the accuracy, latency or cost claim.
06

Validate real-use transfer, downstream outcomes and expiry

Decision question
Does performance survive new speakers, devices, sites, live streaming, code-switching and operational load; does transcript quality support the actual user task?
Minimum evidence
Held-out site and device trials; streaming and concurrency tests; language-by-language and subgroup results; downstream task measures; monitoring, fallback and human review; change triggers and dated re-evaluation.
Stop condition
Stop when an offline corpus score is treated as proof of live captions, meetings, calls, voice agents, accessibility or high-consequence decisions.

Minimum speech-claim failure matrix

Check these conditions deliberately before transferring a controlled transcription score into a broad spoken-language claim.

  • 01
    Clean studio or close-mic speech is presented as evidence of field robustness.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

  • 02
    A single average hides weak languages, dialects, speakers or acoustic conditions.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

  • 03
    WER or CER is compared across different references, tokenisation or normalisation rules.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

  • 04
    Code-switching, proper names, numbers, jargon or spontaneous repairs are excluded.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

  • 05
    Voice-activity, endpoint, empty-output and failed-audio cases disappear from the denominator.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

  • 06
    Offline accuracy is presented as real-time streaming performance without latency and packet-loss tests.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

  • 07
    Synthetic noise substitutes for real sites, devices, overlapping speakers and recording chains.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

  • 08
    Transcript accuracy is treated as proof that the downstream dialogue, summary or decision is correct.

    Record the affected language, speaker, device, setting, metric and user task; preserve the narrowest conclusion that remains and specify whether to resample, re-annotate, rescore, include failures, retest or withdraw the claim.

Minimum speech-capability evaluation record

Let the next reviewer reconstruct the conclusion with the same audio, references, normalisation, scoring, failure denominator and operating boundary.

  1. 01Verbatim claim, decision, owner, publication date and evidence cut-off
  2. 02Exact system, model, API, decoder, vocabulary and version
  3. 03Speech task, languages, varieties, users, sites, consequences and quality floors
  4. 04Audio source, sampling, speakers, devices, channels, acoustic conditions and provenance
  5. 05Codec, sampling rate, VAD, denoising, segmentation, endpoint and streaming protocol
  6. 06Reference annotation, adjudication, normalisation, tokenisation and scoring version
  7. 07Complete outcomes, errors, exclusions, invalid runs and subgroup results
  8. 08Latency distribution, dropped audio, retries, concurrency and resource conditions
  9. 09Human corrections, correction time, downstream task outcome and uncertainty
  10. 10Transfer boundary, monitoring, fallback, change triggers and expiry

Common evidence states

Bind conclusions to the exact system, language, population, audio conditions, scoring, user task and date.

Supported

Evidence supports the exact system, speech task, language varieties, audio conditions, scoring protocol and dated operating boundary.

Conditional

Evidence supports a narrower language, speaker group, device, channel, noise condition, batch mode or downstream use.

Mixed

Performance varies materially across languages, varieties, speakers, acoustics, scoring dimensions, latency or failure types.

Insufficient

System identity, representative audio, reference quality, scoring rules, complete failures, latency, correction burden or transfer evidence is missing.

FUURAA analysisThe minimum decision unit for an AI speech recognition and spoken-language claim is exact system and version × speech task, language, variety, population and use decision × acoustic setting, device, channel and preprocessing × reference annotation, normalisation, segmentation and scoring × error distribution, failure denominator, latency and human correction × deployment boundary and cut-off date. AISHELL-1 and WenetSpeech connect Chinese Mandarin corpus development professionally to international research, but corpus scale, language count or one average error rate cannot replace direct validation for the target population and real acoustic chain.

Primary sources and non-transfer boundaries

These sources constrain low-resource languages, scaled weak supervision, multilingual coverage, acoustic robustness and Chinese Mandarin corpora; none independently proves general spoken-language capability.

Sources rechecked 27 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published 22 September 2022; updated 6 December 2022NIST — OpenASR21 Low-Resource Speech Recognition Challenge

Reports a constrained evaluation across fifteen low-resource languages, with case-sensitive scoring and error rates that expose the difficulty of proper names and scarce training data.

BoundaryA challenge result under prescribed data and scoring does not establish performance for every live channel, speaker population, vocabulary or downstream decision.

Open primary source ↗
First submitted 6 December 2022OpenAI — Whisper: Robust Speech Recognition via Large-Scale Weak Supervision

Studies multilingual and multitask speech recognition using 680,000 hours of weakly supervised audio and evaluates zero-shot transfer across diverse datasets.

BoundaryTraining scale and broad benchmark transfer do not independently prove the capability of a particular deployed version, streaming configuration, language variety or acoustic environment.

Open primary source ↗
First submitted 25 May 2022Google — FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Provides a parallel speech benchmark covering 102 languages for recognition, language identification, translation and retrieval, enabling language-by-language comparison.

BoundaryRead or prompted speech in a bounded corpus does not represent every dialect, code-switch, spontaneous conversation, device, noise condition or cultural setting.

Open primary source ↗
First submitted 8 March 2024Speech Robust Bench — Benchmarking Robustness of Speech Recognition

Evaluates speech recognition across 114 perturbations and examines both corruption robustness and subgroup disparities rather than relying on clean-audio averages.

BoundarySynthetic perturbations and selected subgroups cannot reproduce every real site, microphone chain, overlapping speaker, impairment or operational failure.

Open primary source ↗
First submitted 16 September 2017AISHELL Foundation — AISHELL-1 Mandarin Speech Corpus

Makes a Chinese-developed open Mandarin corpus, transcription protocol, lexicon and reproducible Kaldi recipe accessible to international speech research.

BoundaryA documented Mandarin corpus and recipe do not establish universal coverage of Chinese varieties, spontaneous speech, code-switching, field acoustics or current production systems.

Open primary source ↗
First submitted 7 October 2021WenetSpeech Collaboration — A 10,000+ Hours Multi-domain Mandarin Corpus

Brings a Chinese-led 22,400-hour multi-domain Mandarin collection into international evaluation, including matched web audio and mismatched meeting tests.

BoundaryCorpus scale and matched or mismatched test sets do not prove every dialect, speaker, device, conversational setting, label quality or deployment workflow.

Open primary source ↗

Continue checking

Move from speech capability into multilingual, multimodal, factuality and latency evaluation, or the full library.

Enter the AI Knowledge LibraryEvaluate multilingual capabilityEvaluate multimodal capabilityEvaluate factualityEvaluate latency and reliabilityEnter AI Evidence Atlas