FUURAA AI Knowledge Library · machine translation evaluation

How to evaluate AI machine translation capability claims

Use six evidence gates to turn “supports many languages, near-human and ready to use” into exact language directions, versioned sources and references, a complete pipeline, meaning and critical errors, all failures, post-editing, cost and a dated target-workflow boundary.

Published29 August 2026Evidence statusMethod synthesis grounded in primary MT evaluation researchScopeText and document machine translation, quality estimation and post-editing research

Language count is not capability; fluency is not fidelity

Machine translation capability must keep direction, sources, references, pipeline, meaning, critical errors, failures and human correction in one evidence chain.

The same system can perform well in one direction and domain while failing in reverse translation, dialects, terminology-heavy text or long documents. Automatic scores locate signals; they do not replace reproducible configurations, source-grounded expert error review and target-workflow validation.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, language, translator, dataset or leaderboard. It is not a guarantee of accuracy, completeness, fitness or decision quality.

Six rejectable evidence gates

Each gate requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the translation task, direction and decision

Decision question
What does “good translation” mean for this source language, target language, variety, domain, audience and downstream use?
Minimum evidence
Verbatim claim and date; exact system and pipeline versions; source-to-target direction; content type, domain, register, audience and consequence; acceptance thresholds for meaning, terminology, format, latency and cost.
Stop condition
Stop when language count, a selected demo or one average score substitutes for a named direction, task and decision boundary.
02

Version sources, references and evaluation slices

Decision question
Where did eligible source material and references come from, who produced them, and which linguistic and operational variation do they cover?
Minimum evidence
Source manifest and hashes; ownership, licence, dates and collection process; native-writing versus translated origin; reference translators and review; segmentation, context, scripts, dialects, code-switching, length, domains and difficulty slices.
Stop condition
Stop when source provenance, reference quality, contamination, missing documents or material language varieties are unknown.
03

Reproduce the complete translation pipeline

Decision question
What processing, instructions, tools and human work occur between source input and accepted target text?
Minimum evidence
Model, API and prompt versions; OCR, normalisation, segmentation and document context; glossary, translation memory, retrieval and tools; sampling and decoding; safety filters, retries, post-processing, quality estimation and human post-editing.
Stop condition
Stop when a model name hides preprocessing, terminology aids, retries, filtering, selection or human correction.
04

Separate meaning, language, terminology and format

Decision question
Does the target preserve all material meaning while remaining correct, appropriate and usable for the intended audience?
Minimum evidence
Reproducible BLEU or neural-metric signatures; expert MQM-style errors and severity; omissions, additions and mistranslations; names, numbers, negation, terminology and locale formats; fluency, register, discourse, document context and task-specific critical errors.
Stop condition
Stop when fluent wording, round-trip translation, one reference-overlap score or an unvalidated LLM judge stands in for source-grounded expert review.
05

Count every item, failure and correction

Decision question
What happened across all eligible source units, including omissions, refusals, timeouts, malformed output, retries and post-editing?
Minimum evidence
Full denominator and run IDs; empty, truncated and untranslated output; errors by direction, language variety, domain and severity; retries and fallbacks; reviewer agreement; edit distance and minutes; latency, compute, API and total human cost distributions.
Stop condition
Stop when failed items disappear, selected translations replace the full denominator, or human post-editing is treated as free.
06

Validate the target workflow and expire the conclusion

Decision question
Does the controlled pipeline help intended users on their documents, terminology and consequences after sources, prompts, models or workflows change?
Minimum evidence
Shadow or staged use; representative native-speaker and domain-expert review; task outcome and harm checks; monitoring by language slice and critical error; source, terminology, prompt, model, tool and post-editor change triggers; owner and expiry date.
Stop condition
Stop when benchmark evidence transfers to live translation without direct user, document, terminology, monitoring and dated recheck evidence.

Minimum machine-translation claim failure matrix

Check these conditions deliberately before transferring language count, averages or selected translations into real capability.

  • 01
    A high-resource average hides collapse in a low-resource direction, dialect, script, register or code-switched input.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

  • 02
    Tokenisation, casing, punctuation, normalisation or reference choices make two reported BLEU scores incomparable.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

  • 03
    The translation is fluent but omits negation, softens uncertainty, changes agency or invents supporting detail.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

  • 04
    Names, numbers, units, dates, currencies, links, placeholders or document structure are altered or dropped.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

  • 05
    Sentence-level quality looks strong while pronouns, terminology, tone and relationships drift across a document or conversation.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

  • 06
    Round-trip translation returns familiar wording even though the first translation lost source meaning.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

  • 07
    An automatic or LLM judge rewards fluency, misses a critical semantic error or behaves differently by language.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

  • 08
    Selected examples exclude refusals, timeouts, safety blocks, retries, terminology repair and human post-editing effort.

    Record the affected direction, source type, error severity, user and decision; preserve the narrowest conclusion that remains and specify whether to add samples, rerun, review expertly, translate manually, restrict use or withdraw the claim.

Minimum machine-translation capability evaluation record

Let the next reviewer reconstruct the conclusion with the same direction, sources, references, pipeline, metrics, human review and complete runs.

  1. 01Claim, date, owner, source and target direction, task, audience, consequence and thresholds
  2. 02System, model, API, prompt, decoding, tool, filter and deployment versions
  3. 03Source manifest, provenance, licences, dates, native origin, contamination and exclusions
  4. 04Reference production, translator qualifications, review, segmentation and document context
  5. 05Language varieties, scripts, domains, registers, lengths, difficulty and critical-risk slices
  6. 06Metric names, versions, signatures, references, aggregation and uncertainty
  7. 07Expert error taxonomy, severity, reviewer agreement and critical-error results
  8. 08All eligible items, omissions, refusals, failures, retries, fallbacks and corrected outputs
  9. 09Terminology, translation memory, retrieval, post-editing, latency and total cost
  10. 10Target-workflow validation, monitoring, change triggers, residual limits, owner and expiry

Common evidence states

Bind conclusions to the exact direction, pipeline, source distribution, complete denominator, target workflow and date.

Supported

The exact pipeline meets meaning, language, terminology, format, failure, cost and target-workflow thresholds on representative evidence with a current review date.

Conditional

Support holds only for named directions, varieties, domains, source types, terminology controls, reviewer groups or operating conditions.

Mixed

Meaning, critical errors, fluency, terminology, document context, failures, post-editing, latency or cost vary materially across slices.

Insufficient

Direction, pipeline, source provenance, reference quality, reproducible metrics, expert review, full denominator, target transfer or expiry is missing.

FUURAA analysisThe minimum decision unit for a machine translation capability claim is exact system and pipeline version × source language, target language, direction, variety, domain, audience and decision × source provenance, references, segmentation, context and format × prompts, tools, glossaries, translation memory and post-editing × meaning preservation, terminology, names, numbers, format, fluency and critical errors × all items, failures, corrections, latency and cost × target-workflow boundary and cut-off date. BLEU, SacreBLEU, COMET, MQM, FLORES-101 and CCEval illuminate reference overlap, comparable reporting, semantic metrics, expert errors, multilingual long-tail and Chinese-centric evaluation; none independently proves real translation capability.

Primary sources and non-transfer boundaries

These sources constrain automatic scores, configuration reproducibility, semantic metrics, expert errors, language coverage and Chinese-centric evaluation.

Sources rechecked 29 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published July 2002BLEU — A Method for Automatic Evaluation of Machine Translation

Introduces corpus-level modified n-gram precision and a brevity penalty for reproducible comparison with human references.

BoundaryReference overlap does not independently establish meaning preservation, terminology, names, numbers, discourse, user usefulness or safety.

Open primary source ↗
Published October 2018SacreBLEU — A Call for Clarity in Reporting BLEU Scores

Shows that tokenisation and normalisation choices materially change BLEU, and proposes a standard signature for comparable reporting.

BoundaryA reproducible metric configuration improves comparability; it does not make the metric complete or validate the source, references or target workflow.

Open primary source ↗
Published November 2020COMET — A Neural Framework for MT Evaluation

Learns multilingual quality estimates from the source, candidate and reference using several forms of human judgement.

BoundaryCorrelation on WMT data does not guarantee calibration for another language direction, domain, error severity, model version or high-impact decision.

Open primary source ↗
Published 2021MQM — Experts, Errors, and Context

Uses professional translators, explicit error types and severity, and full document context to study human evaluation at scale.

BoundaryTwo WMT language pairs and a particular annotation design do not supply a universal taxonomy, reviewer agreement or acceptance threshold for every use case.

Open primary source ↗
Published 2022FLORES-101 — Low-Resource and Multilingual MT Evaluation

Provides 3,001 professionally translated, fully aligned sentences across 101 languages for controlled many-to-many evaluation.

BoundaryWikipedia-derived sentences do not cover every dialect, register, document, dialogue, terminology system, locale format or production consequence.

Open primary source ↗
Published December 2023CCEval — Chinese-Centric Multilingual Machine Translation

A Chinese-led benchmark traces diverse Chinese source sentences across six domains and eleven target languages, including low-resource directions.

BoundaryA Chinese-centric sentence benchmark does not prove document-level coherence, every Chinese variety, specialist terminology, live input quality or operational post-editing effort.

Open primary source ↗

Continue checking

Move from machine translation into multilingual capability, text summarisation, speech recognition or the full library.

Enter the AI Knowledge LibraryEvaluate multilingual capabilityEvaluate text summarisationEvaluate speech recognitionEnter AI Evidence Atlas