Common evidence states
Bind conclusions to the exact direction, pipeline, source distribution, complete denominator, target workflow and date.
SupportedThe exact pipeline meets meaning, language, terminology, format, failure, cost and target-workflow thresholds on representative evidence with a current review date.
ConditionalSupport holds only for named directions, varieties, domains, source types, terminology controls, reviewer groups or operating conditions.
MixedMeaning, critical errors, fluency, terminology, document context, failures, post-editing, latency or cost vary materially across slices.
InsufficientDirection, pipeline, source provenance, reference quality, reproducible metrics, expert review, full denominator, target transfer or expiry is missing.
FUURAA analysisThe minimum decision unit for a machine translation capability claim is exact system and pipeline version × source language, target language, direction, variety, domain, audience and decision × source provenance, references, segmentation, context and format × prompts, tools, glossaries, translation memory and post-editing × meaning preservation, terminology, names, numbers, format, fluency and critical errors × all items, failures, corrections, latency and cost × target-workflow boundary and cut-off date. BLEU, SacreBLEU, COMET, MQM, FLORES-101 and CCEval illuminate reference overlap, comparable reporting, semantic metrics, expert errors, multilingual long-tail and Chinese-centric evaluation; none independently proves real translation capability.