FUURAA AI Knowledge Library · classification and intent-routing evaluation

How to evaluate AI text-classification and intent-routing capability claims

Use six evidence gates to turn “accurately identifies intent,” “routes every request automatically” or “leads multilingual classification” into a versioned label ontology, representative inputs and annotation, reconstructable scoring and routing, classwise performance, calibration and out-of-scope rejection, a complete failure denominator, real consequences and a dated applicability boundary.

Published31 August 2026Evidence statusMethod synthesis grounded in primary English, Chinese and multilingual NLU and intent researchScopeText classification, intent routing, out-of-scope detection and multilingual NLU

Overall accuracy is not reliable routing; known-label scores are not unknown-input handling

Classification capability must keep label meaning, input population, annotation, the complete pipeline, classwise errors, rejection and real consequences in one evidence chain.

Class prevalence can make overall accuracy look strong while a consequential minority class still fails; the nearest known label may not be a valid answer. Evaluation must test confusion, calibration, low-confidence abstention, out-of-scope and new intents, together with the human and operational cost of misrouting.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, dataset or leaderboard. It does not guarantee routing accuracy, policy compliance, fairness, safety, customer experience or business outcomes.

Six rejectable evidence gates

Each gate requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the task, label ontology and decision

Decision question
What exact single-label, multi-label or hierarchical task, supported labels, exclusions, users, channels and downstream actions does the claim cover?
Minimum evidence
Verbatim claim and date; exact system, model, prompt, feature, label-description, threshold and routing versions; label names, definitions, examples, overlaps and exclusions; users, languages, channels, actions, error costs, abstention and acceptance decision.
Stop condition
Stop when “understands every request” or “high intent accuracy” replaces a versioned ontology, action, threshold and consequence.
02

Version the source population, annotation and label change

Decision question
Who produced each input, when, under which sampling and annotation rules, and how were ambiguous or changing labels resolved?
Minimum evidence
Input provenance, collection window, inclusion and exclusion, prevalence and sampling; annotation guide, annotator qualifications, disagreements and adjudication; duplicates, synthetic or translated data; train-validation-test independence; label additions, merges, splits and retirements.
Stop condition
Stop when duplicates, future records, annotator cues, label names or non-representative balancing leak the answer or distort prevalence.
03

Rebuild input processing, scoring and routing

Decision question
Can every assigned label or route be traced through preprocessing, context, model scores, thresholds, rules, fallback and human handling?
Minimum evidence
Normalization, tokenization, transcription, OCR or translation; context, truncation and retrieval; model, prompt, embeddings, rules, label descriptions and examples; raw scores or probabilities, top-k, thresholds, tie-breaking, abstention, fallback, human route and logs.
Stop condition
Stop when only the final label is retained and the transformed input, score, threshold, rule and pipeline version cannot be reconstructed.
04

Test every class, calibration and unsupported input

Decision question
Does the system beat appropriate baselines across every consequential class and reject ambiguous, out-of-scope, out-of-distribution and new-intent inputs?
Minimum evidence
Stratified and chronological splits; majority, rules and simple model baselines; confusion matrix; per-class precision, recall and F1; macro, micro and weighted aggregates; AUROC or AUPRC where appropriate; Brier score or expected calibration error; selective risk, OOS and OOD tests; rare, ambiguous, adversarial, multilingual and code-switched slices.
Stop condition
Stop when overall accuracy hides imbalance, critical false positives, weak rare-class recall, poor calibration or forced classification of unsupported input.
05

Count every input, failure, correction and cost

Decision question
What happened across the complete target denominator, including empty, uncertain, conflicting, repeated, failed and manually corrected cases?
Minimum evidence
All eligible inputs and labels; empty, malformed, duplicate, low-confidence, multi-label, conflicting and unsupported cases; retries and repeated-run variance; wrong-route outcomes, human overrides, escalation and correction time; latency distribution, compute and total operating cost.
Stop condition
Stop when selected examples or a successful subset hide abstentions, dropped inputs, critical misroutes, manual rescue, tail latency or cost.
06

Transfer through drift and expire the claim

Decision question
Does the evidence survive the target channel, current language mix, new labels, policy change and changing input distribution?
Minimum evidence
Shadow or staged target tests; input, label and prevalence drift; confusion and calibration monitoring; new-intent and OOS rates; threshold recalibration; change triggers, owner, fallback, rollback and expiry date.
Stop condition
Stop when a static benchmark transfers directly to live routing without target validation, drift monitoring, safe fallback and rollback.

Minimum classification and routing claim failure matrix

Check these conditions deliberately before transferring an aggregate score or selected examples into real classification and intent-routing capability.

  • 01
    Labels overlap, contradict one another or omit a valid destination

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

  • 02
    Class imbalance makes high overall accuracy hide rare or critical failures

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

  • 03
    Label names, examples or duplicate records leak the answer

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

  • 04
    Train and test contain near-duplicates, future data or the same user threads

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

  • 05
    Out-of-scope input is forced into the closest supported class

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

  • 06
    Probabilities are poorly calibrated even when top-label accuracy looks strong

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

  • 07
    Translation, code-switching, OCR or speech transcription changes the input

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

  • 08
    New products, policies, seasons or user behaviour invalidate historical labels

    Record the affected label, input, channel, language, threshold and downstream action; preserve the narrowest conclusion that remains and specify whether to revise the ontology, relabel, resplit, recalibrate, restrict use, route to a human or withdraw the claim.

Minimum classification and routing capability evaluation record

Let the next reviewer reconstruct the conclusion with the same labels, inputs, annotation, pipeline, thresholds and complete denominator.

  1. 01Claim, date, owner, task, ontology, users, languages, channels, action and decision
  2. 02System, model, prompt, features, label descriptions, thresholds, rules and deployment versions
  3. 03Input provenance, time, sampling, prevalence, permissions, exclusions and target denominator
  4. 04Annotation guide, annotators, disagreements, adjudication, ambiguity and label changes
  5. 05Normalization, transcription, OCR, translation, context, truncation and routing trace
  6. 06Splits, duplicate and leakage controls, baselines, metric definitions, seeds and reference code
  7. 07Per-class precision, recall, F1, confusion, calibration, OOS/OOD and critical slices
  8. 08All inputs, empty and low-confidence cases, failures, repeats, fallbacks and corrections
  9. 09Wrong-route consequences, human review, latency, compute and total cost
  10. 10Target evidence, drift monitoring, recalibration, triggers, fallback, rollback and expiry

Common evidence states

Bind conclusions to the exact pipeline, label ontology, complete denominator, target action and date.

Supported

The exact pipeline meets declared per-class, calibration, unknown-input, critical-error, complete-denominator, latency, cost and target-routing thresholds, with a current review date.

Conditional

Support holds only for named labels, channels, languages, prevalence, thresholds, fallback controls or input conditions.

Mixed

Classes, languages, channels, calibration, unknown-input rejection, stability, latency or downstream outcomes differ materially.

Insufficient

Ontology, source population, annotation, pipeline trace, leakage control, classwise results, complete denominator, target validation or expiry is missing.

FUURAA analysisThe minimum decision unit for an AI text-classification and intent-routing capability claim is exact system, model, prompt, feature, label-description, threshold and routing version × classification task, label ontology, language, channel, user and decision × input provenance, time, sampling, annotation guide, disagreement, missingness and class prevalence × normalization, context, training, splits, baselines, scoring, calibration and abstention × classwise accuracy, precision, recall, F1, confusion, out-of-scope detection and critical errors × all inputs, empty results, low confidence, failures, variance, human correction, latency and cost × label and distribution drift, target-workflow boundary and cut-off date. Overall accuracy is an observation, not transferable proof of routing capability.

Primary sources and non-transfer boundaries

These sources constrain general English understanding, harder tasks, out-of-scope intent, original-Chinese classification, Chinese few-shot learning and multilingual intent evaluation.

Sources rechecked 31 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published 20 April 2018GLUE — General Language Understanding Evaluation

Combines nine English sentence and sentence-pair tasks with task-specific metrics and a diagnostic suite for transferable language understanding.

BoundaryAn aggregate benchmark score does not establish accuracy for a particular label ontology, live input stream, decision threshold, language or operating cost.

Open primary source ↗
Published 2 May 2019SuperGLUE — A Stickier Benchmark

Introduces harder English understanding tasks, standard formats, evaluation tools and human baselines after models saturated GLUE.

BoundaryHarder curated tasks still do not reproduce a production taxonomy, class imbalance, out-of-scope traffic, policy consequences or domain drift.

Open primary source ↗
Published 4 September 2019CLINC150 — Intent Classification and Out-of-Scope Prediction

Provides 150 intents across ten domains plus out-of-scope examples and shows that strong in-scope classifiers can still struggle to reject unsupported requests.

BoundaryBalanced, written English assistant queries and a fixed out-of-scope set do not represent another taxonomy, channel, language, adversarial input or live novelty rate.

Open primary source ↗
Published 13 April 2020CLUE — Chinese Language Understanding Evaluation

A Chinese-led benchmark brings together nine original-Chinese classification and reading tasks with baselines and linguistic diagnostics.

BoundarySelected simplified-Chinese datasets do not establish transfer to every Chinese variety, code-switching pattern, business taxonomy or current user population.

Open primary source ↗
Published 15 July 2021FewCLUE — Chinese Few-shot Learning Evaluation

A Chinese-led benchmark compares zero-shot, few-shot and fine-tuning schemes across nine Chinese tasks and reports sensitivity to the base model and method.

BoundaryFew-shot benchmark averages do not prove stable labels, representative examples, calibration, rare-class recall or safe production updates.

Open primary source ↗
Published 18 April 2022MASSIVE — Multilingual Intent and Slot Evaluation

Provides one million localized virtual-assistant utterances across 51 languages, 18 domains, 60 intents and 55 slots for multilingual intent and slot evaluation.

BoundaryTranslated parallel utterances and virtual-assistant domains do not establish natural local usage, code-switching, speech errors, new intents or another workflow.

Open primary source ↗

Continue checking

Move from classification and routing into tool use, multilingual capability, question answering or the full library.

Enter the AI Knowledge LibraryEvaluate tool useEvaluate multilingual capabilityEvaluate question answeringEnter AI Evidence Atlas