FUURAA AI Knowledge Library · question-answering evaluation

How to evaluate AI question-answering capability claims

Use six evidence gates to turn “answers anything” or “leading reading comprehension” into an exact task, questions and sources, answer contract, retrieval and context, support, completeness, abstention, all failures, human work, cost and a dated target-workflow boundary.

Published29 August 2026Evidence statusMethod synthesis grounded in primary QA and reading-comprehension researchScopeReading-comprehension, open-domain and multilingual AI QA research

A fluent answer is not support; one score is not complete capability

Question-answering capability must keep questions, sources, retrieval, answer support, completeness, abstention and every failure in one evidence chain.

Extractive spans, open-domain answers, multiple choice and conversational QA are different tasks with different errors. A reviewable capability judgment requires a fixed question and answer contract, then source, process, complete-denominator and real-use checks.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, language, dataset or leaderboard. It is not a guarantee of answer accuracy, completeness or fitness.

Six rejectable evidence gates

Each gate requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the question-answering task and answer contract

Decision question
What exact questions, sources, answer form, users and acceptance decision does the claim cover?
Minimum evidence
Verbatim claim and date; exact system, model, prompt, retrieval and tool versions; closed-book, provided-context, open-domain or conversational task; answer type, language, user group, consequence, scoring and abstention thresholds.
Stop condition
Stop when one benchmark score or demo answer replaces a named task, source condition, answer contract and decision threshold.
02

Version questions, sources and reference answers

Decision question
What did the system receive, what counted as truth and what information was absent or time-sensitive?
Minimum evidence
Question provenance, wording and distribution; source documents, corpus snapshot, retrieval cut-off and permissions; reference-answer authors, evidence spans, aliases, acceptable variants, disagreements, freshness dates, unanswerable labels and exclusions.
Stop condition
Stop when the question set, corpus version, reference process, answerability or fact date is unknown.
03

Separate access, retrieval, reading and reasoning

Decision question
Did the system find the right material and use it, or answer from memory, shortcuts or unsupported inference?
Minimum evidence
Input fidelity; retrieval recall, rank and no-result cases; context packing and truncation; passage, table and multi-document coverage; evidence localisation; coreference, comparison, arithmetic and synthesis slices; contamination and memorisation probes.
Stop condition
Stop when answer accuracy hides missing evidence, failed retrieval, truncated context, memorised answers or shortcut cues.
04

Verify support, completeness and abstention

Decision question
Is each material answer claim supported, sufficiently complete and calibrated when evidence is missing or conflicting?
Minimum evidence
Atomic answer claims mapped to source spans; exact match, token F1, semantic and task-specific scoring; numerical, temporal, entity and unit checks; answer completeness; contradiction tests; verified unanswerable, ambiguous and conflicting-source cases; confidence and refusal quality.
Stop condition
Stop when fluency, overlap or an unverified model judge substitutes for source support, completeness and correct abstention.
05

Count every question, failure, variance and human correction

Decision question
What happened across the full eligible set, repeated runs and important user or question slices?
Minimum evidence
Complete question denominator; empty, malformed, unsupported and partial answers; refusals, timeouts and tool failures; repeated-run variance; language, domain, answer type and subgroup results; retries, human corrections, latency, tokens, source fees and total cost.
Stop condition
Stop when selected successes hide unanswered items, unsupported detail, unstable repeats, human rescue, slow tails or cost.
06

Transfer to the target workflow and expire the claim

Decision question
Does the controlled result remain useful for target users, sources and consequences after information or the system changes?
Minimum evidence
Shadow or staged use; representative questions, source authorities, languages, accessibility needs and escalation routes; human acceptance and downstream outcome checks; monitoring; model, prompt, retrieval, corpus, source, policy and fact-change triggers; owner and expiry date.
Stop condition
Stop when a static benchmark transfers directly to live questions, current facts, new languages or consequential decisions without target validation.

Minimum QA claim failure matrix

Check these conditions deliberately before transferring selected answers or one score into real QA capability.

  • 01
    The answer is correct but unsupported by the supplied source

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

  • 02
    The relevant passage is absent, truncated or ranked below the context limit

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

  • 03
    An unanswerable or ambiguous question receives a confident fabricated answer

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

  • 04
    Exact match rejects a valid paraphrase or accepts an incomplete fragment

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

  • 05
    A numerical, temporal, entity or unit detail is wrong despite fluent wording

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

  • 06
    One language, domain, answer type or user group hides materially weaker results

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

  • 07
    Repeated runs change the answer, citation, refusal or confidence materially

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

  • 08
    A source, fact, corpus, retrieval system or model change invalidates prior evidence

    Record the affected question, source, answer type, language, user and decision; preserve the narrowest conclusion that remains and specify whether to add sources, relabel, rerun, review manually, restrict use or withdraw the claim.

Minimum QA capability evaluation record

Let the next reviewer reconstruct the conclusion with the same questions, sources, answer contract, system and complete runs.

  1. 01Claim, date, owner, QA task, answer contract, users, consequence and thresholds
  2. 02System, model, prompt, retrieval, tools, runtime and deployment versions
  3. 03Question set, provenance, wording, languages, domains, slices and exclusions
  4. 04Source corpus, snapshot, permissions, freshness, retrieval cut-off and preprocessing
  5. 05Reference answers, evidence spans, acceptable variants, disagreements and answerability
  6. 06Retrieval results, ranks, context packing, truncation and evidence localisation
  7. 07Answer claims, support, completeness, contradictions, confidence and abstention
  8. 08Metrics, judge versions, human validation, critical errors and subgroup results
  9. 09All questions, failures, repeats, retries, corrections, latency and total cost
  10. 10Target-workflow evidence, monitoring, change triggers, residual limits and expiry

Common evidence states

Bind conclusions to the exact system, task, sources, complete denominator, target workflow and date.

Supported

The exact system meets answer support, completeness, abstention, reliability, subgroup, human-effort, latency and cost thresholds on representative target questions, with a current review date.

Conditional

Support holds only for named question types, sources, corpus versions, languages, domains, answer forms or operating controls.

Mixed

Retrieval, support, completeness, refusal, critical errors, repeated-run reliability, latency or cost vary materially across slices.

Insufficient

Question or source provenance, answer contract, reference process, support checks, complete denominator, target transfer or expiry is missing.

FUURAA analysisThe minimum decision unit for an AI question-answering capability claim is exact system, model, prompt, retrieval and tool version × QA task, question distribution, language, user and answer contract × source provenance, corpus snapshot, permissions, freshness and reference process × access, retrieval, context, reading, reasoning and evidence localisation × answer support, completeness, critical errors, uncertainty and abstention × all questions, failures, variance, human correction, latency and cost × target-workflow boundary and cut-off date. One correct answer is an observation, not transferable proof of capability.

Primary sources and non-transfer boundaries

These sources constrain extractive QA, unanswerable questions, real search queries, paragraph reasoning, multilingual use and Chinese open-domain reading comprehension.

Sources rechecked 29 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published 16 June 2016SQuAD — 100,000+ Questions for Machine Comprehension of Text

Established large-scale English extractive reading comprehension over Wikipedia passages with span-based answer scoring.

BoundaryCrowd-written questions, Wikipedia passages and extractive spans do not establish open-domain retrieval, synthesis, multilingual use, abstention or production answer quality.

Open primary source ↗
Published 11 June 2018SQuAD 2.0 — Know What You Do Not Know

Adds adversarially written unanswerable questions so systems must answer when supported and abstain when the passage lacks an answer.

BoundaryPassage-level abstention under a fixed dataset does not prove calibrated uncertainty, safe refusal or evidence handling across open and changing information.

Open primary source ↗
Published July 2019Natural Questions — A Benchmark for Question Answering Research

Uses real anonymised search queries with long, short or null answers annotated against retrieved Wikipedia pages.

BoundarySearch queries paired with one selected Wikipedia page do not represent every information need, source set, freshness requirement, user group or consequential decision.

Open primary source ↗
Published 1 March 2019DROP — Discrete Reasoning Over Paragraphs

Tests reference resolution plus counting, addition, sorting and other discrete operations over English paragraphs.

BoundaryParagraph arithmetic and exact-answer scoring do not establish broad reasoning, factual freshness, source selection or reliable multi-step workflow completion.

Open primary source ↗
Published 10 March 2020TyDi QA — Information-Seeking QA in Typologically Diverse Languages

Collects information-seeking questions directly in eleven typologically diverse languages rather than translating an English task.

BoundaryEleven languages and Wikipedia-based annotations do not cover every language variety, culture, domain, source format or production interaction.

Open primary source ↗
Published 14 November 2017DuReader — Chinese Machine Reading Comprehension from Real-world Applications

A Chinese-led open-domain resource uses questions and documents from Baidu Search and Baidu Zhidao with manually generated answers, including yes-no and opinion questions.

BoundaryIts platforms, Chinese data distribution and reference process do not represent all Chinese varieties, regions, source authorities, current facts or deployed answer workflows.

Open primary source ↗

Continue checking

Move from question answering into factuality, RAG, reasoning or the full library.

Enter the AI Knowledge LibraryEvaluate factualityEvaluate RAGEvaluate reasoningEnter AI Evidence Atlas