FUURAA AI Knowledge Library · Factuality and hallucination claim evaluation

How to evaluate AI factuality and hallucination claims

Use six evidence gates to turn “fewer hallucinations” into an exact task, versioned reference facts, atomic claims, independent support relations, complete denominators, uncertainty and a dated conclusion. Apply it to model reports, product documentation, research, media verification, diligence and engineering review.

Published25 August 2026Evidence statusMethod synthesis grounded in primary risk and factuality evaluation researchScopeModel reports, product documentation, research, media verification, diligence and engineering review

A credible tone is not factual evidence

Factuality must return to checkable claims, applicable sources, complete denominators, time boundaries and uncertainty—not how convincing an answer sounds.

“Hallucination” may mean conflict with supplied context, conflict with world facts, unverifiability, citation mismatch, temporal staleness or unsupported inference. Different definitions produce different scores, so the claim must be defined before evidence and metrics are selected.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, fact-checking service, dataset, benchmark or judge. It is not an accuracy guarantee, certification, audit, medical, legal, investment, procurement or compliance opinion.

Six rejectable evidence gates

Each gate answers one decision question, requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the exact factuality claim and decision

Decision question
Does “factual” mean source-grounded, world-correct, internally consistent, current, complete or appropriately uncertain; for which users, tasks and consequences?
Minimum evidence
Verbatim claim, claimant and date; model and system version; task, domain and language; factuality definition, unit of analysis, quality floor, allowed evidence and decision context.
Stop condition
Stop when fluency, confidence, citation count or one aggregate score is used as a universal definition of factuality.
02

Specify questions, references, time and answerability

Decision question
Which questions have stable answers, which change over time, which contain false premises, and which cannot be resolved from the permitted evidence?
Minimum evidence
Prompt set and sampling frame, reference answer and source version, evidence cut-off, jurisdiction and locale, false-premise cases, unanswerable cases, ambiguity policy and update schedule.
Stop condition
Stop when a stale answer key grades current facts, disputed claims are treated as settled, or forced answering turns missing evidence into an error attributed only to the model.
03

Decompose outputs into checkable claims

Decision question
What is the smallest independently verifiable claim, and how are hedges, dates, quantities, entities, relations, causal language and implied facts preserved?
Minimum evidence
Atomic-claim protocol, segmentation examples, inclusion rules, attribution and quote handling, claim identifiers, omitted-content review, annotator training, agreement and adjudication.
Stop condition
Stop when one supported sentence hides unsupported clauses, qualification is stripped, or a passage-level label conceals mixed factuality.
04

Match every claim to independent evidence

Decision question
Does accessible evidence directly support, contradict or leave the exact claim unresolved, and is the evidence authoritative and applicable to its date, population and scope?
Minimum evidence
Exact source, version and location; retrieval query and result set; publication status; support relation; conflicting sources; independence; reviewer rationale and checked date.
Stop condition
Stop when a resolving URL is treated as support, search snippets replace sources, the model verifies itself, or evidence supports a nearby but different claim.
05

Measure errors, abstention and uncertainty on complete denominators

Decision question
How often are claims supported, contradicted, unverifiable, omitted or overqualified; when does the system abstain, and does expressed confidence track observed correctness?
Minimum evidence
Claim-level precision and coverage, response-level success, contradiction and unverifiable rates, abstention and over-refusal, calibration by confidence band, uncertainty, subgroup results and all failed outputs.
Stop condition
Stop when unsupported claims are removed from the denominator, shorter answers win by saying less without disclosure, or refusal is counted as factual success regardless of usefulness.
06

Test real use, change and expiry

Decision question
Does evidence survive authentic prompts, current events, retrieval outages, adversarial framing, multiple languages and system changes; which change requires revalidation?
Minimum evidence
Production-like task distribution, fresh and false-premise probes, source and retrieval logs, incident samples, user correction path, material-change log, monitoring window, expiry and named owner.
Stop condition
Stop when a static benchmark becomes a permanent deployment claim, the evidence index or model changes without retest, or users cannot challenge and correct consequential falsehoods.

Minimum factuality-claim failure matrix

Check these conditions deliberately to expose source mismatch, temporal staleness and denominator errors behind credible tone, real links and attractive scores.

  • 01
    A fluent answer with real citations makes an unsupported central claim

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

  • 02
    A stale answer key marks a current fact wrong—or a stale model answer correct

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

  • 03
    One supported clause hides an invented date, quantity or causal relation

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

  • 04
    The evaluator and generator share the same misconception or source

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

  • 05
    Search snippets are accepted without opening the exact source and passage

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

  • 06
    Unsupported claims disappear from the denominator after filtering

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

  • 07
    Refusing every difficult question produces an apparently perfect score

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

  • 08
    A model, retrieval index, prompt or world fact changes without revalidation

    Record affected claims, sources, dates, users, tasks and decisions; preserve the narrowest statement that remains and specify whether to add evidence, relabel, segment results, retest or withdraw the conclusion.

Minimum factuality evaluation record

Let the next reviewer reconstruct the conclusion under the same system, questions, reference facts, sources, judging and denominator boundaries.

  1. 01Exact claim, claimant, task, decision, checked date and expiry
  2. 02Model, endpoint, prompt, tools, retrieval and system versions
  3. 03Prompt sample, domain, language, population and consequence tier
  4. 04Reference sources, versions, locations and evidence cut-off
  5. 05Stable, fresh, disputed, false-premise and unanswerable labels
  6. 06Atomic claims, qualifiers, omissions, attributions and identifiers
  7. 07Support, contradiction and unresolved judgments with rationale
  8. 08Precision, coverage, abstention, calibration, failures and uncertainty
  9. 09Reviewer fit, independence, agreement and adjudication
  10. 10Narrowest supported conclusion, material unknowns and revalidation owner

Common evidence states

Bind conclusions to the exact system, question distribution, reference facts, support protocol, denominator and date—not a generic “hallucination rate”.

Supported

Exact system and scope meet declared factual precision, coverage, abstention and freshness thresholds under a reproducible protocol.

Conditional

Evidence supports only named domains, question types, sources, languages, dates or consequence levels; transfer remains bounded.

Mixed

Supported and unsupported claims, freshness results, judges or subgroups diverge materially and require segmented reporting.

Insufficient

Fluency, citations, one benchmark, self-consistency or an undocumented demonstration cannot establish factual reliability.

FUURAA analysisThe minimum decision unit for an AI factuality and hallucination claim is exact system and version × questions, domain, language and time × atomic claims and qualifiers × source version and support relation × precision, coverage, abstention and uncertainty × cut-off date. Citation verification asks whether one source supports one sentence; factuality evaluation must also explain how questions were sampled, facts updated, all claims entered the denominator and the system knew when to abstain.

Primary sources and non-transfer boundaries

These sources constrain generative-AI risk, common misconceptions, atomic facts, hallucination recognition, dynamic knowledge and search-augmented judging; none independently proves general factual reliability.

Sources rechecked 25 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published 26 July 2024; updated 8 April 2026NIST AI 600-1 — Generative AI Profile

Identifies confabulation as a generative-AI risk and links measurement, provenance, human review and post-deployment monitoring to risk management.

BoundaryThe profile is voluntary cross-sector guidance; it does not define one universal hallucination metric or certify any system as factual.

Open primary source ↗
First submitted 8 September 2021TruthfulQA — Measuring How Models Mimic Human Falsehoods

Tests whether models reproduce common misconceptions across 817 questions and 38 categories, separating truthfulness from simple imitation.

BoundaryIts curated questions and tested model snapshot do not establish current-world freshness, long-form factuality, every language or deployment domain.

Open primary source ↗
First submitted 23 May 2023FActScore — Fine-grained Atomic Evaluation of Factual Precision

Decomposes long-form generations into atomic facts and measures the proportion supported by a specified knowledge source.

BoundaryAtomic support depends on the chosen source, retrieval and evaluator; factual precision alone does not measure completeness, relevance or decision fitness.

Open primary source ↗
First submitted 19 May 2023Renmin University of China and collaborators — HaluEval

Provides generated and human-annotated hallucination samples for question answering, dialogue, summarisation and hallucination recognition.

BoundarySynthetic generation and annotation choices shape difficulty; recognition of a labelled hallucination is not proof that a system avoids one in live use.

Open primary source ↗
First submitted 5 October 2023Google Research and collaborators — FreshQA / FreshLLMs

Uses dynamic questions, fast-changing knowledge and false premises to separate correctness from hallucination under changing real-world facts.

BoundarySearch augmentation depends on evidence freshness, ranking and prompt order; benchmark gains do not prove every retrieved answer is current or supported.

Open primary source ↗
First submitted 27 March 2024Google DeepMind and collaborators — Long-form Factuality / SAFE

Introduces LongFact and a search-augmented evaluator that decomposes responses, searches for evidence and judges support for individual facts.

BoundaryAutomated agreement is conditional on search results, decomposition, judge and prompt; it does not replace expert review for consequential claims.

Open primary source ↗

Continue checking

Move from factuality into citation verification, research-synthesis review and benchmark-claim reading.

Verify AI citationsReview research synthesesRead benchmark claimsEnter AI Evidence Atlas