FUURAA AI Knowledge Library · Evidence literacy guide

How to decide whether an AI claim deserves belief

A practical route from claim, source, evaluation design and technical artefacts to applicability boundaries and a reviewable judgment. Use it for papers, model releases, leaderboards, demos, policy reports and industry narratives.

Published4 August 2026Evidence statusMethod guide grounded in primary standards and researchScopeInitial and deep review of public AI claims

Set the boundary first

A credible source does not make every conclusion transferable.

Evidence strength depends on the match between claim and record: who measured what, using which version, under what conditions, against which comparison, for how long, and for which decision.

Applicability boundaryThis guide helps organise public records and expose gaps. It does not verify hidden data, private systems or signer identity, and it does not replace domain-expert review, experimental replication, safety testing, legal advice or certification.

Six-pass review

Every pass should leave a handoff-ready record.

Do not begin with agreement or rejection. Make the claim testable, then increase only the conclusion strength the record can carry.

01

Write the smallest testable claim

Core question
What exactly is being claimed about which system, version, task, population, place and time?
Evidence to preserve
A bounded claim statement, named decision and explicit excluded interpretations.
Stop condition
Stop if the subject or comparison can change without changing the sentence.
02

Classify the source before reading the result

Core question
Is this primary research, a standard, official data, developer disclosure, independent replication, journalism or commentary?
Evidence to preserve
Canonical URL, authorship, publication date, version, correction history and funder or developer relationship.
Stop condition
Do not let a summary inherit more authority than the record it summarizes.
03

Audit the evaluation design

Core question
Do task, baseline, sample, metric, uncertainty and failure criteria match the public claim?
Evidence to preserve
Protocol, dataset split, baseline choice, metric definition, confidence or variation, and raw outputs where available.
Stop condition
A leaderboard position alone cannot establish broad capability or deployment fitness.
04

Trace data, model and artefact provenance

Core question
Can another reviewer identify the exact data, model, code, prompt, environment and licence conditions?
Evidence to preserve
Model card, dataset documentation, immutable versions, code or configuration, dependency record and stated rights.
Stop condition
Reproducible code does not cure unsuitable data, and open weights do not imply open training data.
05

Test transfer and deployment boundaries

Core question
What changes when language, population, workflow, incentives, latency, cost or human oversight changes?
Evidence to preserve
Subgroup results, local evaluation, operating constraints, incident modes, oversight design and post-deployment monitoring.
Stop condition
Treat untested transfer as an open question, not as a weaker version of proof.
06

Record a decision-sized evidence state

Core question
What does the record support now, what would change the judgment, and when must it be reviewed again?
Evidence to preserve
Supported, developing, mixed or insufficient status; dated rationale; contradictions; owner; next review trigger.
Stop condition
Do not compress a multidimensional record into one confidence percentage without a defensible model.

Minimum evidence record

Nine fields turn ‘I read it’ into ‘someone else can review it’.

  1. 01claim

    One bounded sentence

  2. 02decision

    The decision this review informs

  3. 03subject

    System, model, version and operator

  4. 04source

    Canonical record, author and date

  5. 05method

    Task, sample, baseline and metrics

  6. 06artefacts

    Data, code, configuration and documentation

  7. 07boundary

    Population, geography, workflow and excluded uses

  8. 08state

    Supported, developing, mixed or insufficient

  9. 09review

    Owner, review date and change trigger

Evidence states

States describe the public record without pretending to mathematical certainty.

supported

Supported

Several relevant records converge, material contradictions are addressed and the conclusion stays within the tested boundary.

developing

Developing

Credible evidence exists, but replication, coverage, duration or deployment observation remains incomplete.

mixed

Mixed

Credible records disagree, or the result changes materially across methods, populations or operating conditions.

insufficient

Insufficient

The public record is too sparse, indirect, outdated or poorly matched to support the decision.

FUURAA analysisA useful evidence state matches the granularity of the decision. A model result on one fixed English benchmark may be supported while ‘fit for multilingual Southeast Asian public services’ remains insufficient. The states do not conflict because they answer different claims.

Fast rejection conditions

Lower the conclusion strength when these signals appear.

  • 01No exact model or system version
  • 02Only the best result is shown; variation and failures are absent
  • 03Test, training or tuning data may contaminate one another
  • 04Baselines are irrelevant or outdated
  • 05Developer disclosure is presented as independent verification
  • 06One metric stands in for cost, fairness, robustness or safety
  • 07English-only, single-region or laboratory results are extrapolated to global deployment
  • 08A source was updated without stating which conclusions changed

Primary sources and boundaries

The method is grounded in public records; FUURAA analysis remains separate from source findings.

Sources checked 4 August 2026. Publication dates, evidence roles and boundaries follow each source's public record.

Official framework · primary sourceNIST AI RMF 1.0

Connects AI evaluation to context, affected people, measurement, governance and continuing risk management across the lifecycle.

BoundaryVoluntary, cross-sector and use-case agnostic; it is not a sector certification or a complete evaluation protocol.

Open primary source ↗
Official GenAI profile · primary sourceNIST AI 600-1 · Generative AI Profile

Adds generative-AI-specific risks and actions to the AI RMF, including measurement, provenance, information integrity and human oversight concerns.

BoundaryA cross-sector companion resource; actions still need prioritisation for the actual system and risk context.

Open primary source ↗
Peer-reviewed research · documentation methodModel Cards for Model Reporting

Shows why intended use, performance conditions and relevant subgroup evaluation should accompany released models.

BoundaryA reporting framework cannot independently verify the developer's measurements or make an unsuitable model safe.

Open primary source ↗
Research paper · dataset documentationDatasheets for Datasets

Provides a structured way to inspect dataset motivation, composition, collection, use and limitations before treating results as transferable.

BoundaryDocumentation quality depends on disclosure quality and does not replace data inspection, rights review or local validation.

Open primary source ↗
Peer-reviewed research · evaluation designHELM · Holistic Evaluation of Language Models

Demonstrates scenario coverage, multi-metric evaluation, standardised comparison and disclosure of missing or underrepresented tests.

BoundaryNo benchmark suite can cover every deployment, and living benchmarks change as models, scenarios and methods evolve.

Open primary source ↗

Continue with the record

Apply the method to a real question, then enter the Atlas.

The AI Knowledge Library helps locate material; AI Evidence Atlas connects consequential claims to evidence, disagreement, limits and unknowns.

Enter AI Evidence AtlasReturn to AI Knowledge LibraryRead FUURAA research methodology