FUURAA AI Knowledge Library · Summarization evaluation

How to evaluate AI text summarization capability claims

Use six evidence gates to turn “summarises well” into an exact task, source population, complete pipeline, source support, salient coverage, critical omissions, compression and presentation quality, all failures, human correction, latency, cost and a dated target-workflow boundary.

Published28 August 2026Evidence statusMethod synthesis grounded in primary summarization research and evaluation benchmarksScopeExtractive, abstractive, single-document and multi-document summarization research

Short, fluent and human-like is not the same as supported, complete or fit for use

Summarization capability must keep source support, salient coverage, critical omissions, conflicts, failures and human correction in one evidence chain.

High ROUGE, fluent tone or one judge score can still hide wrong attribution, omitted conditions and source truncation. Test factual support, coverage, compression and presentation separately before returning to the target user and decision.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, summary, institution, dataset or leaderboard. It is not a guarantee of accuracy, completeness, fitness or decision quality.

Six rejectable evidence gates

Each gate answers one decision question, requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the summarization task and acceptance decision

Decision question
Is the claim about extraction, abstraction, one or many sources, headlines, briefs, meeting minutes, updates, comparison or decision support?
Minimum evidence
Verbatim claim and date; source type and count; target audience and decision; required length, format, language, perspective and citations; inclusion, exclusion, abstention and acceptance rules.
Stop condition
Stop when “summarises well” or one polished example replaces a defined source population, output contract and consequence.
02

Reproduce the complete source-to-summary pipeline

Decision question
Which exact model, prompt, context, retrieval, chunking, ranking, merge, post-processing and human steps produced the accepted summary?
Minimum evidence
Provider, model and API versions; system and user prompts; source ordering and truncation; context window; chunk and overlap rules; retrieval and reranking; tools; decoding; retries; templates; edits and reviewer instructions.
Stop condition
Stop when omitted source sections, hidden retrieval, template rules or human rewriting are credited to the summarizer.
03

Define source provenance, coverage and challenge strata

Decision question
Which languages, domains, lengths, structures, dates, conflicts, redundancies and source qualities are represented?
Minimum evidence
Source set and hashes; origin, date, permission and sampling frame; deduplication; train or benchmark overlap; strata for language, genre, length, chronology, number of documents, tables, dialogue, contradictions, updates and low-quality inputs.
Stop condition
Stop when short English news is transferred to long reports, meetings, scientific material, Chinese content or multi-source synthesis without direct evidence.
04

Separate support, coverage, compression and presentation quality

Decision question
Does evaluation distinguish source support from salient coverage, omission, contradiction, redundancy, coherence, readability, format and audience fit?
Minimum evidence
Atomic summary claims linked to source spans; entity, number, date, negation and attribution checks; required-topic coverage; critical omission and contradiction review; compression ratio; redundancy; structure; human rubric; inter-rater agreement and uncertainty.
Stop condition
Stop when ROUGE, one reference summary, one LLM judge or fluency is treated as proof of factual support, completeness and usefulness.
05

Count every source, output, failure and correction

Decision question
Which sources or sections were missing, truncated, ignored, contradicted, guessed, refused, retried or corrected before acceptance?
Minimum evidence
Complete source and run denominator; unsupported and contradicted claims; critical omissions; duplicate or stale sources; refusals; empty and malformed outputs; timeouts; retries; selection; edits; reviewer time; latency and cost distributions.
Stop condition
Stop when failed runs, omitted documents or human corrections disappear from the reported success, quality and cost denominator.
06

Test the target workflow, change triggers and expiry

Decision question
Does the result survive target sources, users, deadlines, review procedures, updates and downstream decisions without exceeding its evidence?
Minimum evidence
Target-workflow sample; critical-content map; reviewer and escalation rules; citation and source-opening procedure; monitoring; model, prompt, source and policy change triggers; fallback; owner; review date and expiry.
Stop condition
Stop when benchmark or demo evidence is used for consequential decisions without target-source validation and source-linked human review.

Minimum summarization-capability claim failure matrix

Check these conditions deliberately before transferring overlap, fluency or one reference summary into real summarization capability.

  • 01
    A fluent summary invents a cause, quote, number, date, actor or relationship absent from the source

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

  • 02
    The summary is factually supported but omits the exception, uncertainty, dissent or condition that changes the decision

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

  • 03
    One wrong negation, unit, denominator or attribution reverses the meaning while average overlap remains high

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

  • 04
    Early or retrieved sections dominate while late, low-ranked, tabular or cross-document evidence disappears

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

  • 05
    Conflicting sources are silently merged into one confident statement instead of preserving disagreement and dates

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

  • 06
    A headline-style result is presented as evidence for long-form, meeting, scientific or multilingual summarization

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

  • 07
    Failed generations, refusals, truncation, retries, selected candidates and human rewrites vanish from the denominator

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

  • 08
    An automatic metric or LLM judge grades style highly while missing unsupported claims or critical omissions

    Record the affected claim, source, language, position, summary layer and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample sources, rerun, open the source for review, correct manually or withdraw the claim.

Minimum text-summarization capability evaluation record

Let the next reviewer reconstruct the conclusion with the same sources, pipeline, versions, layer-specific measures, complete runs, human protocol and target boundary.

  1. 01Exact claim, date, summarization task, source population, audience, decision and acceptance rubric
  2. 02Provider, model, API, prompt, tool, retrieval, reranking, post-processing and human versions
  3. 03Source ordering, context limit, truncation, chunk, overlap, merge, decoding, retry and output settings
  4. 04Source set, provenance, dates, permission, hashes, sampling, deduplication and benchmark overlap
  5. 05Language, domain, genre, length, chronology, source-count, conflict and input-quality strata
  6. 06Atomic source support, entities, numbers, dates, negation, attribution and contradiction results
  7. 07Required-topic coverage, critical omissions, compression, redundancy, coherence, format and audience fit
  8. 08All sources and runs, truncations, omissions, refusals, malformed outputs, retries, selections and edits
  9. 09Human rubric, source-opening protocol, reviewer agreement, correction time, latency, cost and uncertainty
  10. 10Target boundary, monitoring, escalation, fallback, change triggers, owner, review date and expiry

Common evidence states

Bind conclusions to the exact pipeline, source population, summary contract, complete denominator, human review, target workflow and date.

Supported

Evidence supports the exact pipeline, source population, summary contract, source-linked claims, critical coverage, complete denominator, human review and dated target boundary.

Conditional

Evidence supports a narrower language, domain, source length, summary format, compression target, review process or downstream use.

Mixed

Results vary materially across factual support, coverage, source position, length, language, domain, conflict, reviewer, latency or cost.

Insufficient

Pipeline identity, representative sources, source-linked review, critical-omission checks, full denominators, corrections, costs or target transfer evidence is missing.

FUURAA analysisThe minimum decision unit for an AI text-summarization claim is exact pipeline and version × summarization task, source population, audience and decision × source provenance, language, domain, length, structure, time and conflict × prompt, context, chunking, retrieval, ranking, merging and post-processing × source support, salient coverage, critical omissions, contradictions, compression and presentation quality × all sources, runs, failures, abstentions, human correction, latency and cost × target-workflow boundary and cut-off date. ROUGE, LCSTS, XSum, SummEval, FRANK and AGGREFACT illuminate overlap, Chinese short text, extreme summarization, broad evaluation and factuality; none independently represents real summarization capability.

Primary sources and non-transfer boundaries

These sources constrain overlap, Chinese short text, one-sentence news, human evaluation and factuality detection; none independently proves real-workflow capability.

Sources rechecked 28 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published July 2004ROUGE — A Package for Automatic Evaluation of Summaries

Defines recall-oriented overlap measures that became a widely used baseline for comparing system and reference summaries.

BoundaryN-gram or sequence overlap does not independently establish factual support, coverage of decision-critical content, coherence, usefulness or safety.

Open primary source ↗
First submitted 19 June 2015LCSTS — Large-Scale Chinese Short Text Summarization

Introduces more than two million real Chinese short-text and summary pairs from Sina Weibo, plus 10,666 manually relevance-labelled pairs.

BoundaryShort social-media text does not establish performance on long reports, meetings, specialised domains, every Chinese variety or production review workflows.

Open primary source ↗
First submitted 27 August 2018XSum — Extreme Summarization

Introduces a one-sentence, highly abstractive news summarization task built from BBC articles and their short summaries.

BoundaryOne-sentence BBC news summarization does not establish multi-document synthesis, faithful long-form compression, meeting minutes, scientific summaries or target-workflow quality.

Open primary source ↗
First submitted 24 July 2020; revised 1 February 2021SummEval — Re-evaluating Summarization Evaluation

Compares 14 automatic metrics and 23 summarization models with expert and crowd human judgements on CNN/DailyMail outputs.

BoundaryMetric correlations on English news and a fixed model collection do not guarantee validity for new systems, languages, domains, summary lengths or high-stakes uses.

Open primary source ↗
First submitted 27 April 2021FRANK — Factuality in Abstractive Summarization

Provides a typology and human annotations for factual errors across model-generated CNN/DailyMail and XSum summaries, enabling metric comparison by error type.

BoundaryNews-focused factuality annotations do not measure salience, completeness, audience fit, decision usefulness or every error pattern from newer models and domains.

Open primary source ↗
First submitted 25 May 2022; revised 26 May 2023AGGREFACT — Understanding Factual Errors in Summarization

Aggregates nine annotated factuality datasets and stratifies results by summarizer generation and error type, showing that detector performance varies materially.

BoundaryA factuality detector benchmark does not establish overall summary quality, support every current generator, replace source-linked human review or transfer automatically to a target domain.

Open primary source ↗

Continue checking

Move from summarization into factuality, long context, multilingual and document-understanding evaluation, or the full library.

Enter the AI Knowledge LibraryEvaluate factualityEvaluate long contextEvaluate multilingual capabilityEvaluate document understandingEnter AI Evidence Atlas