FUURAA AI Knowledge Library · Document intelligence evaluation

How to evaluate AI document understanding and OCR capability claims

Use six evidence gates to turn “understands documents” into an exact task, document population, acquisition and rendering pipeline, OCR and model version, transcription, reading order, layout, tables, forms, key information, document VQA, complete pages and fields, human correction, latency, cost and a dated target-workflow boundary.

Published28 August 2026Evidence statusMethod synthesis grounded in primary Document AI and OCR researchScopeOCR, layout, tables, forms, key-information extraction and document-VQA research

Reading characters is not the same as understanding structure or supporting a decision

A document-capability claim must keep pages, fields, layout relations, failures, abstentions, human correction and downstream use in one evidence chain.

A page-level average can hide one critical digit, wrong reading order or an omitted page; a fluent answer can also lack document support. OCR, layout analysis, table reconstruction, key-information extraction and document VQA should be tested separately before returning to the real workflow.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, document, institution, dataset or leaderboard. It is not an accuracy guarantee, document-authenticity or completeness determination, identity confirmation, rights clearance, professional judgment, procurement, audit, investment, legal or compliance opinion.

Six rejectable evidence gates

Each gate answers one decision question, requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the document task and downstream decision

Decision question
Is the claim about text detection, transcription, reading order, layout, tables, forms, key information, document VQA, classification or an end-to-end workflow?
Minimum evidence
Verbatim claim and date; document task; input and output contract; required fields, units and relations; downstream user and decision; quality floor, abstention rule and acceptance rubric.
Stop condition
Stop when “reads documents” or one demo replaces a defined document population, output schema and decision consequence.
02

Reproduce acquisition, rendering and preprocessing

Decision question
How did native files, scans or photographs become pages, regions, tokens, tables and answers, under which exact system versions?
Minimum evidence
Provider, model, OCR engine, parser, prompt and API versions; native PDF or raster path; resolution and DPI; orientation, crop, dewarp, denoise, contrast, compression, page split, language detection, table reconstruction, retrieval and post-processing.
Stop condition
Stop when native PDF text, high-resolution scans and phone photographs are pooled, or hidden rules and human edits are attributed to the model.
03

Define document provenance, coverage and challenge strata

Decision question
Which languages, scripts, templates, page qualities, layouts, handwriting, tables, stamps and multi-page dependencies are represented?
Minimum evidence
Document set and hashes; source, permission and sampling frame; deduplication; train or benchmark overlap; strata for language, script, genre, template, scan quality, font size, rotation, handwriting, table density, page count and temporal drift.
Stop condition
Stop when clean English pages or one stable template are presented as evidence for multilingual, handwritten, historical or open-world documents.
04

Separate transcription, structure, extraction and answer validity

Decision question
Does evaluation distinguish text detection and characters from reading order, layout, tables, entity relations, key fields and evidence-supported answers?
Minimum evidence
Detection precision and recall; CER or WER with normalisation rules; reading-order and layout metrics; table cell and structure scores; entity and linking results; document-VQA scoring; field-level exactness, units, confidence, unsupported-answer rate and human validation.
Stop condition
Stop when page-level accuracy, edit distance, one VQA score or one multimodal judge is treated as proof of every document-understanding layer.
05

Count every page, omission, abstention and correction

Decision question
Which pages, fields and cells were skipped, duplicated, truncated, guessed, sent to review or corrected before the accepted output?
Minimum evidence
Complete page and field denominator; blank, corrupt, encrypted, unsupported and timed-out inputs; duplicate or missing pages; empty and low-confidence outputs; abstentions, unsupported answers, retries, corrections, reviewer time, latency and cost distributions.
Stop condition
Stop when failed pages disappear, only populated fields are scored, or human correction is omitted from the success and cost denominator.
06

Test target workflow, sensitive handling and expiry

Decision question
Does the result survive target documents, access controls, review procedures, template changes and downstream use without exceeding its evidence?
Minimum evidence
Target-workflow sample; field-criticality map; reviewer and escalation rules; access, retention and redaction configuration; prompt-injection and hidden-text tests where relevant; monitoring, change triggers, fallback, owner, review date and expiry.
Stop condition
Stop when a benchmark average is transferred into identity, authenticity, financial, medical, contractual or other consequential decisions without direct target-context validation.

Minimum document-understanding claim failure matrix

Check these conditions deliberately before transferring clean pages, character averages or one VQA result into end-to-end document understanding.

  • 01
    High character accuracy hides one wrong digit, date, name, unit or decimal that changes the downstream decision

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

  • 02
    Text is transcribed correctly but reading order, columns, key–value links or table cells are wrong

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

  • 03
    A page average hides small print, handwriting, stamps, rotated text or a minority script

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

  • 04
    A fluent document answer is unsupported, copied from the wrong page or detached from its source region

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

  • 05
    Native PDFs perform well while photographs, compressed scans or multi-page packets fail

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

  • 06
    Skipped, blank, encrypted, duplicate or truncated pages disappear from the reported denominator

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

  • 07
    Template rules, retrieval, manual correction or spreadsheet cleanup are credited to the model

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

  • 08
    A benchmark result is transferred to invoices, applications, records or agreements without field-level target review

    Record the affected page, field, language, template, processing stage and downstream decision; preserve the narrowest conclusion that remains and specify whether to resample, rerender, rerun, review at field level, correct manually or withdraw the claim.

Minimum document-understanding and OCR capability evaluation record

Let the next reviewer reconstruct the conclusion with the same documents, pipeline, versions, layer-specific metrics, complete pages and fields, human protocol and target boundary.

  1. 01Exact claim, date, document task, output schema, downstream user, decision and acceptance rubric
  2. 02Provider, model, OCR engine, parser, API, prompt, retrieval and post-processing versions
  3. 03Native or raster path, DPI, resolution, orientation, crop, dewarp, denoise, compression and page handling
  4. 04Document set, provenance, permission, hashes, sampling, deduplication and benchmark overlap
  5. 05Language, script, genre, template, quality, handwriting, table and multi-page strata
  6. 06Separate detection, transcription, reading-order, layout, table, entity, linking and answer results
  7. 07Normalisation, field criticality, confidence, unsupported-answer checks, human protocol and uncertainty
  8. 08All pages and fields, omissions, duplicates, corrupt inputs, abstentions, retries and corrections
  9. 09Field-level error, subgroup, reviewer time, latency, cost and downstream outcome distributions
  10. 10Target boundary, sensitive handling, escalation, monitoring, fallback, change triggers, owner and expiry

Common evidence states

Bind conclusions to the exact pipeline, document population, task layers, complete denominator, human review, target workflow and date.

Supported

Evidence supports the exact pipeline, document population, task layers, complete pages and fields, human review and dated target-workflow boundary.

Conditional

Evidence supports a narrower file type, template, language, script, image quality, field set, extraction task, review process or downstream use.

Mixed

Results vary materially across transcription, layout, tables, fields, questions, languages, templates, page qualities, confidence thresholds or reviewers.

Insufficient

Pipeline identity, representative documents, layer-specific metrics, full denominators, unsupported-answer controls, corrections, costs or target transfer evidence is missing.

FUURAA analysisThe minimum decision unit for an AI document-understanding and OCR claim is exact pipeline and version × document task, output schema, downstream user and decision × document type, language, script, template, quality, layout and cross-page relations × acquisition, rendering, preprocessing, OCR, model, retrieval and post-processing × detection, transcription, order, structure, fields, answers and critical errors × all pages, omissions, abstentions, human correction, latency and cost × target-workflow boundary and cut-off date. FUNSD, PubLayNet, DocVQA, LayoutXLM/XFUND, LayoutLMv3 and OCRBench illuminate different layers; none independently represents end-to-end document intelligence.

Primary sources and non-transfer boundaries

These sources constrain forms, layout, document VQA, multilingual documents, unified text–image modelling and broad OCR testing; none independently proves real-workflow capability.

Sources rechecked 28 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

First submitted 27 May 2019; revised 29 October 2019FUNSD — Form Understanding in Noisy Scanned Documents

Provides 199 annotated, noisy scanned forms for text detection, OCR, spatial layout analysis and entity labelling or linking.

BoundaryA small form corpus does not establish performance on every template, script, handwriting style, image quality, table or production document stream.

Open primary source ↗
First submitted 16 August 2019PubLayNet — Large-Scale Document Layout Analysis

Builds layout annotations from more than one million PubMed Central PDFs, with over 360,000 document images for common scientific-article elements.

BoundaryScientific-article layout detection does not establish reading order, OCR transcription, tables, forms, receipts, handwriting, multilinguality or downstream correctness.

Open primary source ↗
First submitted 1 July 2020; revised 5 January 2021DocVQA — Visual Question Answering on Document Images

Introduces more than 50,000 questions over 12,000 document images and shows that structure-sensitive questions remain difficult relative to human performance.

BoundaryQuestion-answer accuracy on its documents does not verify complete transcription, unsupported-answer control, every page type or the correctness of a downstream decision.

Open primary source ↗
First submitted 18 April 2021; revised 9 September 2021LayoutXLM and XFUND — Multilingual Visually-Rich Document Understanding

Combines text, layout and image pre-training and introduces XFUND with manually labelled key–value pairs across seven languages, including Chinese and Japanese.

BoundarySeven-language form results do not establish every script, locale, code-switching pattern, document genre, translation step or real multilingual workflow.

Open primary source ↗
First submitted 18 April 2022; revised 19 July 2022LayoutLMv3 — Unified Text and Image Masking for Document AI

Uses unified text and image masking plus word–patch alignment across form, receipt, document VQA, classification and layout-analysis tasks.

BoundaryBenchmark gains from a pre-trained architecture do not establish a complete OCR pipeline, document authenticity, field-level reliability, human correction burden or deployment fitness.

Open primary source ↗
First submitted 13 May 2023; revised 26 August 2024OCRBench — OCR in Large Multimodal Models

Brings 29 datasets together across text recognition, scene-text and document VQA, key-information extraction and handwritten mathematical expressions.

BoundaryA broad benchmark aggregate can still hide critical-character errors, reading-order failures, template drift, omitted pages, unsupported answers and human review cost.

Open primary source ↗

Continue checking

Move from document intelligence into multimodal, factuality, multilingual and benchmark evaluation, or the full library.

Enter the AI Knowledge LibraryEvaluate multimodal capabilityEvaluate factualityEvaluate multilingual capabilityRead benchmark claimsEnter AI Evidence Atlas