FUURAA AI Knowledge Library · Image generation evaluation

How to evaluate AI image generation capability claims

Use six evidence gates to turn “high-quality text-to-image” into an exact use, prompt distribution, system version, generation settings, image–text alignment, composition, visual defects, human preference, complete sampling, selection and editing, safety, latency, cost and a dated deployment boundary. Apply it to research on text-to-image, marketing assets, illustration, product visuals and content production.

Published27 August 2026Evidence statusMethod synthesis grounded in primary image-generation evaluation researchScopeText-to-image, marketing assets, illustration, product visuals and content-production research

Beautiful is not the same as correct, controllable or usable

An image-generation claim must keep prompts, run configuration, all candidates, dimensional evaluation, human selection and real use in one evidence chain.

Showcase images and one score can hide prompt misunderstanding, relation errors, failed seeds, selection labour, safety problems and post-editing at the same time. A model can be more attractive but less faithful, or score higher automatically while being harder to use in production.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, artist, dataset or leaderboard. It is not a quality guarantee, originality decision, rights clearance, safety certification, procurement, audit, investment, legal or compliance opinion.

Six rejectable evidence gates

Each gate answers one decision question, requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the use case, prompt and acceptance decision

Decision question
Is the claim about text-to-image generation, variation, inpainting or an edited workflow; for which audience, channel and consequence?
Minimum evidence
Verbatim claim and date; intended task and output format; prompt language and policy; target audience, usage rights review, quality floor and acceptance rubric.
Stop condition
Stop when a showcase image, generic “photorealistic” label or model name replaces a defined task, prompt population and acceptance decision.
02

Preserve the exact generator and run configuration

Decision question
Which model, API, checkpoint and complete generation pipeline produced the images, under what settings and resource conditions?
Minimum evidence
Provider and model version; checkpoint and adapters; sampler, steps, guidance, seed and resolution; negative prompt; safety filters; reference images; upscalers, face restoration, editing tools and compute environment.
Stop condition
Stop when outputs from different versions, hidden reranking, undisclosed editing or incomparable resolutions are pooled into one capability claim.
03

Define prompt distribution, provenance and coverage

Decision question
How were prompts sampled, translated or rewritten; which languages, cultures, subjects, styles and difficulty levels are represented?
Minimum evidence
Prompt set and hashes; source and licence; sampling frame; deduplication; translation and rewriting steps; subject, language and complexity strata; reference-image provenance and possible benchmark overlap.
Stop condition
Stop when hand-picked English prompts or an undisclosed prompt optimiser are presented as evidence for broad multilingual, cultural or open-world generation.
04

Separate fidelity, composition, quality and preference

Decision question
Does evaluation distinguish prompt following, object count and relations, visual defects, readability, aesthetics and user preference instead of collapsing them?
Minimum evidence
Metric names, versions and references; per-skill results; object and relation checks; artefact rubric; typography tests; qualified human ratings, instructions, blinding, agreement and uncertainty; metric–human validation.
Stop condition
Stop when FID, CLIP-like similarity, one aesthetic score or one judge model is treated as proof of every image-quality dimension.
05

Count every sample, rejection, retry and edit

Decision question
How many candidates were generated, rejected, regenerated, reranked or edited before the displayed or accepted result?
Minimum evidence
All prompts, seeds and outputs; invalid and blocked runs; candidate count per prompt; selection and reranking rules; human selection time; edit history; acceptance, abstention and failure rates; latency and cost distributions.
Stop condition
Stop when only the best image survives, failed or unsafe outputs disappear, or success is reported without the total generation and human-work denominator.
06

Test harms, deployment transfer and expiry

Decision question
What happens across protected groups, sensitive subjects, misleading contexts, real channels and future model or policy changes?
Minimum evidence
Subgroup and red-team prompt sets; stereotype, sexual and violent-content tests; identity and consent controls; provenance or disclosure checks; incident and fallback procedures; target-channel review; change triggers, owner and expiry date.
Stop condition
Stop when a benchmark average or filter label is transferred into legal, cultural, safety or production suitability without direct review in the target context.

Minimum image-generation claim failure matrix

Check these conditions deliberately before transferring selected examples or one automated score into a broad image-generation claim.

  • 01
    Prompt optimiser or translation silently changes the requested subject, relation, culture or style

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

  • 02
    One object is correct but counts, attributes, spatial relations or negation fail

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

  • 03
    High alignment hides malformed anatomy, unreadable text, repeated detail or implausible geometry

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

  • 04
    A distribution or judge score improves while human acceptance for the target task declines

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

  • 05
    Best-of-many selection hides rejected, blocked, unsafe or low-quality generations

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

  • 06
    Results shift materially across seeds, aspect ratios, resolutions, languages or API revisions

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

  • 07
    People, cultures or sensitive subjects receive stereotyped, sexualised or misleading depictions

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

  • 08
    A benchmark result is transferred to advertising, editorial, education or high-stakes use without channel review

    Record the affected prompt, language, setting, metric, audience and channel; preserve the narrowest conclusion that remains and specify whether to resample, regenerate, rescore, include failures, review manually or withdraw the claim.

Minimum image-generation capability evaluation record

Let the next reviewer reconstruct the conclusion with the same prompts, version, settings, complete outputs, metrics, human rubric and operating boundary.

  1. 01Exact claim, date, use case, audience, channel, consequence and acceptance rubric
  2. 02Provider, model, API or checkpoint version and complete generation pipeline
  3. 03Sampler, steps, guidance, seed, resolution, filters, adapters and resource conditions
  4. 04Prompt set, language, source, licence, sampling, translation, rewriting and hashes
  5. 05Reference-image provenance, benchmark overlap and data-contamination controls
  6. 06Separate alignment, composition, artefact, typography, aesthetic and preference results
  7. 07Metric and judge versions, human-review protocol, agreement and uncertainty
  8. 08All outputs, invalid runs, blocks, retries, candidate counts, selections and edits
  9. 09Acceptance, abstention, subgroup, safety, latency, cost and human-effort distributions
  10. 10Deployment boundary, disclosure, monitoring, fallback, change triggers, owner and expiry

Common evidence states

Bind conclusions to the exact generator, prompt distribution, configuration, evaluation dimensions, sampling process, use and date.

Supported

Evidence supports the exact generator, prompt distribution, run configuration, evaluation dimensions, complete sampling process and dated deployment boundary.

Conditional

Evidence supports a narrower prompt class, language, resolution, generation mode, selection process, audience or channel.

Mixed

Results vary materially across alignment, composition, visual quality, people, languages, seeds, judges, safety or production acceptance.

Insufficient

System identity, representative prompts, run settings, separate metrics, human validation, complete outputs, safety tests, effort, cost or transfer evidence is missing.

FUURAA analysisThe minimum decision unit for an AI image-generation claim is exact generator and version × use, audience, prompt distribution and language × complete generation pipeline, settings and resources × alignment, composition, visual defects, typography and human preference × all outputs, failures, selection, editing, safety, latency and cost × deployment boundary and cut-off date. FID, CLIPScore, T2I-CompBench++, GenEval, HEIM and Gecko illuminate different parts; none independently represents real image-production capability.

Primary sources and non-transfer boundaries

These sources constrain distributional quality, image–text compatibility, compositionality, object-level alignment, holistic risks and human ratings; none independently proves general image-generation capability.

Sources rechecked 27 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

First submitted 26 June 2017; revised 12 January 2018Heusel et al. — Fréchet Inception Distance (FID)

Introduces FID to compare generated and reference-image feature distributions, creating an influential distribution-level signal for image generation.

BoundaryA distribution score does not verify whether an individual image follows its prompt, contains the right objects, is safe, is original or is useful for a target workflow.

Open primary source ↗
First submitted 18 April 2021; revised 23 March 2022CLIPScore — Reference-Free Image–Text Compatibility

Studies a reference-free signal based on image–text compatibility and shows where it complements reference-based evaluation.

BoundaryCorrelation on evaluated captioning corpora does not make CLIPScore a complete judge of prompt fidelity, counting, spatial relations, typography, aesthetics, harm or human preference.

Open primary source ↗
First submitted 12 July 2023; revised 8 March 2025T2I-CompBench++ — Compositional Text-to-Image Benchmark

Provides structured prompts for attribute binding, object relationships, generative numeracy and complex compositions, exposing failures hidden by broad similarity scores.

BoundaryPerformance on its prompt taxonomy and automated evaluators does not establish aesthetics, cultural fit, safety, provenance, editing quality or real production acceptance.

Open primary source ↗
First submitted 17 October 2023GenEval — Object-Focused Text-to-Image Alignment

Uses object detection and colour verification to test co-occurrence, position, count and colour at a more fine-grained level than holistic alignment metrics.

BoundaryDetector agreement on selected object tasks does not prove open-world correctness, subtle relations, readable text, physical plausibility, cultural accuracy or safe use.

Open primary source ↗
First submitted 7 November 2023Stanford CRFM — Holistic Evaluation of Text-to-Image Models (HEIM)

Frames twelve evaluation aspects across alignment, quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality and efficiency.

BoundaryA benchmark snapshot across selected models, prompts, metrics and raters is not a universal ranking or evidence for every language, policy, audience and deployment.

Open primary source ↗
First submitted 25 April 2024; revised 17 March 2025Google DeepMind — Gecko Text-to-Image Evaluation

Studies prompt skills, human-rating templates, annotator reliability and metric correlation using more than 100,000 annotations.

BoundaryImproved metric–human correlation on evaluated prompts and rating templates does not remove ambiguity, rater variation, domain shift or the need for task-specific human review.

Open primary source ↗

Continue checking

Move from image generation into multimodal, factuality, multilingual and benchmark evaluation, or the full library.

Enter the AI Knowledge LibraryEvaluate multimodal capabilityEvaluate factualityEvaluate multilingual capabilityRead benchmark claimsEnter AI Evidence Atlas