FUURAA AI Knowledge Library · Coding capability claim evaluation

How to evaluate AI coding capability claims

Use six evidence gates to turn “can code” into a defined task, repository and environment, reproducible tools and budgets, independent tests and review, complete failure denominators, real-workflow controls and a dated conclusion. Apply it to model reports, code assistants, coding agents, benchmarks, procurement, diligence and engineering review.

Published26 August 2026Evidence statusMethod synthesis grounded in primary measurement and code-evaluation researchScopeCode assistants, coding agents, benchmarks, procurement, diligence and engineering review

Writing code is not completing a software change

A coding claim must preserve the task, repository state, context, tools, validation, failures, human work and operating controls together.

Code that parses, passes examples or scores well on short-function benchmarks answers only a local question under a specific protocol. A real software change may also require cross-file understanding, requirement clarification, test changes, environment handling, data protection, review and safe rollback.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, code assistant, agent, repository or benchmark. It is not a capability guarantee, certification, audit, security conclusion, procurement, investment, legal or compliance opinion.

Six rejectable evidence gates

Each gate answers one decision question, requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the exact coding claim and user outcome

Decision question
Is the claim about completion, generation, explanation, review, debugging, repository issue resolution, migration or an autonomous software change—and for which language, codebase and consequence?
Minimum evidence
Verbatim claim and date; exact model, agent and toolchain version; task definition; language, framework, repository and user; acceptance, security and maintainability thresholds.
Stop condition
Stop when “can code” collapses materially different tasks into one label or is inferred from a polished demonstration.
02

Define task provenance, repository state and contamination

Decision question
Where did tasks, tests and fixes come from; what commit and environment do they target; could problems, solutions or near-duplicates have entered training or prompt examples?
Minimum evidence
Sampling frame, task and repository licences, commit hashes, issue and test provenance, time split, duplicate search, contamination analysis, hidden-test custody and evidence cut-off.
Stop condition
Stop when memorised public solutions, benchmark-specific patches or a curated easy subset are silently transferred to unseen repositories.
03

Reproduce context, tools, permissions and budgets

Decision question
Which files, documentation, retrieval, terminal, tests, network, package registries, credentials, retries, tokens, time and human hints were available?
Minimum evidence
Exact prompt and context selection; repository visibility; tool versions; sandbox and network policy; command log; token, time and cost budgets; retries, candidate selection and human assistance.
Stop condition
Stop when a tool-rich multi-attempt agent result is compared with an unaided one-shot model, or privileged access and human rescue are omitted.
04

Verify behaviour beyond visible tests

Decision question
Does the change satisfy the stated requirement, hidden and adversarial tests, static and security checks, interfaces, data migrations, performance budgets and neighbouring behaviour?
Minimum evidence
Independent oracle; new and held-out tests; diff review; build, lint, type and security results; regression and property tests; dependency and licence checks; runtime observations and reviewer adjudication.
Stop condition
Stop when success means only that generated code parses, visible tests pass, or the system grades tests and patches that it created itself.
05

Count complete outcomes, side effects and human work

Decision question
Across all assigned tasks, how often is the first attempt accepted, how much review and repair remains, and what failures, insecure changes or unintended effects occur?
Minimum evidence
Task-level denominator; pass@1 and explicitly bounded pass@k; compile and test failures; invalid or abandoned runs; review findings; human minutes; reverted changes; latency, token and cost distributions.
Stop condition
Stop when best-of-many output, selected demos or accepted snippets hide failed attempts, review burden, regressions, security findings or total cost.
06

Test transfer, operating controls and expiry

Decision question
Does evidence survive new repositories, languages, frameworks, dependency states, ambiguous requests and team workflows; what approval, rollback and monitoring control real changes?
Minimum evidence
Production-like shadow tasks, repository and language strata, maintainer review, least privilege, branch protection, isolated execution, rollback rehearsal, monitoring period, material-change triggers, expiry and owner.
Stop condition
Stop when a static benchmark becomes permanent authority to merge or deploy, or the model, agent, tools, repository, dependencies or policy change without revalidation.

Minimum coding-claim failure matrix

Check these conditions deliberately to distinguish local code generation from a reviewable, mergeable and operable software change.

  • 01
    Generated code passes examples but fails held-out edge cases

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

  • 02
    A public benchmark solution or near-duplicate appears in training data

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

  • 03
    A one-file score is presented as repository-level engineering ability

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

  • 04
    The patch fixes the visible test while breaking a neighbouring interface

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

  • 05
    A package or command introduces an insecure dependency or side effect

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

  • 06
    Best-of-many sampling hides failed attempts, latency and cost

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

  • 07
    Human hints, context selection or repair are omitted from the result

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

  • 08
    Benchmark success is converted into unattended merge or deploy authority

    Record affected tasks, repositories, environments, tests, users and decisions; preserve the narrowest conclusion that remains and specify whether to add tests, segment results, restrict permissions, repair manually, retest or withdraw the claim.

Minimum coding-capability evaluation record

Let the next reviewer reconstruct the conclusion under the same repository, environment, context, tools, budgets, tests and review boundaries.

  1. 01Exact claim, claimant, system, task, user, date and expiry
  2. 02Language, framework, repository, commit, environment and consequence
  3. 03Task, issue, test, reference-fix and contamination provenance
  4. 04Prompt, selected context, retrieval, documentation and tools
  5. 05Permissions, network, dependencies, tokens, time, cost and retries
  6. 06Acceptance oracle, held-out tests, build, lint and type results
  7. 07Diff review, security, licences, performance and side effects
  8. 08Complete task denominator, failures, abandoned runs and variance
  9. 09Human hints, review minutes, repairs, reverts and escalation
  10. 10Narrowest supported conclusion, operating controls and owner

Common evidence states

Bind conclusions to the exact system, task, repository, tools, budgets, validation, human work and date—not a generic claim that a system “can code”.

Supported

The exact system meets declared functional, review, security, resource and operating thresholds on a reproducible and representative task distribution.

Conditional

Evidence supports only named tasks, repositories, languages, tools, permissions, budgets and review controls; transfer remains bounded.

Mixed

Results diverge across task types, repositories, languages, tests, reviewers, security findings or resource levels and require segmented reporting.

Insufficient

A demonstration, visible-test pass, short-function benchmark, selected patch or undocumented agent run cannot establish coding capability.

FUURAA analysisThe minimum decision unit for an AI coding capability claim is exact system, agent and version × task, language, repository commit and environment × context, tools, permissions and budgets × independent tests, review, security and side effects × complete failure denominator, human work and cost × operating controls and cut-off date. Generated code is the start of a candidate change; only independent validation, complete outcomes and a controlled workflow can support the narrower conclusion that it is usable for a named class of software change.

Primary sources and non-transfer boundaries

These sources constrain statistical transfer, short-function correctness, code-task breadth, cross-file context, real repository issues and contamination control; none independently proves end-to-end software-engineering capability.

Sources rechecked 26 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published 17 February 2026NIST AI 800-3 — Expanding the AI Evaluation Toolbox with Statistical Models

Distinguishes fixed-benchmark accuracy from generalized accuracy and shows why repeated trials, item difficulty, assumptions and uncertainty matter.

BoundaryStatistical modelling improves interpretation; it does not make coding tasks representative of a repository, workflow or deployment by itself.

Open primary source ↗
First submitted 7 July 2021OpenAI — Evaluating Large Language Models Trained on Code / HumanEval

Introduces HumanEval and pass@k for functional correctness of Python programs synthesised from docstrings, while discussing repeated sampling and deployment impacts.

BoundaryShort self-contained functions with hidden unit tests do not establish repository understanding, maintainability, security or end-to-end software-engineering ability.

Open primary source ↗
First submitted 9 February 2021Microsoft Research and collaborators — CodeXGLUE

Collects ten code-understanding and generation tasks across fourteen datasets, making task type and metric choice visible rather than treating coding as one capability.

BoundaryBreadth across curated datasets is not evidence of current production performance, secure changes or successful collaboration inside a live repository.

Open primary source ↗
First submitted 5 June 2023UC San Diego — RepoBench

Separates repository-level retrieval, code completion and the combined pipeline, exposing the need for relevant cross-file context in Python and Java.

BoundaryNext-line completion with supplied repository context does not measure issue diagnosis, multi-step editing, test repair, review or operational safety.

Open primary source ↗
First submitted 10 October 2023Princeton University — SWE-bench

Frames real GitHub issues as repository-level software-engineering tasks requiring codebase understanding, coordinated edits and executable validation.

BoundaryA resolved benchmark instance remains conditional on repository snapshot, environment, tests, issue selection, scaffolding and the benchmark's definition of success.

Open primary source ↗
First submitted 12 March 2024LiveCodeBench collaboration — Holistic and Contamination-Free Evaluation

Uses newly released contest problems and covers generation, self-repair, execution and test-output prediction to reduce contamination and broaden code evaluation.

BoundaryFresh contest tasks reduce one contamination path but do not reproduce private codebases, ambiguous product requirements, team review, dependencies or production operations.

Open primary source ↗

Continue checking

Move from coding capability into reasoning, benchmarks, web agents, latency and reliability, and meaningful human oversight.

Evaluate reasoning capabilityRead benchmark claimsEvaluate web-using agentsEvaluate latency and reliabilityEvaluate human oversightEnter AI Evidence Atlas