FUURAA AI Knowledge Library · Agent evaluation protocol

How to evaluate an AI agent before deployment

A bilingual protocol for systems that act: go beyond final answers by freezing models, tools, authority, environment and human controls, then measure outcomes, process, variance, recovery, safety, transfer and handoff.

Published4 August 2026Evidence statusFUURAA method synthesis grounded in primary frameworks and measurement researchScopeAI agents with tools, memory or external actions

Object of evaluation

Evaluate the acting system—not only its underlying model.

The same model becomes a different agent system when tools, permissions, memory, prompts, environment or human controls change. Public benchmarks provide signals, but cannot replace system-level tests matched to a deployment decision.

Applicability boundaryThis is a public evaluation-design method—not safety certification, a compliance conclusion, legal advice, procurement advice or a FUURAA product-capability claim. High-impact, financial, medical, public-service or physical actions still require use-specific professional accountability, local requirements and independent validation.

Freeze the eight-part system identity first

If the version is unclear, the score is not comparable.

  • 01
    agent release
    Exact agent package, orchestration loop and configuration hash
  • 02
    model stack
    Provider, model, snapshot, reasoning settings and fallback order
  • 03
    tools
    Schemas, versions, side effects, credentials and approval rules
  • 04
    environment
    Sandbox, network, files, services, fixtures and reset procedure
  • 05
    memory
    Read/write scope, retention, retrieval policy and initial state
  • 06
    authority
    Permitted actions, budgets, exclusions, expiry and revocation
  • 07
    human controls
    Approval points, intervention channel, escalation owner and stop path
  • 08
    evaluation build
    Dataset, scorer, evaluator, seeds, repetitions and run timestamp

Seven evaluation gates

Every gate preserves a question, evidence and a blocking condition.

01

Define the decision and task distribution

Core question
What deployment decision will this evaluation inform, and which real tasks does the sample represent?
Evidence to preserve
Intended users, workflows, task taxonomy, prevalence, consequence and a non-agent baseline.
Blocking condition
Stop if a convenient benchmark is being used as a substitute for the actual decision context.
02

Freeze the system and affordances

Core question
Could another team reconstruct exactly what the agent could see, remember, call and change?
Evidence to preserve
The eight-part system identity, dependency locks, fixtures, permissions and environment reset evidence.
Blocking condition
Do not compare runs when model, tools, permissions or hidden environment state changed silently.
03

Score outcome and process separately

Core question
Did the task finish correctly, and did the path respect authority, safety, cost and evidence requirements?
Evidence to preserve
Final state, partial credit, tool trace, policy decisions, unsupported claims, cost, latency and interventions.
Blocking condition
A plausible final answer cannot erase prohibited actions, hidden retries or an unreconstructable trace.
04

Measure variance, recovery and long-horizon work

Core question
How often does the agent succeed across repetitions, longer tasks, partial failure and changed initial conditions?
Evidence to preserve
Per-run results, confidence intervals, failure clusters, checkpoints, recovery attempts and human-time baseline.
Blocking condition
Do not turn one successful trajectory—or a 50% task horizon—into a reliability or universal autonomy claim.
05

Challenge tools, instructions and authority

Core question
What happens under prompt injection, deceptive content, unavailable tools, stale credentials and conflicting instructions?
Evidence to preserve
Adversarial cases, containment evidence, denied actions, escalation records and residual-risk owners.
Blocking condition
Block release when the agent can cross a consequential boundary without an independent control or visible record.
06

Test transfer and human handoff

Core question
Does performance survive new users, languages, data, environments and real escalation pressure?
Evidence to preserve
Held-out and field results, subgroup slices, abstentions, operator workload, handoff quality and affected-party feedback.
Blocking condition
Do not transfer a laboratory, English-only or expert-operated result to a broader operating claim.
07

Make a scoped release decision

Core question
What exact use is ready, conditional, blocked or invalidated—and what evidence reopens the decision?
Evidence to preserve
Decision owner, threshold table, accepted uncertainty, monitoring, rollback, expiry and re-evaluation triggers.
Blocking condition
No aggregate score may override a failed critical gate or justify a use outside the tested boundary.

Minimum test matrix

An average success rate cannot replace layered testing.

Preserve per-run traces, failure classes and run configuration for every layer. Judge critical gates separately for the use case; do not average them away.

LayerWhat to testMinimum passing evidence
NominalRepresentative tasks under documented conditionsTask success, quality, cost, latency and complete trace
BoundaryRare, ambiguous and high-consequence casesSafe abstention, clarification and correct escalation
FaultUnavailable tools, timeouts, partial writes and stale stateNo duplicate consequence; recover or stop visibly
AdversarialInjection, malicious artefacts and authority confusionContainment, denial, evidence preservation and escalation
TransferNew users, languages, data and operating environmentsReport degradation and retain the original claim boundary
RegressionFrozen critical set on every material system changeNo silent loss against the approved release baseline

Release language

Release state must bind to a use, version and date.

supported

Ready for scoped trial

All critical gates meet named thresholds inside one bounded, monitored use.

developing

Conditional

Useful evidence exists, but compensating controls or missing coverage must be named.

mixed

Blocked

A critical safety, authority, reliability or handoff gate fails despite other strengths.

insufficient

Invalidated

The evaluated system, environment or decision boundary changed materially; prior results no longer govern release.

FUURAA analysisThe most common agent-evaluation error is translating ‘a model scored well on a benchmark’ directly into ‘this agent can reliably do the work.’ Model capability is one system input. Tool side effects, authority boundaries, environment state, long-horizon error accumulation, human takeover and runtime monitoring jointly determine deployment evidence.

Minimum evaluation run record

Twelve fields make each result reconstructable, comparable and expirable.

  1. 01eval_id + decision

    Record an inspectable identifier, value, path or owner.

  2. 02system identity

    Record an inspectable identifier, value, path or owner.

  3. 03task distribution

    Record an inspectable identifier, value, path or owner.

  4. 04baseline

    Record an inspectable identifier, value, path or owner.

  5. 05dataset + fixtures

    Record an inspectable identifier, value, path or owner.

  6. 06scorers + thresholds

    Record an inspectable identifier, value, path or owner.

  7. 07repetitions + seeds

    Record an inspectable identifier, value, path or owner.

  8. 08per-run traces

    Record an inspectable identifier, value, path or owner.

  9. 09failures + interventions

    Record an inspectable identifier, value, path or owner.

  10. 10boundary + exclusions

    Record an inspectable identifier, value, path or owner.

  11. 11release state + owner

    Record an inspectable identifier, value, path or owner.

  12. 12expiry + rerun trigger

    Record an inspectable identifier, value, path or owner.

Primary sources and boundaries

Use frameworks and research to design the method—not to impersonate deployment approval.

Sources checked 4 August 2026. Each source retains its publication date, role and non-transfer boundary.

Published 26 January 2023NIST AI RMF 1.0

Connects context, measurement, governance and continuing risk management across the AI lifecycle.

BoundaryVoluntary and use-case agnostic; it does not prescribe these gates or certify an agent.

Open primary source ↗
Living resource · checked 4 August 2026NIST AI Metrology Center

Organises metrics, methods and tools for testing, evaluation, validation and verification.

BoundaryNIST states that inclusion is not endorsement, validation or a suitability decision.

Open primary source ↗
Published 26 July 2024 · updated 8 April 2026NIST AI 600-1 · Generative AI Profile

Adds generative-AI risks around measurement, provenance, human oversight and information integrity.

BoundaryCross-sector guidance still requires prioritisation for the actual agent and decision context.

Open primary source ↗
Open-source framework · released May 2024UK AI Security Institute · Inspect

Provides composable tasks, datasets, agents, tools, scorers, sandboxes, logs and analysis.

BoundaryA framework makes runs reproducible; it does not make a weak task set representative or a scorer correct.

Open primary source ↗
Last updated 8 May 2026METR · Task-Completion Time Horizons 1.1

Relates agent success probability to the human-expert duration of coherent software tasks.

BoundaryMETR warns that the metric is not wall-clock autonomy, all-task coverage or evidence that entire jobs can be automated.

Open primary source ↗
v1.0 released 4 December 2024MLCommons · AILuminate

Separates demo, practice and official safety tests and documents result-publication conditions.

BoundaryIts published benchmark focuses on defined language-model hazards; it is not a complete agent deployment evaluation.

Open primary source ↗