FUURAA AI Knowledge Library · Browser-local verification tool

AI Agent Evaluation Run Record Verifier

Check a FUURAA Agent evaluation run-record JSON locally in your browser for system identity, tasks, design, per-run evidence, transfer boundaries, release state and expiry. The file is not uploaded to FUURAA.

Tool statusPublicly usable · browser-local processingVerification rulefuuraa.ai-agent-evaluation-run-record/v1Sources checked4 August 2026

Verification workspace

Load a run record, then close ten structural checks.

Choose a local JSON file, paste content, or load a blank template or clearly marked complete fictional example. Verification reads only the text in this page.

Checks cover Schema identity, record and decision identity, the complete acting system, task context, evaluation design, process evidence, applicability boundary, Release decision and controls, Decision expiry and Critical-failure consistency.

Privacy boundary: this tool does not submit, store or transmit the selected file. Continue to handle sensitive or restricted material under your organisation's rules.

Applicability boundary

This tool checks whether a record is reviewable—not whether the run occurred.

It can check

  • JSON, v1 rule identity and ten required structural groups
  • Release state, decision expiry and an obvious critical-failure conflict
  • Whether system, task, evaluation, result and transfer-boundary fields are missing

It cannot prove

  • That a run occurred, traces are untampered or owner identities valid
  • That tasks are representative, scorers valid, results reproducible or conclusions transferable
  • That a system is safe, deployment-ready or legally, audit or certification compliant

FUURAA analysisStructural verification should first reject records missing complete system identity, predeclared thresholds, per-run traces, critical failures, excluded uses, stop plans or expiry. A real release judgment still requires independent reviewers to open raw logs, inspect tasks and scorers, repeat consequential measurements, challenge failure classification and invalidate the decision when the system, authority, environment or use changes.

Standards and method sources

Use primary specifications to design inspectable records while preserving each source's boundary.

Internet-Draft · 16 June 2022JSON Schema 2020-12 · Validation Vocabulary

Provides concepts for required properties, types, enumerations and format-oriented structural validation.

Boundary: The document is an expired Informational Internet-Draft. This tool applies a focused FUURAA v1 rule set, not a general JSON Schema implementation.

Internet Standard · December 2017IETF RFC 8259 · JSON Data Interchange Format

Defines the JSON grammar and interoperable representation parsed by this browser-local verifier.

Boundary: Valid syntax cannot establish that a run occurred, a trace is authentic or a decision is sound.

W3C Recommendation · 30 April 2013W3C PROV-DM · The PROV Data Model

Frames provenance through entities, activities, agents, derivations and responsibility.

Boundary: The FUURAA run record borrows provenance concepts but is not a normative PROV serialisation.

Released May 2024 · documentation checked 4 August 2026UK AI Security Institute · Inspect

Shows how tasks, datasets, agents, tools, scorers, sandboxes and logs can make evaluation runs inspectable.

Boundary: Inspectable runs do not make a weak task distribution representative or an invalid scorer trustworthy.

Published 26 January 2023NIST AI RMF 1.0

Connects context, measurement, governance and continuing risk management across the AI lifecycle.

Boundary: Voluntary and use-case agnostic; it does not prescribe this verifier or certify a deployment.