FUURAA AI Knowledge Library · Commissioning protocol

AI agent commissioning and authority handover

A bilingual protocol connecting pre-deployment evaluation to real operation: freeze the release object, reconcile evidence, bind identity to least authority, rehearse human control and rollback, use a bounded canary to reach an independent expiring decision, then hand responsibility and monitoring evidence to the operating team.

Published7 August 2026Evidence statusFUURAA method synthesis grounded in primary risk, secure-development, deployment and provenance sourcesScopeCommissioning decisions for AI agents with tools, external actions or persistent operation

Core principle

Evaluation shows how a system performed; commissioning decides what it may do now, who owns it and until when.

Commissioning is not the act of handing an evaluation report to operations. It connects exact system identity, evidence, authority, stopping conditions and ownership into a decision that can be enforced, reviewed and expired.

Applicability boundaryThis is a public research method, not a FUURAA product-capability claim, safety certification, audit opinion, legal or compliance conclusion. It does not replace sector-specific human-safety, medical, financial, employment, privacy, cybersecurity, consumer-protection or regulatory approval.

Four separated responsibilities

Make delivery, acceptance and authorisation visibly different decisions.

  • 01
    release owner
    Own the declared use, release scope, acceptance boundary, exceptions and expiry.
  • 02
    system custodian
    Freeze the system identity, bind authority and demonstrate controls, rollback and denied actions.
  • 03
    operating owner
    Accept day-to-day responsibility, monitoring thresholds, escalation duties and user communication.
  • 04
    independent authoriser
    Challenge the evidence, exclusions, unresolved risk and conditions before granting a bounded decision.

Eight commissioning gates

Every gate must preserve a question, evidence and a stopping condition.

01

Freeze the release object and intended use

Core question
Which exact agent, model, prompts, tools, memory, data, environment, users and tasks are being commissioned?
Evidence to preserve
Stable release ID, component and dependency manifest, task distribution, user cohort, languages, geography, excluded uses and change diff.
Stopping condition
Do not authorise a product label when the acting system and use boundary cannot be reconstructed.
02

Reconcile evaluation evidence and unresolved findings

Core question
Does the evidence package cover this release, task mix, authority and operating context—and what remains unknown?
Evidence to preserve
Evaluation record and verifier result, failed runs, subgroup slices, transfer limits, red-team findings, accepted deviations and evidence review date.
Stopping condition
Block commissioning when evidence belongs to another version, hides failures or excludes a consequential task or cohort.
03

Bind identity to the least authority needed

Core question
Which actions, data, tools, networks, spend, schedules and delegations may this release use—and which must be denied?
Evidence to preserve
Credential and tool inventory, scoped roles, approved action classes, rate and value limits, network paths, denied-action tests, delegation rules and emergency revocation.
Stopping condition
Do not rely on policy text when technical controls permit broader action or credentials are shared, recoverable or unowned.
04

Verify environment, provenance and supply-chain closure

Core question
Can every consequential model, dataset, prompt, package, API, tool and configuration be traced to an approved source and version?
Evidence to preserve
Checksums, provenance links, licences and notices, software bill of materials, provider contracts, runtime configuration, secret references and known external dependencies.
Stopping condition
Stop for unexplained assets, mutable unpinned dependencies, missing rights review or a provider change that invalidates the evaluation.
05

Rehearse human control, fallback and rollback

Core question
Can people detect, interrupt, assume control, correct and recover the service under realistic time pressure?
Evidence to preserve
Named responders, escalation paths, user recourse, takeover rehearsal, rollback target, restore test, safe fallback, communications template and maximum response time.
Stopping condition
Do not commission when the stop control is inaccessible, rollback is untested or affected people have no usable route to challenge or obtain help.
06

Run a bounded canary with predeclared exit criteria

Core question
What smallest live slice can test the operating assumptions without silently becoming general release?
Evidence to preserve
Canary cohort, duration, task and authority limits, baseline, success and stop thresholds, action logs, human interventions, comparison group and retained failures.
Stopping condition
Stop expansion when thresholds were chosen after results, cohorts changed without review, monitoring has gaps or failures cannot be attributed to the release.
07

Issue an independent, scoped commissioning decision

Core question
What exactly is authorised, restricted, deferred or rejected, by whom, on which evidence and until when?
Evidence to preserve
Decision state, authorised release and cohort, authority envelope, exclusions, conditions, dissent, accountable owners, effective time, expiry and reopen triggers.
Stopping condition
A meeting note, successful demo or unsigned checklist is not a commissioning decision.
08

Transfer responsibility into an operating evidence cycle

Core question
Can the operating team prove what it accepted, what to monitor, when to restrict action and what invalidates the decision?
Evidence to preserve
Signed handover, monitoring owner, signal definitions, thresholds, change classes, incident route, review cadence, record location, access test and replacement decision link.
Stopping condition
Do not begin normal operation when ownership, evidence access, monitoring or escalation disappears at the handover boundary.

Commissioning states

States describe the decision that can be relied on—not unfinished work made to look ready.

not ready

not ready

One or more release, evidence, authority or recovery gates remain blocked.

canary only

canary only

A narrow cohort and authority envelope may operate while predeclared evidence is gathered.

commissioned with conditions

commissioned with conditions

The exact release and use are authorised with named limits, owners, monitoring and expiry.

expired or invalidated

expired or invalidated

Time, change, drift, incident or missing review means the decision can no longer support operation.

FUURAA analysisThe most dangerous commissioning gap is often not the absence of evaluation, but the failure to bind the evaluated object, deployed release and real authority in one decision. Models, tools, prompts, data, memory, runtime and provider dependencies jointly form the acting system; a material post-evaluation change can leave the evidence covering a different system. Commissioning should therefore be an enforceable authority contract and an expiring evidence decision—not permanent permission created by a meeting.

Minimum commissioning and handover record

Fifteen fields turn release approval into a reconstructable, revocable record of responsibility.

  1. 01release and component identity

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  2. 02intended use and excluded uses

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  3. 03users, tasks, languages and geography

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  4. 04evaluation package and unresolved findings

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  5. 05authority envelope and denied actions

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  6. 06provenance and dependency manifest

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  7. 07rights, security and privacy escalation flags

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  8. 08human control and user recourse

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  9. 09rollback, restore and fallback evidence

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  10. 10canary design and results

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  11. 11monitoring signals and thresholds

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  12. 12incident and emergency-revocation route

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  13. 13decision, conditions and dissent

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  14. 14release and operating owners

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

  15. 15effective time, expiry and reopen triggers

    Record an inspectable identifier, definition, value, evidence path, owner, time or explicit unknown.

Primary sources and evidence boundaries

Use standards and guidance to design commissioning controls—not to impersonate operating proof.

Sources rechecked 17 August 2026. Every source states its publication timing, methodological role and non-transfer boundary.

Published 26 January 2023NIST · AI RMF 1.0

Provides voluntary lifecycle-wide outcomes for governing, mapping, measuring and managing AI risk.

BoundaryUse-case agnostic; it does not authorise a release, define agent permissions or certify a deployment.

Open primary source ↗
Final published 26 July 2024NIST · SP 800-218A

Adds AI-specific secure-development practices to the SSDF for model producers, system producers and acquirers.

BoundaryCentred on secure development; it is not an agent commissioning procedure or proof that controls work in operation.

Open primary source ↗
Published 27 November 2023UK NCSC · Secure deployment

Covers infrastructure protection, continuous model protection, incident preparation, responsible release and usable guidance.

BoundaryHigh-level secure-development guidance; it does not evaluate a particular agent, environment or release decision.

Open primary source ↗
Published 27 November 2023UK NCSC · Secure operation and maintenance

Connects behaviour and input monitoring, secure updates and lessons learned to the deployed lifecycle.

BoundaryIt does not set task-specific thresholds, sector-specific legal duties or universal evidence-retention periods.

Open primary source ↗
Draft EN dated September 2025 · publication announced 13 November 2025ETSI · EN 304 223 V2.0.0

Provides baseline cybersecurity requirements across secure design, development, deployment, maintenance and end of life.

BoundaryA baseline security standard, not a complete safety, privacy, sectoral compliance or commissioning record.

Open primary source ↗
W3C Recommendation 30 April 2013W3C · PROV-DM

Supplies a general model for relating entities, activities and agents in a provenance record.

BoundaryA provenance data model does not establish truth, control effectiveness, rights clearance or release suitability.

Open primary source ↗