FUURAA AI Knowledge Library · Change decision protocol

AI agent material change and recommissioning

A bilingual protocol for deciding whether an update remains covered by an existing commissioning decision: freeze baseline and candidate, inventory model, data, prompt, tool, authority and use differences, determine whether prior evidence still transfers, then choose bounded verification, partial re-evaluation, full recommissioning or reversal.

Published7 August 2026Evidence statusFUURAA method synthesis grounded in primary risk, secure-development, operations and provenance sourcesScopeCommissioned AI agents undergoing technical, authority or use changes

Core principle

A version-number change is not necessarily material—and an acting system can become different without one.

Materiality is not determined by changed lines, release labels or provider wording. It depends on whether a change alters the use, people, authority, consequences, controls or evidence coverage in the commissioning decision.

Applicability boundaryThis is a public research method, not a FUURAA product-capability claim, safety certification, audit opinion, legal or compliance conclusion. Regulatory definitions and duties for substantial or material changes vary by jurisdiction, sector and use and require separate determination.

Seven change surfaces

Locate where the acting system changed before deciding whether evidence and authority still hold.

Change surfaceMinimum inventory
model or inference pathProvider version, weights, routing, decoding, safety layer or fallback model.
data, memory or retrievalTraining or reference data, embeddings, memory rules, retention, indexes or retrieval ranking.
instructions and orchestrationSystem prompts, policies, planners, agent graph, retries, delegation or termination logic.
tools and external dependenciesAPIs, schemas, packages, connectors, providers, response contracts or network paths.
identity and authorityCredentials, roles, spend, action limits, schedules, delegation or emergency revocation.
users, tasks and environmentCohorts, languages, geography, consequentiality, operating context or human reliance.
controls and evidenceMonitoring, thresholds, logging, recourse, rollback, retention or review ownership.

Eight change gates

Every gate must preserve a question, evidence and a stopping condition.

01

Freeze the authorised baseline and candidate

Core question
Which exact operating system is authorised now, and which exact system would replace it?
Evidence to preserve
Baseline and candidate IDs, manifests, timestamps, environments, authority envelopes and machine-readable diffs.
Stopping condition
Do not classify a change from a product name, release note or provider label alone.
02

Inventory every changed surface and hidden dependency

Core question
What changed directly, transitively or outside the deployer's control?
Evidence to preserve
Component diff, dependency graph, provider notices, tool schemas, data and prompt lineage, configuration and secret references.
Stopping condition
Treat an unobservable provider or dependency update as unresolved, not immaterial.
03

Map the change to use, people, authority and harm

Core question
Could the change alter who is affected, what the agent may do or the severity and reversibility of failure?
Evidence to preserve
Affected tasks and cohorts, action classes, data sensitivity, consequence pathways, human reliance and excluded uses.
Stopping condition
Any plausible expansion of authority, population, purpose or consequence defeats an immaterial classification.
04

Determine which prior evidence remains transferable

Core question
Which evaluation, canary, control and operating findings still cover the candidate—and why?
Evidence to preserve
Evidence-to-component map, unchanged assumptions, invalidated findings, transfer rationale, unknowns and expiry dates.
Stopping condition
Similarity is not transfer evidence; block reliance where the causal link to a finding has changed.
05

Select proportionate verification before seeing results

Core question
What regression, differential, adversarial, subgroup and denied-action tests does this change require?
Evidence to preserve
Predeclared tests, metrics, baselines, slices, success and stop thresholds, seeds, environments and retained failures.
Stopping condition
Do not reduce evidence because a candidate demo looks better or expand thresholds after failure.
06

Verify control continuity, rollback and coexistence

Core question
Will least authority, human intervention, logging, rollback and data integrity survive the transition?
Evidence to preserve
Negative-action tests, takeover rehearsal, rollback and restore evidence, migration checks, shadow or canary isolation and credential revocation.
Stopping condition
Do not proceed when rollback is nominal, states cannot be reconciled or old authority survives unintentionally.
07

Issue an independent change decision

Core question
Is the candidate covered, conditionally verified, partially re-evaluated, fully recommissioned or rejected?
Evidence to preserve
Decision class, evidence, dissent, conditions, owners, authorised scope, effective time, expiry and reopen triggers.
Stopping condition
The builder of the change must not silently convert its release note into operating authority.
08

Update the operating record and watch the transition

Core question
Can operators prove which version acts, under which decision, and detect transition-specific failure?
Evidence to preserve
Deployment event, active version proof, updated inventory, monitoring deltas, canary results, user communication, old-version disposition and follow-up review.
Stopping condition
Stop expansion if version attribution, evidence continuity or transition monitoring is incomplete.

Five change decisions

The decision sets both the evidence burden and whether the candidate may act now.

covered maintenance

covered maintenance

No authorised boundary or evidence-bearing behaviour changed; document the diff and routine verification.

bounded verification

bounded verification

A known low-impact surface changed; targeted regression and control checks support continued authority.

partial re-evaluation

partial re-evaluation

One or more prior findings no longer transfer; refresh the affected evidence before release.

full recommissioning

full recommissioning

Purpose, people, authority, consequentiality or multiple coupled surfaces changed materially.

reject, hold or revert

reject, hold or revert

Evidence, controls, rollback or independent ownership is insufficient for the candidate.

FUURAA analysisAI-agent change risk comes from coupling. An apparently local model update can alter tool selection; a tool-schema adjustment can change retries, spend and external action; an authority expansion can leave prior safety tests covering the wrong consequences. Materiality is therefore not a single technical property but a relational judgement across change, use, authority and evidence. Begin with the candidate unapproved; authority continues only when traceable differences, transferable evidence and an independent decision close together.

Minimum change decision record

Sixteen fields turn “just an update” into a reconstructable evidence decision.

  1. 01baseline decision and release ID

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  2. 02candidate release and change owner

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  3. 03machine-readable component diff

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  4. 04direct and transitive dependencies

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  5. 05provider-originated changes and unknowns

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  6. 06affected users, tasks and environments

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  7. 07authority and data-boundary delta

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  8. 08harm, reliance and reversibility analysis

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  9. 09prior evidence retained

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  10. 10prior evidence invalidated

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  11. 11predeclared verification plan

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  12. 12test results and retained failures

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  13. 13control, rollback and migration evidence

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  14. 14decision class, conditions and dissent

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  15. 15authoriser, effective time and expiry

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

  16. 16deployment proof and transition review

    Record an inspectable identifier, definition, delta, evidence path, owner, time or explicit unknown.

Primary sources and evidence boundaries

Use frameworks to design change control—not to impersonate change verification.

Sources rechecked 7 August 2026. Every source states its publication timing or living-document status, methodological role and non-transfer boundary.

Published 26 January 2023NIST · AI RMF 1.0

Provides voluntary lifecycle outcomes for governing, mapping, measuring and managing AI risk, including continuous review.

BoundaryIt does not define a universal material-change threshold or authorise a particular release.

Open primary source ↗
Living resource · checked 7 August 2026NIST AIRC · AI RMF Playbook: Map

Suggests reviewing third-party release schedules, hotfixes, updates and compatibility guarantees when mapping risk.

BoundarySuggested actions are informative and context dependent; the Playbook is being updated with the AI RMF.

Open primary source ↗
Living resource · checked 7 August 2026NIST AIRC · AI RMF Playbook: Manage

Connects post-deployment monitoring, change management, renewed TEVV, risk response and decommissioning.

BoundaryIt does not prescribe this protocol, sector duties or evidence sufficient for every agent.

Open primary source ↗
Final published 26 July 2024NIST · SP 800-218A

Adds AI-specific secure-development practices for model producers, system producers and acquirers.

BoundarySecure development guidance does not establish operating safety, change materiality or release approval.

Open primary source ↗
Published 27 November 2023UK NCSC · Secure operation and maintenance

Calls for behaviour and input monitoring, secure-by-design updates and lessons learned during operation.

BoundaryHigh-level security guidance does not set task-specific tests, legal duties or universal thresholds.

Open primary source ↗
W3C Recommendation 30 April 2013W3C · PROV-DM

Supplies a general model for relating entities, activities and agents across a provenance record.

BoundaryProvenance can show lineage; it does not prove truth, control effectiveness or release suitability.

Open primary source ↗