Deployment evaluation · 13

Task-definition protocol

Fix start state, completion, time, damage, assistance, reset and exclusion rules before testing.

Evidence statusEngineering evidence framework · task and site validation requiredLast reviewed: 14 August 2026
13

FUURAA thesis

A result becomes comparable only when the protocol makes success and failure reproducible.

Engineering map

Turn the title into engineering objects that can be observed, measured and reviewed.

Each layer states the boundary to establish and the evidence needed for the next decision.

01

Conditions

Version the robot, objects, site and people.

02

Scoring

Define completion, quality, failure and intervention.

03

Repetition

Randomise relevant variation and preserve every run.

Verification questions

Write the questions first, then decide whether a demo, test, pilot or operating record can answer them.

Each question needs an object, conditions, denominator, threshold and accountable decision owner.

  1. 01

    Can another team run the same protocol?

  2. 02

    Are resets and retries counted?

  3. 03

    Does scoring reward unsafe speed?

  4. 04

    Are exclusions decided before results?

Evidence to preserve

Enable the next reader to reconstruct conditions, results, failures and the decision.

A conclusion alone loses reviewability; raw records, configuration and exclusions matter too.

  1. 01

    Evidence package 1

    Versioned test protocol

  2. 02

    Evidence package 2

    Complete trial-level dataset

  3. 03

    Evidence package 3

    Predefined scoring and exclusions

Scope boundary

State what this evidence still cannot be generalised to.

Sources and evidence status

Read standards scope, measurement evidence and application conclusions separately.

Source dates and review status remain visible; external sources open in a new tab.