Deployment evaluation · 13

Measure useful task outcome

Use service-specific quality and completion measures with every attempt in the denominator.

Evidence statusService decision framework · site and user validation requiredLast reviewed: 14 August 2026
13

FUURAA thesis

A fair metric counts what users receive, what staff repair and what the system abandons.

Engineering map

Turn the title into engineering objects that can be observed, measured and reviewed.

Each layer states the boundary to establish and the evidence needed for the next decision.

01

Quality

Set observable service acceptance criteria.

02

Completeness

Count success, partial success, retry, intervention and damage.

03

Human impact

Measure waiting, workload, dignity and accessibility.

Verification questions

Write the questions first, then decide whether a demo, test, pilot or operating record can answer them.

Each question needs an object, conditions, denominator, threshold and accountable decision owner.

  1. 01

    Does the metric reflect recipient need?

  2. 02

    Are all attempts included?

  3. 03

    Is manual finishing counted?

  4. 04

    What threshold supports expansion?

Evidence to preserve

Enable the next reader to reconstruct conditions, results, failures and the decision.

A conclusion alone loses reviewability; raw records, configuration and exclusions matter too.

  1. 01

    Evidence package 1

    Outcome specification

  2. 02

    Evidence package 2

    Attempt-level dataset

  3. 03

    Evidence package 3

    Decision threshold record

Scope boundary

State what this evidence still cannot be generalised to.

Sources and evidence status

Read standards scope, measurement evidence and application conclusions separately.

Source dates and review status remain visible; external sources open in a new tab.