FUURAA Robotics Observatory · Embodiment and autonomy evaluation

How to evaluate embodied AI and robot autonomy claims

Use six evidence gates to turn “the robot can work autonomously” into an exact embodiment, task and site, hardware and calibration, human contribution, complete physical trials, safety, recovery, simulation transfer, operating controls and a dated conclusion. Apply it to robot demonstrations, papers, product documentation, pre-procurement research, pilots and engineering review.

Published26 August 2026Evidence statusMethod synthesis grounded in primary robot measurement and embodied-learning researchScopeRobot demonstrations, papers, product documentation, pre-procurement research, pilots and engineering review

Motion is not sufficient evidence of autonomy

A robot claim must keep the model, hardware, environment, human contribution, physical outcome and operating time in one evidence chain.

A compelling video only proves that an action was recorded. A robot may depend on a choreographed environment, teleoperation, prior maps, special fixtures or undisclosed resets; a model score cannot replace physical-system evidence for safety, reliability, recovery and maintenance.

Applicability boundaryThis is a public research method, not a FUURAA, FUUVO or AFUU product-capability claim or an assessment of any robot, model, provider, country, laboratory, dataset or leaderboard. It is not a safety certification, procurement, audit, investment, legal or compliance opinion.

Six rejectable evidence gates

Each gate answers one decision question, requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the embodiment, autonomy claim and work outcome

Decision question
Which exact robot is claimed to perceive, plan or act autonomously; on what task, for whom, under what authority and with what physical consequence?
Minimum evidence
Verbatim claim and date; robot serial or configuration; morphology, payload and end effector; software, model, policy and firmware versions; task, site, user and acceptance thresholds.
Stop condition
Stop when a video, model name or generic label such as “general-purpose robot” substitutes for a bounded task and complete system identity.
02

Reproduce hardware, calibration, authority and human contribution

Decision question
What sensors, actuators, compute, tools, maps, fixtures and calibrations were active; what could the robot do; and when did a human plan, teleoperate, approve, reset or rescue it?
Minimum evidence
Bill of materials and configuration hashes; calibration and maintenance records; power, network and tool dependencies; autonomy mode; permission envelope; intervention, reset and teleoperation logs.
Stop condition
Stop when hidden teleoperation, pre-positioning, human resets or privileged infrastructure are excluded from the capability claim or denominator.
03

Define environment, task distribution and data provenance

Decision question
Which objects, people, lighting, surfaces, clutter, dynamics, failure states and task sequences are represented; where did demonstrations and test scenarios come from; could they overlap training?
Minimum evidence
Sampling frame; initial-state randomisation; environment and object strata; training and evaluation lineage; perceptual and semantic duplicate checks; operator identities; held-out sites and time splits.
Stop condition
Stop when a choreographed scene, memorised layout or selected successful trajectory is presented as broad autonomy or real-world generalisation.
04

Separate perception, planning, control and physical execution

Decision question
Where did each trial fail: sensing, state estimation, language grounding, planning, control, grasping, locomotion, contact, tool use, verification or stopping?
Minimum evidence
Time-synchronised sensor, state, plan, action and force traces; subsystem ablations; counterfactual instructions; operator and safety-controller events; end-state inspection independent of the policy.
Stop condition
Stop when a task-level success score hides wrong perception, unsafe motion, accidental completion, fixture dependence or a human-supplied recovery.
05

Count every trial, intervention, near miss and recovery

Decision question
Across all assigned starts, how often did the robot finish to specification without damage or intervention; how long did it take; what resources, retries and maintenance were required?
Minimum evidence
Predeclared trial count; continuous runs; success criteria; aborts, collisions, drops, unsafe forces and near misses; human interventions; recovery time; latency, energy, consumables, wear and confidence intervals.
Stop condition
Stop when aborted attempts, resets, safety stops, damaged objects, maintenance or slow failures disappear from the denominator or operating cost.
06

Validate simulation transfer, site transfer and operating controls

Decision question
Does performance survive equivalent physical tests, new sites, hardware variation, people, disturbances and time; what authority, monitoring, fallback, exclusion zones and expiry govern operation?
Minimum evidence
Paired simulation and physical results; held-out hardware and sites; canary deployment; shift-level and fleet logs; safety case; geofencing; fallback and manual recovery tests; change triggers and dated review.
Stop condition
Stop when simulation, one laboratory or one robot is treated as proof of production fleets, unattended operation, public-space safety or durable economics.

Minimum robot-claim failure matrix

Check these conditions deliberately to distinguish motion demonstrations and local capability from sustained autonomous work.

  • 01
    A polished demonstration uses a fixed camera, marked floor, pre-positioned objects and selected successful takes.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

  • 02
    A human silently plans, teleoperates, resets or rescues the robot while the result is labelled autonomous.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

  • 03
    Simulation success is transferred to hardware without paired physics, sensing, latency and contact validation.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

  • 04
    One narrow manipulation or navigation skill is described as general-purpose embodied intelligence.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

  • 05
    Aborts, failed starts, safety stops and operator resets are excluded from the reported trial count.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

  • 06
    Task completion ignores collisions, excessive force, drops, damage, near misses or inaccessible emergency stops.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

  • 07
    Cross-robot training gains are assumed to apply to an untested embodiment, gripper, sensor suite or control rate.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

  • 08
    Short laboratory trials are extrapolated to new sites, shifts, fleets, maintenance cycles or public interaction.

    Record the affected robot, hardware, task, site, people and consequence; preserve the narrowest conclusion that remains and specify whether to recalibrate, add physical trials, include failures, narrow authority, validate recovery, retest or withdraw the claim.

Minimum embodiment and autonomy evaluation record

Let the next reviewer reconstruct the conclusion under the same embodiment, site, initial states, authority, trial denominator and physical end states.

  1. 01Verbatim claim, decision, owner, publication date and evidence cut-off
  2. 02Exact robot, embodiment, payload, tools, sensors, actuators, compute, software, model, policy and firmware
  3. 03Task, site, user, work system, authority, consequence and acceptance thresholds
  4. 04Training, demonstration, simulation, environment, object and evaluation provenance
  5. 05Initial-state, reset, calibration, maintenance and randomisation protocol
  6. 06Human planning, teleoperation, approval, safety control, intervention and rescue
  7. 07Predeclared trials, continuous-run protocol and time-synchronised traces
  8. 08Complete outcomes: success, abort, collision, drop, damage, near miss, retry and invalid run
  9. 09Latency, energy, wear, consumables, maintenance, recovery and uncertainty
  10. 10Simulation-to-real evidence, transfer boundary, operating controls, change triggers and expiry

Common evidence states

Bind conclusions to the exact robot, task, site, physical trials, safety envelope and date—not a generic claim of “autonomy”.

Supported

Repeated physical evidence supports the exact robot, task, site and operating boundary with complete trials and controls.

Conditional

Evidence supports a narrower embodiment, environment, operator, autonomy mode, duration or safety envelope.

Mixed

Performance varies materially across tasks, starts, sites, hardware, people or failure types.

Insufficient

System identity, physical trials, denominators, human contribution, safety outcomes, recovery or transfer evidence is missing.

FUURAA analysisThe minimum decision unit for an embodied AI and robot-autonomy claim is exact robot, hardware and software version × task, work system, site, population and consequence × initial state, environment, object and people distribution × perception, planning, control, execution and human contribution × complete trials, intervention, recovery, safety and resource cost × operating boundary and cut-off date. RoboMIND brings a Chinese-led large-scale multi-embodiment manipulation and failure collection into international research, but dataset size, a digital twin or model success rate cannot replace direct validation of the target robot at the real site.

Primary sources and non-transfer boundaries

These sources constrain repeatable physical tests, long-horizon tasks, vision-language-action transfer, data diversity, cross-embodiment transfer and Chinese multi-embodiment research; none independently proves general robot autonomy.

Sources rechecked 26 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Created 18 February 2010; updated 27 May 2026NIST — Standard Test Methods for Response Robots

Defines repeatable apparatus, procedures and quantitative metrics for robot mobility, endurance, sensing, dexterity and operational tasks, often using continuous repetitions.

BoundaryA standard elemental test supports comparable capability measurement; it does not certify an entire autonomous work system or every deployment environment.

Open primary source ↗
First submitted 6 December 2021CALVIN — Language-Conditioned Long-Horizon Robot Manipulation Benchmark

Tests composition of language-conditioned manipulation tasks over longer horizons, including transfer to novel instructions, environments and objects in simulation.

BoundaryA simulated benchmark does not establish physical safety, actuator reliability, sensor calibration, real-world recovery or sustained field operation.

Open primary source ↗
First submitted 28 July 2023Google DeepMind — RT-2: Vision-Language-Action Models

Reports 6,000 evaluation trials for a vision-language-action approach that transfers web-scale visual and language knowledge into tested robotic control tasks.

BoundaryImproved generalisation on selected tabletop tasks does not establish universal embodiment, fleet reliability, safe contact or unbounded physical reasoning.

Open primary source ↗
First submitted 24 August 2023UC Berkeley and collaborators — BridgeData V2

Provides 60,096 manipulation trajectories across twenty-four environments on a documented low-cost robot and evaluates multiple imitation and offline-RL methods.

BoundaryMore diverse demonstrations can improve tested transfer; they do not remove embodiment mismatch, dataset bias, hardware wear, unsafe states or site-specific failure.

Open primary source ↗
First submitted 13 October 2023Open X-Embodiment Collaboration — Robotic Learning Datasets and RT-X Models

Standardises data from twenty-two robots and reports cross-robot transfer across a large multi-institution manipulation collection.

BoundaryPositive average transfer does not mean every robot, task, sensor, gripper or control rate benefits; each target embodiment still requires direct validation.

Open primary source ↗
First submitted 18 December 2024RoboMIND Collaboration — Multi-Embodiment Robot Manipulation Data

Brings a Chinese-led multi-embodiment contribution into international evaluation through 107,000 demonstrations, 479 tasks, four embodiments and 5,000 labelled failure demonstrations.

BoundaryA large unified collection and digital twin do not independently establish unbiased coverage, production safety, long-duration reliability or transfer beyond tested platforms.

Open primary source ↗

Continue checking

Move from robot autonomy into multimodal capability, reasoning, web agents, human oversight and the full Robotics Observatory.

Enter the Robotics ObservatoryEvaluate multimodal capabilityEvaluate reasoning capabilityEvaluate web agentsEvaluate human oversightEnter AI Evidence Atlas