Robotics technology stack · Layer 02

Embodied foundation models

A practical map of vision-language-action models, cross-embodiment data, robot policies and the gap between model demos and dependable work.

Evidence statusActive research; selective real-world demonstrationsLast reviewed: 11 August 2026
02

FUURAA thesis

Foundation models can broaden task priors and interfaces, but the deployed robot still needs bounded actions, real-time control, safety supervision and evidence on its own embodiment.

Reading method

Turn a technology label into an engineering chain, testable claims and explicit boundaries.

01Architecture

See the interfaces among data, control, hardware and people.

02Measures

Translate capability into task, latency, failure and recovery.

03Evidence

Separate standards, independent measurement, research and first-party claims.

04Boundary

State what cannot be inferred from a demo, benchmark or interface.

System breakdown

Four interdependent layers determine whether the technology can enter real work.

Each layer shows its role and the failure signal most worth watching.

01

Multimodal representation

Images, language, robot state and action histories are encoded into a shared policy context.

02

Robot data

Demonstrations and trajectories teach contact-rich actions that web data cannot directly supply.

03

Policy and action interface

A model may output waypoints, skills or low-level actions, each with different latency and safety implications.

04

Runtime guardrails

Skill libraries, validators, monitors and human escalation bound what the model may do in the physical world.

Engineering evaluation

Five checks turn abstract capability into reviewable system evidence.

Record normal performance, failure, recovery and human cost—not only the best-looking result.

  1. 01

    Separate model from system

    Report model capability, controller, sensors, embodiment and operator assistance separately.

  2. 02

    Hold out environments and objects

    Evaluate on genuinely unseen combinations, not reordered scenes from the training distribution.

  3. 03

    Count interventions and resets

    Success rate should include retries, teleoperation, prompt changes and manual scene preparation.

  4. 04

    Evaluate recovery

    Introduce dropped objects, changed goals and partial failures rather than scoring only clean execution.

  5. 05

    Constrain physical authority

    Verify speed, force, workspace and tool constraints outside the generative model.

Scope boundaries

State what the evidence supports—and what it does not.

Sources and evidence status

Keep the source, date, evidence identity and reading boundary visible.

This page prioritises standards bodies, public measurement programmes, official project documentation and original research disclosures.