Vision-language-action models
Policies map multimodal observations and instructions toward robot actions.
Robotics & Embodied AI · Desk 02
Follow how perception, language, planning and action become policies that operate physical machines.
Desk thesis
Reader outcomes
Understand VLA, world models, imitation learning, reinforcement learning and classical control in one map.
Separate semantic reasoning from motor-control evidence.
Evaluate transfer across tasks, robots, sites and failure conditions.
Professional map
Every topic carries an engineering meaning and a question that must be answered.
Policies map multimodal observations and instructions toward robot actions.
Predictive representations support planning, counterfactuals and physical-scene understanding.
Demonstrations and rewards shape behaviour, coverage and failure modes.
Task planning, motion planning and feedback control operate at different time scales.
Tactile sensing, force control and bimanual coordination expose long-tail physical complexity.
Useful autonomy requires recognising uncertainty, requesting help and recovering from execution errors.
Source gateways
FUURAA connects readers to originals through professional translation and structured analysis; it neither reproduces source sites nor hides evidence boundaries.
Open models, datasets and tools for real-world robot learning.
Recent robotics preprints across manipulation, control, perception, planning and systems.
Broader AI preprints relevant to planning, reasoning, agents and embodied systems.
A shared format and collection spanning data from many robot embodiments.
Accessible explanations of research from BAIR researchers.
Open research on interactive agents, 3D understanding and embodied intelligence.
Scope & limits
Benchmark gains may not transfer to a different robot, camera, gripper or environment.
Language fluency is not evidence of physical reliability or safe control.
Public model reports often omit complete failure, intervention and duty-cycle distributions.