Frontier ObservationEmbodied Intelligence · 2 February 2026

Research domain 04
We follow the convergence of perception, spatial reasoning, manipulation, locomotion, world models and human-aware design as AI moves from screens into machines and shared environments.
FUURAA conceptual visualOpen research questions
What representations allow machines to reason robustly about space, force and uncertainty?
How can robots learn transferable skills without hiding the reality gap?
Which human controls and social norms are needed when autonomous systems share our spaces?
Reference points
These sources are independent references. Listing does not imply collaboration, approval or endorsement.
Multimodal reasoning and action in the physical world.
Open canonical source ↗02Carnegie Mellon Robotics InstituteLong-running work across autonomy, perception, manipulation and human-robot interaction.
Open canonical source ↗03IEEE SpectrumEngineering reporting that separates demonstrations from deployable systems.
Open canonical source ↗Related publications
Frontier ObservationEmbodied Intelligence · 2 February 2026
Research BriefMultimodal Intelligence · 11 February 2025
Live evidence radar
Figure demonstrates two humanoids completing a shared room-reset task without explicit message passing or a central planner.
→Observation-based teamwork can be flexible, but shared spaces introduce uncertainty about intent, timing and collision risk.
→π0.7 accepts varied conditioning such as language, visual subgoals and performance metadata to steer task strategy as well as task identity.
→The reported system recombines learned skills and follows coaching for tasks not directly represented by matched demonstrations.
→π0.7 reports transfer of a manipulation task to a robot configuration without matched task demonstrations on that embodiment.
→The research describes distilling experience from optimized specialist policies into a broader model while preserving strong task performance.
→The reported RL-token method keeps the main VLA fixed while training a smaller actor and critic on a compressed internal representation.
→Broad VLA competence can cover many tasks, while fine alignment and speed may still benefit from task-local reinforcement learning.
→Keep the question open
Researchers, engineers, universities and public institutions may propose sources, corrections or areas for structured review.