Caltech, Stanford unveil HomeBody system to give GPT Astra spatial memory for humanoid robots

Researchers at Caltech and Stanford have unveiled HomeBody, a system that gives GPT Astra persistent spatial memory and the ability to execute compound humanoid actions for long-horizon tasks. In a kitchen demo, a Unitree G1 robot guided by Astra cleaned the room and retrieved remembered items from ambiguous requests. The team said the system requires no environment-specific training or additional policy learning.

HomeBody

HomeBody replaces the trained intermediate VLA layer in conventional three-part architectures—System2 VLM for vision and instructions, System1 VLA for control commands, and System0 for coordinated motion. As frontier models such as Astra improve, the researchers argue trained VLAs are no longer necessary and System2 can directly manage reusable motion skills. A plug-and-play VLM calls a composable skill library.

HomeBody

Cleaning a kitchen requires deciding what to keep or discard and where to move each item. Because the scene changes with every movement, the robot relies on memory and action feedback to track completed and pending work. In videos, HomeBody placed coffee bags on a counter and discarded a designated milk carton, coordinating multiple trips, grasps and placements. It did not perform common rummaging actions and took a long time overall. In another demo, the robot used stored keyframes to locate a drawer, retrieve medicine, hand it to a user and discard a drink carton. It showed two-handed coordination, but the video was heavily accelerated and efficiency remains limited.

HomeBody

To develop the system, the team let the humanoid explore an unseen environment. Spatial data came from iPhone 0.5x video, D435i camera observations, SLAM lidar scans, joint poses and Astra-selected waypoints. HomeBody recorded the room from the robot’s first-person view, retaining context even after objects left its field of view. Astra then acted as a Real2Sim agent, building a digital twin in Isaac Sim. In the virtual world, commands such as “tidy the kitchen” used spatial context to select actions and targets without action-level scripting.

The team said HomeBody’s skills can run on a Razer Blade laptop with an RTX 4090 GPU, supporting perception and motion planning. GPT Astra runs remotely, sending skill requests and targets to the laptop and receiving execution results. The lightweight setup allows local skill execution while accessing frontier models over the network. But the team acknowledged that Astra’s inference latency causes pauses between skills, and more complex perception or skills may require more computing power. Direct large-model control of robots still has a long way to go.

HomeBody

Like(0)

Comments Get first!

RobotHOT - China & Global Robotics Insights

RobotHOT brings China & Global Robotics Insights. Explore robotics, embodied‑intelligence updates, startup news and industry analysis for professionals worldwide.

About USContact US