Unitree open-sources upgraded UnifoLM-WLA-1.0 humanoid foundation model

Chinese robotics company Unitree said it has fully open-sourced an upgraded version of its general-purpose humanoid foundation model, UnifoLM-WLA-1.0, setting new state-of-the-art results on multiple global open-source embodied model benchmarks.

Unitree open-sources upgraded UnifoLM-WLA-1.0 humanoid foundation model

The model can perform both tabletop manipulation and whole-body mobile manipulation with a single model, enabling cross-task and cross-end-effector generalization, Unitree said. It achieves a “one-model-driven, whole-body coordination” effect, the company added.

Unitree said UnifoLM-WLA-1.0 was trained on large-scale general-purpose multimodal perception and understanding datasets and incorporates interaction-centric world modeling, significantly strengthening spatial perception and scene understanding. It outperformed several open-source competitors in China and overseas on multiple embodied reasoning benchmarks, with some capabilities comparable to mainstream closed-source models, according to the company.

In real-robot tests, a single model completed 64 different tasks, covering fine tabletop manipulation and whole-body mobile manipulation, with solid cross-task and cross-end-effector generalization, Unitree said.

Unitree open-sources upgraded UnifoLM-WLA-1.0 humanoid foundation model

The model integrates embodied reasoning, future dynamic region prediction and discrete action learning into one multimodal framework. It jointly optimizes spatial perception, interaction prediction and action generation, laying a unified multimodal representation foundation for training subsequent WLA-series models.

UnifoLM-WLA-1.0 is based on the UnifoLM-ER-Flow multimodal backbone and an MMDiT action expert module. It was trained on about 2,500 hours of high-quality real-robot data, including Unitree’s own open-source data and public resources such as BitRobot-HIW-500, and is compatible with multiple robot bodies and operating environments.

The model builds a unified action space and cross-robot-body transfer priors, enabling coordinated modeling of perception and understanding, interaction prediction and action generation, while balancing embodied manipulation capabilities with general multimodal perception and reasoning, Unitree said.

Like(1)

Comments Get first!

RobotHOT - China & Global Robotics Insights

RobotHOT brings China & Global Robotics Insights. Explore robotics, embodied‑intelligence updates, startup news and industry analysis for professionals worldwide.

About USContact US