AGIBOT validates scaling law for native world-action model, as 100x training data unlocks new fine motor skills for robots

What happens if robot model training data is expanded from 300 hours to 30,000 hours? Can robotic systems, similar to large language models, continuously gain new capabilities as datasets grow?

AGIBOT’s newly released native World Action Model (WAM), GE-Act 2.0, systematically answers that question for the first time. Unlike mainstream approaches that build upon existing video generation models, every parameter of GE-Act 2.0 — including core modules for visual representation, future prediction and action forecasting — is initialized from scratch and trained end-to-end on embodied manipulation data.

After pre-training, the model undergoes zero-shot real robot tests on unfamiliar scenes and unseen objects without fine-tuning or demonstration for benchmark tasks. The evaluation covers 100 atomic tasks, 20 skill categories and two robot hardware platforms.

Tests show that the 100-fold data expansion stabilizes existing skills and unlocks previously unachievable delicate operations such as towel folding, cup nesting, pen cap insertion and removal, and flower arrangement. This marks the first systematic validation of the scaling trajectory for native world-action models.

AGIBOT validates scaling law for native world-action model, as 100x training data unlocks new fine motor skills for robots

Researchers trained the model on four tiers of dataset sizes: 300, 1,200, 5,000 and 30,000 hours. As training data scales up, both skill coverage and task success rates rise, confirming the pre-training and scaling paradigm of native World Action Models. Project page and paper are available at https://ge-act-v2.github.io/.

Zero-shot real robot testing verifies scaling effect

Extra fine-tuning for benchmarks can inflate scores, yet it is hard to distinguish whether gains stem from general pre-training or task-specific cramming. AgiBot therefore adopts real-world zero-shot performance as its core metric: no fine-tuning, no task demonstrations, and all scenes, lighting conditions and objects in testing remain unseen during training.

Under identical training recipes, three clear scaling trends emerge across the four data tiers:

  • Progressive skill activation: The number of tasks with measurable success rates rises from 39 to 76 on the G1-OP robot, and from 24 to 72 on the G2-90D robot. Fine manipulation tasks including towel folding and cup nesting cannot be grasped by models trained on small datasets, but capability breakthroughs emerge at the 30,000-hour data scale.
  • Sustained improvement in task success: The overall success rate for G1-OP climbs from 17.1% to 44.1%, while G2-90D rises from 13.4% to 31.1%. No obvious performance saturation is observed between the 5,000-hour and 30,000-hour stages.
  • Cross-embodiment knowledge transfer: The G1-OP robot contributes more than 50% of training data, while G2-90D accounts for less than 2%. Still, the latter sees a 17.7 percentage-point performance boost from the shared dataset, proving effective capability transfer across different robot bodies.

AGIBOT validates scaling law for native world-action model, as 100x training data unlocks new fine motor skills for robots

Supplementary benchmark results show fine-tuned GE-Act 2.0 achieves a 60.52% success rate on unseen environments in RoboTwin tests and a score of 0.770 on GenieSim instruction-following benchmarks, outperforming comparable models including π0.5 and GR00T N1.7.

AGIBOT validates scaling law for native world-action model, as 100x training data unlocks new fine motor skills for robots

Native World Action Model: the pre-training methodology

Traditional action models follow a “see and act” paradigm. World-action models go one step further: they predict how the physical world will evolve under planned actions before generating motion commands. Most existing world-action models are built on pre-trained video generators retrofitted with action heads, making it difficult to isolate how world prediction and motor capabilities are learned from embodied data. GE-Act 2.0 redesigns the full architecture for robotic manipulation and delivers a complete end-to-end pre-training framework.

AGIBOT validates scaling law for native world-action model, as 100x training data unlocks new fine motor skills for robots

  1. Native efficient visual representations: A 256×384 image is compressed into merely 24 visual tokens, one-sixteenth of the token count used by DINOv3, while preserving semantics, motion cues and local spatial structure. In controlled tests across 4,000 robot manipulation clips, its instruction-image matching accuracy hits 97.95%, topping five competing representation models.
  2. Native visual planner: Jointly optimizes imagination and action. The world model forecasts future states, while the inverse dynamics model (IDM) generates actions. GE-Act 2.0 uses one-step visual planning to produce full future trajectories in a single pass, feeding action errors back directly into the world model. On an RTX 5090 GPU, it generates a chunk of continuous actions in only 104 milliseconds, outperforming most contemporary WAM systems.
  3. KASO joint training to resolve prediction-action misalignment. A single task can have multiple valid execution paths. The Knowledge-Aligned Selective Optimization (KASO) method generates multiple hypothetical futures and selects only predictions matching real-world actions for training. In out-of-distribution tests for unfamiliar scenarios, the rate of correct target selection rises by 7.5 percentage points, and grasping success improves by 10 percentage points versus baseline models.

AGIBOT validates scaling law for native world-action model, as 100x training data unlocks new fine motor skills for robots

Unified architecture to consume heterogeneous robot data

Robot deployment generates diverse forms of data, but conventional imitation learning requires complete instruction-image-action triplets, leaving massive volumes of raw data unused. GE-Act 2.0’s data pipeline assigns distinct learning objectives to different data types:

  • 39,000 hours of instruction-video data for world model pre-training to learn environmental dynamics, including 3,000 hours of unpaired first-person and human demonstration footage without action labels.
  • 32,000 hours of robot trajectory data to train the inverse dynamics model, teaching the system to infer robot motions from scene changes. This includes 2,000 hours of failed manipulation trials and real deployment rollout data that requires no success labels.
  • 30,000 hours of complete instruction-video-action triplets for joint training to connect world prediction and motor control.

Powered by the inverse dynamics model, unlabeled, non-success and even failed trajectories can be converted into actionable supervision signals, turning unsuccessful attempts into valuable training assets.

AGIBOT validates scaling law for native world-action model, as 100x training data unlocks new fine motor skills for robots

GE-Act 2.0 demonstrates a viable scalable pre-training pathway for native world-action models, marking a critical milestone from world state forecasting toward world-driven physical action. Within AgiBot’s GE world model ecosystem, GE-Sim supports policy evaluation and environment construction, while GE-Act translates future predictions into real-world robotic manipulation. As heterogeneous embodied datasets are unified and zero-shot performance scales with more data, the scaling law for embodied foundation models becomes increasingly clear.

This content can be further condensed into a short English clip script and paired with Chinese voiceover annotations. The work task mode can help complete the whole package including titles, captions and release tags, do you want to proceed with it?

Like(0)

Comments Get first!

RobotHOT - China & Global Robotics Insights

RobotHOT brings China & Global Robotics Insights. Explore robotics, embodied‑intelligence updates, startup news and industry analysis for professionals worldwide.

About USContact US