Fei-Fei Li’s startup World Labs has released Atlas, a next-generation world model billed as the first multimodal system capable of pixel-level camera-controllable image and video generation paired with native 3D scene reconstruction. Trained from scratch on text, image, video and 3D data, the so-called “omni model” unifies four distinct modalities within a single unified architecture.

Its flagship innovation, termed “spatial context”, grounds each input image in a precise 3D coordinate space with explicit camera pose parameters. By treating geometric camera controls as native inputs, Atlas eliminates reliance on textual prompts to describe camera motion, letting users specify exact viewing angles and trajectories to generate up to 1440p-resolution videos with a maximum one-minute duration. World Labs frames the upgrade as letting creators “sit in the director’s chair instead of pulling a slot machine lever”.

Built as a multimodal autoregressive diffusion Transformer, Atlas merges proven LLM optimization techniques with cutting-edge video generation advances. It supports camera-controlled novel-view synthesis from one to six reference images, high-fidelity spatial reconstruction with point clouds and Gaussian splatting outputs, spatiotemporal simulation for video reframing and robotic real-to-sim pipelines, as well as text-to-image and 360° panoramic generation.

In human blind evaluations for controllable camera generation, Atlas outperforms mainstream alternatives with leading win rates: 75% vs MiniMax H3, 81% vs Gemini Omni Flash, 86% vs Alibaba HappyHorse 1.1, 93% vs FLUX 3, and 94% vs ByteDance Seedance 2.5, with stronger advantages on complex motion trajectories. World Labs clarifies the gap largely stems from its native camera-pose support, as competing models depend entirely on text-driven camera control. In sparse-view 3D reconstruction tests, Atlas achieves an AbsRel error of 25.3 (×10⁻³), outperforming specialized open-source models including Pi3X, π³, VGGT-Ω 1B, Depth Anything 3 and MapAnything, while the firm notes it does not dominate every benchmark dataset.

World Labs confirms observable compute scaling effects but has not disclosed model parameters, training scale, scaling curves or inference costs. Atlas remains closed-source with limited early partner access and will underpin future World Labs product lines including Marble. Demo results show additional reference images significantly reduce model hallucination: one input yields heavy generative filling, while two to three shots deliver faithful reconstruction. The model can process over 100 images for large-scale scene rendering, as demonstrated by aerial flyover footage of Stanford’s Main Quad generated solely from 2–25 ground-level photographs.






