
World Labs unveils Atlas, a 4‑modal world model with spatial context
Fei-Fei Li’s startup World Labs has released Atlas, a next-generation world model billed as the first multimodal system capable of pixel-level camera-controllable image and video generation paired with native 3D scene reconstruction. Trained from scratch on text, image, video and 3D data, the so-called “omni model” unifies four distinct modalities within a single unified architecture. Its flagship innovation, termed “spatial context”, grounds each input image in a precise 3D coordinate space with explicit camera pose parameters. By treating geometric camera controls as native inputs, Atlas eliminates reliance on textual prompts to describe camera motion, letting users specify exact viewing angles and trajectories to generate up to 1440p-resolution videos with a maximum one-minute duration. World Labs frames the upgrade as letting...





