World Labs releases Atlas, an omni world model trained from scratch on text, images, video and 3D in one shared spatial context
World Labs introduced Atlas on 2026-09-01, a multimodal autoregressive diffusion transformer that takes text, images, video and 3D as native input types and generates what comes next while staying 3D-consistent with everything already seen. Concrete capabilities include camera-controlled generation of up to one minute of 1440p video from one to six reference images with explicit camera geometry as an input type rather than a text instruction, spatial reconstruction from one to dozens of images that outputs explicit 3D and beats specialized reconstruction models, and space-time simulation for real-to-sim robotics workflows. Atlas will back future versions of Marble and is currently early access only.
Source
↳ Follow the thread