Reconstructing a 3D scene used to require hundreds of carefully captured photographs, expensive equipment, and a healthy dose of patience. World Labs, the spatial intelligence company co-founded by Stanford AI luminary Fei-Fei Li, just compressed that workflow down to two or three snapshots.
The company’s new model, called Atlas, is a pretrained multimodal autoregressive diffusion transformer. In plainer terms: it’s an AI system that understands text, images, video, and 3D data all at once, and can use that understanding to fill in everything a camera didn’t capture. The result is full 3D scene reconstruction, novel view synthesis, and camera-controlled video generation, all from minimal input.
What Atlas actually does
Atlas operates as what World Labs calls an “omni world model” for spatial intelligence. Feed it between one and six reference images, and it can generate up to one minute of video at 1440p resolution while maintaining precise camera control along exact paths you specify.
The 3D reconstruction capability is where things get genuinely impressive. Traditional photogrammetry pipelines typically demand anywhere from 100 to 300+ images shot from carefully planned angles to build a usable 3D model. Atlas collapses that requirement to as few as two or three input images, producing explicit 3D representations like point clouds and Gaussian splats with what the company describes as remarkable fidelity.
On sparse-view 3D reconstruction benchmarks, including DTU and ETH3D, Atlas posted a mean absolute-relative pointmap error of 25.3. Its closest open-source competitor scored 28.7.
World Labs’ co-founders frame new view prediction as “an AI-complete primitive for spatial intelligence.” Translation: the ability to imagine what a scene looks like from angles you’ve never photographed is a foundational capability, one that unlocks everything from robotics navigation to visual effects pipelines.
The human verdict
In consumer tests evaluating camera control quality, human raters preferred Atlas over competing models between 75% and 94% of the time.
Atlas also supports dynamic space-time simulations, including bullet-time effects. Think the iconic slow-motion rotating camera move from The Matrix, but generated from AI rather than an elaborate rig of synchronized cameras.
The model handles all of this within a unified spatial context. Rather than stitching together separate tools for video generation, 3D reconstruction, and novel view synthesis, Atlas consolidates these functions into a single system.
From lab to limited release
World Labs has been building toward this moment since the company’s founding in 2024. The startup attracted significant attention from the outset, largely because of Li’s reputation as one of the architects of modern computer vision. Her ImageNet dataset and associated challenges played a central role in catalyzing the deep learning revolution over a decade ago.
Atlas was announced on September 1, 2026, and is currently entering an early access phase for select partners. World Labs has not released the model’s weights or code publicly, keeping the technology behind a controlled access wall for now.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
20









English (US) ·