TL;DR — Key Takeaways
- World Labs’ Atlas model is designed to generate and reconstruct 3D environments while reasoning about camera position, motion and physical space.
- Atlas can combine text, images, video and 3D data into a shared spatial representation of a scene.
- The model can generate new viewpoints, depth maps, point clouds and 3D Gaussian splats, making it useful beyond traditional video generation.
World Labs, the AI startup co-founded by computer vision pioneer Fei-Fei Li, has unveiled Atlas, a new AI model that can generate and reconstruct 3D environments and simulate how objects and machines operate within them.
The company is targeting a major challenge in AI development: giving world models a deeper understanding of physical space. Atlas can work with text, images, video and 3D data, combining those inputs into a shared spatial context that represents a scene.
The technology could have significant implications for robotics, where developers need realistic virtual environments to train and test machines before deploying them in the physical world.
World Labs has raised $1.2 billion from investors that include NVIDIA, AMD and Autodesk. Li launched the company in 2024 after previously serving as director of Stanford University’s AI Lab and co-founding the Stanford Institute for Human-Centered AI. Her work at World Labs is centered on spatial intelligence, an approach that seeks to give AI the ability to reason about environments and physical interactions.
Uses Camera Geometry and Trajectories
Atlas is based on a highly complex idea: multimodal autoregressive diffusion transformer architecture. In simple terms, this is an AI design that builds complex outputs step by step, using earlier information to guide each new piece while progressively refining the result.
Unlike AI video tools that rely largely on text instructions to determine a virtual camera’s movement, the model can use camera geometry and trajectories as direct inputs. This allows developers to specify where a virtual camera is located and how it moves through a generated environment.
A key feature is the amount of visual information Atlas can create from limited source material. The model can use a single image to produce video lasting as long as one minute at 1440p resolution. It can also create views of areas that were not visible in the original image, although these unseen areas are generated rather than reconstructed from actual visual data.
World Labs said Atlas also reconstructs scenes captured with three to five ordinary cameras, allowing video to be viewed from new angles without specialized capture systems.
Atlas also produces point clouds, depth maps and 3D Gaussian splats, giving developers outputs that can be used beyond conventional video. These capabilities broaden its potential role in gaming, visual effects, design and robotics.
Robotics may prove to be the more important enterprise use case. World Labs demonstrated a Real-to-Sim workflow in which developers record a physical location with a smartphone and convert it into a virtual environment. A simulated robot can then navigate the resulting environment while Atlas generates the RGB images and depth information that its sensors would encounter.
Atlas is currently available through early access for selected enterprise users. World Labs has not announced a date for general availability. The model is also expected to provide underlying technology for future versions of the company’s Marble platform.