cd /news/artificial-intelligence/world-labs-launches-atlas-for-video-… · home topics artificial-intelligence article
[ARTICLE · art-118146] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

World Labs launches Atlas for video, 3D reconstruction and robot simulation

World Labs, the San Francisco AI company co-founded by Fei-Fei Li, launched Atlas on September 1, a single model for generating controlled video, reconstructing 3D scenes, and producing simulated camera views for robots, with access initially limited to selected partners. Atlas accepts text, images, camera poses, and 3D depth maps, and can generate up to one minute of 1440p video from one to six reference images and a camera path.

read6 min views2 publishedSep 1, 2026
World Labs launches Atlas for video, 3D reconstruction and robot simulation
Image: Runtimewire (auto-discovered)

Fei-Fei Li's company is putting camera control, scene reconstruction and synthetic robot views into one model, with access initially limited to selected partners.

By RuntimeWire Staff · Published

Primary source: World Labs

Why it matters #

Atlas turns World Labs' spatial-intelligence thesis into infrastructure for creators and robotics teams. Its value depends on whether company-run results hold across partner workloads and independent testing.

World Labs, the San Francisco AI company co-founded by Fei-Fei Li, launched Atlas on September 1, presenting one model for generating controlled video, reconstructing 3D scenes and producing simulated camera views for robots.

Atlas turns Li's spatial-intelligence thesis into a base-model strategy. She has argued that AI needs to reason about objects, places and consequences across three-dimensional space. World Labs has built an architecture meant to handle those tasks within the same model instead of dividing them among separate video, reconstruction and simulation systems.

Li brings an unusually long view of computer vision to that bet. She helped create ImageNet, directed the Stanford AI Lab from 2013 to 2018 and spent a Stanford sabbatical as vice president and chief scientist of AI and machine learning at Google Cloud. She is also a founding co-director of Stanford's Institute for Human-Centered Artificial Intelligence. The move from ImageNet to World Labs follows a clear technical progression: teaching machines to recognize what appears in an image, then building systems that can represent the space beyond its visible pixels.

Her co-founders cover the other pieces Atlas needs. Justin Johnson completed his Stanford Ph.D. under Li and worked on visual reasoning, image generation and 3D reasoning at the University of Michigan and Facebook AI Research. Ben Mildenhall, a former Google research scientist, co-created Neural Radiance Fields, or NeRF, an influential method for synthesizing new views of a scene from images. Christoph Lassner held research roles at Meta Reality Labs and Epic Games, where he led research teams, and worked at Body Labs before its acquisition by Amazon.

One context for several kinds of spatial data

World Labs describes Atlas as a multimodal autoregressive diffusion transformer pretrained from scratch. It accepts text, images, camera poses and 3D depth maps, with videos represented as sequences of images. Those inputs are combined into what the company calls a shared spatial context, grounding each image at a position in three-dimensional space.

The explicit treatment of camera geometry is the central technical choice. Conventional video generators commonly receive camera instructions through text prompts such as pan, crane or truck. Atlas receives a camera path as native model input, giving users direct control over the position and angle of generated views.

World Labs says Atlas can generate up to one minute of 1440p video from one to six reference images and a manually designed camera path. It can also place unrelated reference images at chosen positions inside a scene, then generate the visual transitions between them. That gives filmmakers and game developers a way to stage shots and environments with repeatable geometry instead of repeatedly prompting for an acceptable camera movement.

Atlas also works in the other direction. World Labs says it can reconstruct a real location from one or more images, generate views from unseen angles and export explicit geometry as point clouds or 3D Gaussian splats. The company says two or three images are typically enough for a faithful reconstruction, while additional images reduce the amount of scenery the model must invent.

That distinction matters in production work. A plausible building placed behind the camera may be useful for a fictional environment and unacceptable when reconstructing a factory, store or real estate property. Atlas is designed to shift between those cases by using the amount of source material as a control over how much it generates.

The robotics plan is coming into view

World Labs has described a real-to-sim workflow for reconstructing physical tasks as simulations and varying appearance, object configuration, clutter, physics, robot state and camera viewpoint in a technical post.

Atlas supplies part of that pipeline. World Labs says the model can reconstruct an environment from phone footage and generate the RGB images and depth readings that a robot-mounted camera would observe while moving through the simulated space. For manipulation tasks, the company says Atlas contributes visual and geometric reconstruction to simulations involving rigid, articulated and deformable objects.

Physical robot training consumes hardware time, requires repeated resets and exposes equipment to failures. A sufficiently accurate world model could let robotics developers test more conditions in simulation before returning to the machine. World Labs still has to show that Atlas-generated sensor views and scene geometry remain accurate across customer hardware, unfamiliar spaces and long sequences of interaction.

Creative tools provide an earlier route into the market. World Labs says Atlas will power future versions of Marble, its product for generating persistent 3D environments, along with other products. Marble already has public plans ranging from a free tier to Standard at $20 a month, Pro at $35 and Max at $95, plus custom enterprise pricing. Atlas itself has no disclosed public price or general-availability date.

As RuntimeWire reported on August 19, Li has been positioning persistent, navigable environments as the bridge between creative software and systems that act in physical space.

Atlas enters a field with several definitions of a world model. Odyssey is developing interactive world simulation, Niantic Spatial focuses on geospatial mapping and visual positioning, and Yann LeCun's AMI Labs is pursuing controllable models designed to understand and plan in the physical world. World Labs' pitch centers on a single model that accepts spatially grounded multimodal inputs and produces video, explicit 3D geometry and simulated sensor views.

World Labs sets its own test

World Labs published internal evaluations for camera-controlled generation and sparse-view 3D reconstruction. For the camera test, the company paired an input image with one to three cinematic camera motions. Atlas received camera geometry directly, while competing models received text descriptions of those paths. Third-party human raters judged how well each output followed the intended camera movement, and World Labs reported that Atlas's advantage increased with more complex trajectories.

World Labs says Atlas can reconstruct scenes from sparse views and generate explicit 3D outputs, but the supplied materials do not establish independent benchmark results against competing systems.

Atlas is entering early access with selected partners. The number and identities of those partners, along with pricing and a general-availability date, remain undisclosed. Access is available through a request form, so independent developers cannot yet broadly reproduce the company-run results or test how the model behaves outside World Labs' selected demonstrations. The company has not disclosed customer, usage or revenue figures for Atlas.

The launch follows a large buildout. World Labs raised $1 billion in February from investors including AMD, Autodesk, Emerson Collective, Fidelity Management & Research, Nvidia and Sea. The company had previously raised $230 million in financing backed by Andreessen Horowitz, NEA and Radical Ventures, bringing disclosed funding to roughly $1.23 billion.

Atlas shows where that capital is going. World Labs says Atlas performed better as the company trained models with increasing size and compute, and it expects further scaling to add capabilities. That is the company's own scaling result, not independent evidence. The potential customers range from directors controlling virtual cameras to robotics engineers generating synthetic sensor data.

Early access now shifts the burden from polished examples to partner workloads. Atlas has to preserve geometry when reference images disagree, separate reconstruction from plausible invention and produce simulations that remain useful after a robot begins interacting with the scene. Those tests will determine whether Li's spatial-intelligence program can support a distinct model category or remains a collection of adjacent computer-vision tasks.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @world labs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/world-labs-launches-…] indexed:0 read:6min 2026-09-01 ·