For three years, the story of AI has been written mostly in language — models that read, summarize, argue and write. In 2026 that story is turning. The next architecture, called a world model, learns physics the way a child learns it: by watching, predicting and being surprised. Instead of predicting the next token, it predicts the next state of a room, a road, a robot arm — the world itself, one step ahead. The money agrees. Fei-Fei Li’s World Labs just closed a $1 billion round. Yann LeCun’s AMI Labs raised $1.03 billion. Odyssey is valued at $1.45 billion. Decart hit nearly $4 billion with Nvidia writing checks. In a single quarter, four world-model companies raised roughly $3.5 billion, a scale the language-model boom took three years to reach.
The product surface is arriving fast. Google DeepMind’s Genie 3 generates persistent, controllable 3D environments from a single sentence. Runway’s GWM-1 does the same for cinematic scenes. World Labs’ Marble lets users step inside a generated world in real time. NVIDIA’s Cosmos platform is packaging world-model training for the robotics industry, and Meta’s V-JEPA 2 is doing the same for embodied agents. Waymo and Waabi have quietly rebuilt their entire autonomous-driving stacks around them.
The catch — and this report does not minimize it — is that the field cannot yet agree on what counts as a world model. Some are neural physics simulators. Some are generative video engines. Some are memory-augmented policies. On the What-If World benchmark, no leading system exceeds 52% on causal reasoning; open-source models cluster near 28%. The gap between “looks right” and “behaves like physics” is still the story.
From Words to Worlds walks through the architectures behind the shift, the leaders and the challengers, the four commercial fronts already emerging — robotics, autonomy, media and enterprise simulation — and the honest questions about scaling, safety and evaluation. If language models taught machines to speak, world models are teaching them where they are.