cd /news/artificial-intelligence/world-model-artificial-intelligence · home topics artificial-intelligence article
[ARTICLE · art-123161] src=en.wikipedia.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

World Model (Artificial Intelligence)

World models are machine learning systems that build internal representations of environments to predict changes over time, enabling agents to plan and act without real-world trial and error. Jürgen Schmidhuber introduced the term in 1990, and David Ha and Schmidhuber revived it in a 2018 paper. Google DeepMind's Genie 3, introduced in August 2025, generates photorealistic, real-time interactive worlds from text prompts at 24 frames per second, and Waymo adopted it in February 2026 to create the Waymo World Model for autonomous driving simulation.

read11 min views1 publishedSep 8, 2026

A world model in artificial intelligence is a machine learning system that builds an internal representation of an environment. Often this is via understanding objects within video, which predictive LLMs cannot. The model predicts how that environment changes over time in response to actions. Researchers design world models to help agents plan, reason, and act without constant real-world trial and error. World models differ from systems that merely classify or generate outputs. They simulate dynamics such as physics, object interactions, and causality. Early ideas date to the 1990s. Modern versions power robots, autonomous driving, and interactive video generation.

History #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=1>)]

Jürgen Schmidhuber introduced the term world model in machine learning in 1990.<sup>[1]</sup> He proposed recurrent neural networks that predict future states from observations and use those predictions to train agents. David Ha and Schmidhuber revived the concept in a 2018 paper. Their agents learned to drive virtual cars and play video games inside self-generated simulations.[2]

Yann LeCun advanced the idea in a 2022 position paper titled "A Path Towards Autonomous Machine Intelligence".<sup>[3]</sup> He argued that intelligence requires predictive models of the world rather than pure pattern matching. LeCun proposed the joint embedding predictive architecture (JEPA) as a practical foundation. LeCun and collaborators developed several JEPA variants. V-JEPA 2 reached state-of-the-art performance on video understanding and physical reasoning at the time.<sup>[4]</sup> It supports zero-shot robot control in unfamiliar environments.<sup>[4]</sup> Introduced in March 2026, LeWorldModel trains stably end-to-end from raw pixels and uses two loss terms and avoids hand-crafted heuristics.<sup>[5]</sup> LeCun founded Advanced Machine Intelligence Labs in 2026 to further develop world models.[6][7]

Google DeepMind introduced Genie in 2024. The model learned interactive environments from unlabeled internet videos. Genie 2 followed in late 2024 and added three-dimensional generation. The Genie series set benchmarks for general-purpose simulation.

Genie 3 was introduced in August 2025. It produces photorealistic, real-time interactive worlds from text prompts which are displayed at 24 frames per second and explored in real time with text or image prompts. The model supports persistent three-dimensional worlds and real-time interaction.[8]Waymo adopted Genie 3 in February 2026 and used it to create a specialized world model for autonomous driving simulation, called the Waymo World Model. It produces synchronized camera and lidar outputs and creates edge cases that real robotaxis rarely encounter. The edge cases were reported to be unusual by PCMag.[9]

General Intuition announced a $133.7 million seed round.<sup>[10]</sup> World Labs raised $1 billion. AMI raised $1.03 billion.[11]

In April 2026, Alibaba announced Happy Oyster, its world model designed for real-time and “flowy” world model. It includes a directing mode for world building based on text and image prompts and a wandering mode for exploring the resulting world. It can generate 3-minute in-world video clips.[12]

Also in April, World Labs, co-founded by Li Fei Fei, unveiled Spark 2.0, an open-source 3D Gaussian splatting rendering engine that targets smartphone-class devices.[12]

In June 2026, Nvidia released Cosmos 3, a family of open-weight models. It combines previously independent physical reasoning, world simulation, and action generation. Cosmos 3 integrates can process and generate text, image, video, audio, and action sequences. The model employs a Mixture-of-Transformers" (MoT) approach. An autoregressive (AR) transformer handles reasoning and next-token prediction, while a diffusion transformer (DT) does multimodal generation. Encoders (ViT for vision, VAE for visual/audio, and domain-specific for actions) and generate a shared representation space using 3D multi-dimensional rotary position embedding (mRoPE) for spatial and temporal information. The family includes Cosmos3-Nano (16B parameters) for workstations; Cosmos3-Super (64B parameters) for research.[13]

Conceptualization #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=2>)]

World models broadly refer to internal representations used to predict or simulate an environment, but the exact definition remains unsettled and continues to evolve as understeanding grows across artificial intelligence, robotics, cognitive science and computational neuroscience.[14][15]<sup>[16]</sup> World-model learning and inference were used broadly to study how systems predicts states, interpret sensory information and guide action.[16][17]<sup>[18]</sup> Predictive coding, where predictions are checked against observations, and differences betweeen them used to update internal moels, provides one framework for comparing world-model learning across artificial intelligence, robotics and neuroscience.[16][17][18]

In machine learning, a world model can be learned from experience and used to simulate prossible outcomes, allowing an agent to plan or learn behavior without testing every possibility directly in the environment.[19][2][20]<sup>[21]</sup> World modeling therefore describes a general approach to representation, learning and decision-making rather than a particular model architecture.[15]<sup>[3]</sup> Additional examples of this approach include planning and behavior learning within learned latent models.[22][23]

Architecture #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=3>)]

World models process raw sensory data such as video frames or lidar scans. They compress this input into compact latent representations. The system then predicts future representations rather than pixel-by-pixel reconstructions.

Many modern world models use joint embedding predictive architecture (JEPA). An encoder turns observations into embeddings. A predictor estimates one or a suite of embeddings from the current one and an action. In some cases a critic chooses one embedding as the best result. A regularizer keeps embeddings well-behaved.[4]

The model trains by minimizing prediction error in embedding space. This approach avoids the high cost of generating every detail. Some architectures add explicit components. A fast reactive path handles immediate responses. A slower deliberative path performs longer-horizon planning. Video prediction accuracy or robot success rates are key metrics, but do not always predict real-world performance.

Generative world models such as Genie 3 combine these with a simulator. They accept text prompts or layouts and output consistent video, lidar, or three-dimensional scenes. World models often train with self-supervised learning. They use large unlabeled datasets of video or robot interactions. Self-supervised learning can speed learning. Reinforcement learning can fine-tune a model for specific tasks.

Applications #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=4>)]

World models support robot learning. Agents train inside simulations and transfer skills to the physical world. This reduces the need for dangerous or expensive real-world trials. Autonomous vehicles use world models to test rare events.<sup>[24]</sup> Waymo's system simulates tornadoes or unusual pedestrian behavior. Companies train planners without putting vehicles on public roads. Interactive entertainment benefits from world models. Genie 3 lets users generate playable environments from simple descriptions. Game studios prototype levels faster. Scientific simulation gains from these models. Researchers model physical systems or biological processes at scale. Planners in logistics or urban design test strategies inside accurate digital twins.

Comparison with large language models #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=5>)]

Both world models and large language models (LLMs) use inferencing on their inputs to make predictions.

LLMs operate on textual inputs. They predict the next token in text sequences. They excel at language-oriented tasks such as translation or summarization. However, they lack understanding of physics.

World models operate on sensor inputs such as pixels. They predict state changes in that data in latent space. This design supports planning and causal reasoning.

LLMs generate fluent text but often fail at consistent physical predictions. Their architecture employs transformers with refinements such as mixture of experts.

World models divide an inferencing task into work performed by encoders, predictors, simulators, and other pieces. They typically handle multimodal inputs such as video, lidar, radar, and audio, guided by textual prompting.

LLMs power chatbots and code assistants. World models drive embodied agents that act in dynamic environments, such as autonomous driving. The two may be combined in hybrid systems. For example, a LLM handles instructions, while a world model manages low-level control. World model proponents such as LeCun claim that because LLMs are trained only on text, they have no ability to predict anything beyond text, such as real-world events.[3][25]

Benchmarks #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=6>)]

World model benchmarks test physical understanding, long-term consistency, planning, and generalization from sensor data.

Meta introduced three benchmarks for V-JEPA 2.[26]

  • IntPhys 2 measures a model's ability to detect physics violations. It presents pairs of videos that diverge when one breaks physical rules. Humans score near 100% accuracy. V-JEPA 2 achieves little better than random chance on many conditions.<sup>[27]</sup>
  • Minimal Video Pairs (MVPBench) tests physical understanding through multiple-choice questions based on short video clips. It probes object interactions and causality.<sup>[28]</sup>
- Something-Something tests action recognition.<sup>[\[29\]](#cite_note-29)</sup>
- Epic-Kitchens-100 tests human action anticipation.

DeepMind benchmark:

  • Interactive evaluation measures consistency over minutes of interaction, memory of off-screen objects, and response to user actions or text prompts.<sup>[30]</sup> Waymo benchmark:

  • Output generation quality: Metrics include realism, controllability (via text prompts), and usefulness for training planners in simulated worlds. However, pixel reconstruction error rate with episodic rewards often fails.

Other:

  • Epic-Kitchens-100 (often measured with Recall@5)<sup>[31]</sup>
  • Ego4D
  • 50 Salads, Breakfast, etc.

Potential benchmarks:

- Zero-shot transfer to robots
- Long-horizon planning
  • Implausible prediction rate

See also #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=7>)]

References #

[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=8>)]
  1. "1990: Planning & Reinforcement Learning with Recurrent World Models and Artificial Curiosity" .people.idsia.ch . Archived fromthe original on 2026-03-08. Retrieved 2026-03-25.
  2. 12 Ha, David; Schmidhuber, Jürgen (2018)."Recurrent World Models Facilitate Policy Evolution" .Advances in Neural Information Processing Systems 31 (NeurIPS 2018) . Curran Associates. pp. 2451–2463.arXiv :1809.01999 .
  3. 123 LeCun, Yann (June 27, 2022)."A Path Towards Autonomous Machine Intelligence Version 0.9.2" (PDF).
  4. 123"Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning" . Meta AI. June 11, 2025. Retrieved March 25, 2026.
  5. Maes, Lucas; Lidec, Quentin Le; Scieur, Damien; LeCun, Yann; Balestriero, Randall (2026-03-13). "LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels".arXiv :2603.19312 [cs.LG ].
  6. Metz, Cade (March 10, 2026)."Former Meta A.I. Chief's Start-Up Is Valued at $3.5 Billion" .The New York Times – via NYTimes.com.
  7. Orru, Mauro (March 10, 2026)."Former Meta AI Pioneer Yann LeCun Raises Over $1 Billion for New Startup" .The Wall Street Journal .
  8. Whitwam, Ryan (2025-08-05)."DeepMind reveals Genie 3 "world model" that creates real-time interactive simulations" .Ars Technica .
  9. Martindale, Jon (February 6, 2026)."Waymo Is Using Google's Genie 3 AI to Practice Handling Tornadoes, Elephants" .PCMag . Retrieved April 6, 2026.
  10. Bellan, Rebecca (October 16, 2025)."General Intuition lands $134M seed to teach agents spatial reasoning using video game clips" .TechCrunch . Retrieved June 30, 2026.
  11. McCormick, Packy (19 March 2026)."World Models: Computing the Uncomputable" .www.notboring.co . Retrieved 2026-04-09.
  12. 12"Chinese tech giants, AI 'godmother' Li Fei-Fei race into world models" .South China Morning Post . 2026-04-16. Retrieved 2026-04-18.
  13. "Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action" .huggingface.co . 2026-06-01. Retrieved 2026-06-01.
  14. Sakagami, Ryo; Lay, Florian S.; Dömel, Andreas; Schuster, Martin J.; Albu-Schäffer, Alin; Stulp, Freek (2023-11-02)."Robotic world models—conceptualization, review, and engineering best practices" .Frontiers in Robotics and AI .10 1253049. Frontiers.doi :10.3389/frobt.2023.1253049 .ISSN2296-9144 .PMC10652279 .PMID38023585 .
  15. 12 Moerland, Thomas M.; Broekens, Joost; Plaat, Aske; Jonker, Catholijn M. (2023-01-04)."Model-based Reinforcement Learning: A Survey" .Foundations and Trends in Machine Learning .16 (1): 1–118.doi :10.1561/2200000086 .ISSN1935-8237 .
  16. 123 Taniguchi, Tadahiro; Murata, Shingo; Suzuki, Masahiro; Ognibene, Dimitri; Lanillos, Pablo; Ugur, Emre; Jamone, Lorenzo; Nakamura, Tomoaki; Ciria, Alejandra; Lara, Bruno; Pezzulo, Giovanni (2023-07-03)."World models and predictive coding for cognitive and developmental robotics: frontiers and challenges" .Advanced Robotics .37 (13): 780–806.doi :10.1080/01691864.2023.2225232 .ISSN0169-1864 .
  17. 12 Friston, Karl; Moran, Rosalyn J.; Nagai, Yukie; Taniguchi, Tadahiro; Gomi, Hiroaki; Tenenbaum, Josh (2021-12-01)."World model learning and inference" .Neural Networks .144 : 573–590.doi :10.1016/j.neunet.2021.09.011 .ISSN0893-6080 .PMID34634605 .
  18. 12 Ohmae, Shogo; Ohmae, Keiko (August 2026)."Brain-AI convergence: Generative world models and hierarchical attention for human intelligence" .Patterns .7 (8) 101593.doi :10.1016/j.patter.2026.101593 .ISSN2666-3899 .PMC13494634 .PMID42630565 .
  19. Sutton, Richard S. (1990)."Integrated Architectures for Learning, Planning, and Reacting Based on Approximating Dynamic Programming" .Machine Learning Proceedings 1990 . pp. 216–224.doi :10.1016/B978-1-55860-141-3.50030-4 .ISBN978-1-55860-141-3 .
  20. Schrittwieser, Julian; Antonoglou, Ioannis; Hubert, Thomas; Simonyan, Karen; Sifre, Laurent; Schmitt, Simon; Guez, Arthur; Lockhart, Edward; Hassabis, Demis; Graepel, Thore; Lillicrap, Timothy; Silver, David (December 2020)."Mastering Atari, Go, chess and shogi by planning with a learned model" .Nature .588 (7839). Nature Publishing Group: 604–609.arXiv :1911.08265 .Bibcode :2020Natur.588..604S .doi :10.1038/s41586-020-03051-4 .ISSN1476-4687 .PMID33361790 .
  21. Hafner, Danijar; Pasukonis, Jurgis; Ba, Jimmy; Lillicrap, Timothy (April 2025)."Mastering diverse control tasks through world models" .Nature .640 (8059). Nature Publishing Group: 647–653.Bibcode :2025Natur.640..647H .doi :10.1038/s41586-025-08744-2 .ISSN1476-4687 .PMC12003158 .PMID40175544 .
  22. Hafner, Danijar; Lillicrap, Timothy; Fischer, Ian; Villegas, Ruben; Ha, David; Lee, Honglak; Davidson, James (2019-05-24)."Learning Latent Dynamics for Planning from Pixels" .Proceedings of the 36th International Conference on Machine Learning . PMLR: 2555–2565.ISSN2640-3498 .
  23. Hafner, Danijar; Lillicrap, Timothy; Fischer, Ian; Villegas, Ruben; Ha, David; Lee, Honglak; Davidson, James (2019-05-24)."Learning Latent Dynamics for Planning from Pixels" .Proceedings of the 36th International Conference on Machine Learning . PMLR: 2555–2565.ISSN2640-3498 .
  24. Yin, Tenny; Mei, Zhiting; Zheng, Zhonghe; Yamane, Miyu; Wang, David; Sceats, Jade; Bateman, Samuel M.; Zha, Lihan; Badithela, Apurva; Shorinwa, Ola; Majumdar, Anirudha."PlayWorld: Learning Robot World Models from Autonomous Play" .arxiv.org . Retrieved 2026-03-25.
  25. Weldon, Marcus (April 2, 2025)."AI 'Godfather' Yann LeCun: LLMs Are Nearing the End, but Better AI Is Coming" .Newsweek . Retrieved June 30, 2026.
  26. "Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning" . Meta AI. June 11, 2025. Retrieved March 26, 2026.
  27. Bordes, Florian; Garrido, Quentin; Kao, Justine T.; Williams, Adina; Rabbat, Michael; Dupoux, Emmanuel (2025),IntPhys 2: Benchmarking Intuitive Physics Understanding in Complex Synthetic Environments ,arXiv :2506.09849
  28. Krojer, Ben (2025). "A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs".arXiv :2506.09987 [cs.CV ].
  29. Goyal, Raghav (2017). "The "something something" video database for learning and evaluating visual common sense".Proceedings of the IEEE International Conference on Computer Vision (ICCV) .arXiv :1706.04261 .
30. [↑](#cite_ref-30)["Genie 3: A new frontier for world models"](https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/) . Google DeepMind. August 5, 2025. Retrieved March 26, 2026.
31. [↑](#cite_ref-31)["EPIC-KITCHENS-100 Dataset"](https://epic-kitchens.github.io/) . Retrieved March 26, 2026.
[
[edit](</w/index.php?title=World_model_(artificial_intelligence)&action=edit§ion=9>)]
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @jürgen schmidhuber 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/world-model-artifici…] indexed:0 read:11min 2026-09-08 ·