A look at spatial intelligence and world models Spatial intelligence, an AI capability that enables models to reason about three-dimensional space, is emerging as the next frontier beyond large language models, according to Dr. Fei-Fei Li, Stanford professor and AI pioneer. These models can generate 3D scenes and connect natural language with physical environments, with applications in robotics, manufacturing, and infrastructure maintenance, where the American Society of Civil Engineers estimates a $9.1 trillion investment is needed from 2024 through 2033 for repairs. It’s been several years since generative AI https://www.infoworld.com/article/2338115/what-is-generative-ai-artificial-intelligence-that-creates.html and large language models https://www.understandingai.org/p/large-language-models-explained-with LLMs took the world by storm. LLMs surpassed earlier natural-language systems at generating text, while diffusion models https://www.technologyreview.com/2025/09/12/1123562/how-do-ai-models-generate-videos/ enabled generating images, music, and videos. These generative AI models work well in the digital world, but on their own, they have limited capabilities to comprehend the three-dimensional physical world and other spaces. This includes the objects occupying an area, how they relate to each other, tracking movement, and answering complex questions requiring an understanding of dimensions, distances, motion, and collisions. Spatial intelligence is an AI capability that allows models to reason about three-dimensional space. These models can generate 3D scenes of the world and other spaces. This content can then be displayed through traditional renderers, game engines, or AR/VR systems that use spatial computing https://builtin.com/hardware/spatial-computing techniques. But it’s the spatial intelligence model’s ability to connect natural language with 3D models that has the most applications in robotics, manufacturing, construction, and other physical environments. Dr. Fei-Fei Li, often called the godmother of AI https://profiles.stanford.edu/fei-fei-li , published a manifesto on how spatial intelligence is AI’s next frontier https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence , contrasting it with LLMs. “While current state-of-the-art AI can excel at reading, writing, research, and pattern recognition in data, these same models bear fundamental limitations when representing or interacting with the physical world,” wrote Dr. Li. “Our view of the world is holistic—not just what we’re looking at, but how everything relates spatially, what it means, and why it matters. Understanding this through imagination, reasoning, creation, and interaction—not just descriptions—is the power of spatial intelligence.” The concept of spatial intelligence isn’t new and was described in Howard Gardner’s book, Frames of Mind , in 1983. Recent breakthroughs, including the launch of It’s important to develop AI literacy https://drive.starcio.com/2026/02/ai-literacy-a-leadership-guide/ and understand the terminology and concepts related to the physical world and 3D AI technologies: “Spatial intelligence models go beyond pixels to understand the 3D structure of the world—how objects are positioned, how they move, and how they interact,” says David Fattal, founder and CTO at Leia https://immersity.ai/ . “This enables applications like more realistic video generation, spatial computing interfaces, and AI systems that can reason about physical environments. As real-world 3D data becomes more available, these models will become foundational to the next generation of visual AI.” To better understand spatial intelligence, let’s consider physical infrastructure such as bridges and buildings. The American Society of Civil Engineers estimates a $9.1 trillion investment https://www.enr.com/articles/62214-infrastructure-gains-in-new-asce-report-cardbut-progress-hinges-on-post-2026-funds is needed from 2024 through 2033 to achieve a state of good repair. When maintenance and monitoring lag, it can lead to major failures such as the 2022 collapse of the Fern Hollow Bridge in Pittsburgh https://www.ntsb.gov/news/press-releases/Pages/NR20240221.aspx . Spatial intelligence and the development of digital twins may help identify issues earlier and prioritize where investments are needed. “Spatial intelligence models serve as the 4D digital blueprints for our built environment, allowing us to visualize and predict the complex interactions between aging assets and the shifting ground beneath them,” says Patrick Cozzi, chief platform officer at Bentley Systems https://www.bentley.com/ . “By synthesizing disparate geospatial data into a living digital twin, these models provide the foresight necessary to mitigate the hidden risks of structural fatigue and subsurface instability.” There’s a significant challenge in bridge health monitoring https://www.mdpi.com/1424-8220/21/13/4336 and transitioning from manual, infrequent structural inspections to leveraging sensors, digital twins, and spatial intelligence. Cozzi adds, “This integration of continuous field data moves beyond static documentation, empowering agencies to evolve from reactive repairs to proactive, resilient asset management that safeguards the long-term integrity of our most critical public systems.” Bridges are largely static, but the real world is increasingly being occupied by autonomous systems such as self-driving cars, robots, and drones. And where there are moving systems, there is a risk of collisions. “Spatial intelligence models are AI systems that reason about the physical world by combining vision, sensor data, and contextual cues to understand space, motion, and object relationships,” says Sudeep George, CTO at iMerit https://imerit.net/ . “The value of spatial intelligence models lies not just in perceiving an environment, but in enabling machines to act within it safely and in real time. That is especially important in robotics and autonomous systems, where decisions must be made in complex, multimodal, fast-changing settings.” To see one example, this tutorial for simulating robotic environments https://developer.nvidia.com/blog/simulate-robotic-environments-faster-with-nvidia-isaac-sim-and-world-labs-marble combines Nvidia Isaac Sim https://developer.nvidia.com/isaac/sim?size=n 6 n&sort-field=featured&sort-direction=desc , an open source robotics reference framework, with spatial intelligence in Marble from World Labs. Today’s collision detection systems, such as those used in autonomous vehicles https://arxiv.org/html/2508.20892v1 , typically rely on modules for sensing, perception, planning, and control. Spatial intelligence models may offer improvements by assessing the collision risks of unidentified objects or by tracking objects that move out of sensor view. For example, Waymo’s World Model https://waymo.com/blog/2026/02/the-waymo-world-model-a-new-frontier-for-autonomous-driving-simulation/ , built on Genie 3, is a simulator that generates complex weather conditions and other critical safety events. For an out-of-this-world example, James Urquhart, field CTO and technology evangelist at Kamiwaza https://www.kamiwaza.ai , has delivered several examples of spatial intelligence applications, including one for satellite collision detection and conflict analysis. Urquhart says, “Models that specialize in these types of data sets, as well as the physics and geography of the real world, enable faster and more accurate decision-making for tasks that depend on them.” Recent spatial intelligence announcements include creating 3D worlds from image or text prompts with Marble https://www.worldlabs.ai/blog/marble-world-model and simulating water physics, lighting, weather, and animal behavior with Genie 3 https://wavespeed.ai/blog/posts/google-deepmind-genie-3-world-model-2026/ . But difficulties remain in bringing spatial intelligence to physical-world use cases. “Spatial intelligence and world models are laying the groundwork for future AI agents that will be able to interact in and with our physical world,” says Jason Corso, cofounder and chief scientist at Voxel51 https://voxel51.com/ . “These models are significantly more challenging to develop and test, largely because the data underlying their development is complex, and it’s hard to handle all of the combinatorics involved in the physical world.” In addition to learning the models and prototyping with them, development and data leaders need to review the data assets that will feed spatial intelligence models. “Spatial intelligence models translate location signals into a structured understanding of the real world, but they’re only as reliable as the data beneath them,” says Dan Adams, executive vice president and general manager of Enrich at Precisely https://www.precisely.com/ . “The real unlock isn’t the model—it’s the reference layer with persistent identifiers, confidence metadata, and source lineage that lets AI reason about places, not just match strings.” Even once applications are developed, there will be infrastructure challenges in deploying them at the edge. Ali Kayyam, principal research scientist at BrainChip https://brainchip.com/ , says, “The key to unlocking spatial intelligence at scale is having the low-power, event-driven hardware that can run it at the sensor in real time where it matters most.” My suggestions for developers looking to get hands-on with spatial intelligence and world models: