cd /news/artificial-intelligence/spatial-iq-deconstructing-spatial-in… · home topics artificial-intelligence article
[ARTICLE · art-76270] src=research.nvidia.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

A new diagnostic framework called Spatial-IQ, introduced by researchers using NVIDIA Isaac Sim, decomposes object counting in stacked 3D structures into 9 perceptual and cognitive sub-tasks to evaluate spatial reasoning in multimodal large language models (MLLMs). Testing on roughly 80,000 procedurally generated structures reveals that top-performing models often succeed at the target task without mastering lower-level sub-tasks, exposing shortcut behaviors. Training with chain-of-thought supervision over hierarchical sub-tasks and reinforcement learning with verifiable rewards significantly improves spatial consistency and target-task accuracy.

read1 min views1 publishedJul 28, 2026

Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as black boxes, limiting their ability to identify the underlying causes of lower performance: when a model fails a spatial reasoning task, it remains difficult to ascertain whether the hurdle is perceptual, such as recognizing object boundaries, or cognitive, such as reasoning about occlusion to infer hidden geometry. We introduce Spatial-IQ, a hierarchical diagnostic framework that decomposes object counting in stacked 3D structures into 9 perceptual and cognitive sub-tasks organized by the developmental stages of human spatial cognition, with mental rotation as an additional target probe. Using NVIDIA Isaac Sim, we procedurally generated a diverse dataset of roughly 80,000 stacked 3D structures with per-task ground truth. We evaluate models across three output formats (free-response text, multiple-choice images, and image editing) alongside a human baseline. The Spatial-IQ framework shows that top-performing models often succeed at the target task (object counting) without succeeding on the lower-level sub-tasks intended to support it, and that models differ in how much of these hierarchical chains they preserve, often revealing shortcut behavior that raw target-task accuracy alone would obscure. Finally, we demonstrate that training models with chain-of-thought (CoT) supervision over our hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition as both a diagnostic tool and a training signal.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @spatial-iq 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/spatial-iq-deconstru…] indexed:0 read:1min 2026-07-28 ·