arXiv:2609.16233v1 Announce Type: new Abstract: Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat objects in isolation while ignoring real-world hierarchical organization (scenes, rooms, functional areas, object groups). Third, evaluation tasks focus narrowly on basic recognition rather than multi-step spatial reasoning. In this context, we introduce SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects. These annotations are produced through a human-in-the-loop pipeline combining vision-language models with roughly 1,500 human-hours of iterative refinement and verification, producing over 183K annotated nodes with textual descriptions and 3D bounding boxes. Building on this representation, we define three evaluation tasks: Existence-Based Questions probing object attributes, Spatial Intelligence Questions covering counting, size comparison, distance, and directional relations, and Grounded Question-Reasoning-Answer (QRA) triplets requiring multi-step reasoning across semantic levels. Experiments with state-of-the-art vision-language models show that while models perform well on basic recognition tasks (e.g., up to 85% accuracy for detection), performance drops substantially on hierarchical and compositional reasoning (e.g., down to 60% for counting), revealing limitations not captured by existing benchmarks. SceneBench provides a realistic testbed for developing and evaluating models capable of fine-grained spatial reasoning in photorealistic 3D environments.
SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes
Researchers introduced SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and annotated with over 183K hierarchical nodes, built through a human-in-the-loop pipeline using roughly 1,500 human-hours of refinement. Experiments with state-of-the-art vision-language models showed strong basic recognition (up to 85% accuracy for detection) but sharp drops on hierarchical and compositional reasoning, falling to 60% for counting. SceneBench defines three evaluation tasks — Existence-Based Questions, Spatial Intelligence Questions, and Grounded Question-Reasoning-Answer triplets — to test fine-grained spatial reasoning in photorealistic 3D environments.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.