cd /news/computer-vision/scenebench-a-hierarchical-benchmark-… · home topics computer-vision article
[ARTICLE · art-131017] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

Researchers introduced SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and annotated with over 183K hierarchical nodes, built through a human-in-the-loop pipeline using roughly 1,500 human-hours of refinement. Experiments with state-of-the-art vision-language models showed strong basic recognition (up to 85% accuracy for detection) but sharp drops on hierarchical and compositional reasoning, falling to 60% for counting. SceneBench defines three evaluation tasks — Existence-Based Questions, Spatial Intelligence Questions, and Grounded Question-Reasoning-Answer triplets — to test fine-grained spatial reasoning in photorealistic 3D environments.

by read1 min views1 publishedSep 16, 2026

arXiv:2609.16233v1 Announce Type: new Abstract: Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat objects in isolation while ignoring real-world hierarchical organization (scenes, rooms, functional areas, object groups). Third, evaluation tasks focus narrowly on basic recognition rather than multi-step spatial reasoning. In this context, we introduce SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects. These annotations are produced through a human-in-the-loop pipeline combining vision-language models with roughly 1,500 human-hours of iterative refinement and verification, producing over 183K annotated nodes with textual descriptions and 3D bounding boxes. Building on this representation, we define three evaluation tasks: Existence-Based Questions probing object attributes, Spatial Intelligence Questions covering counting, size comparison, distance, and directional relations, and Grounded Question-Reasoning-Answer (QRA) triplets requiring multi-step reasoning across semantic levels. Experiments with state-of-the-art vision-language models show that while models perform well on basic recognition tasks (e.g., up to 85% accuracy for detection), performance drops substantially on hierarchical and compositional reasoning (e.g., down to 60% for counting), revealing limitations not captured by existing benchmarks. SceneBench provides a realistic testbed for developing and evaluating models capable of fine-grained spatial reasoning in photorealistic 3D environments.

── more in #computer-vision 4 stories · sorted by recency
── more on @scenebench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scenebench-a-hierarc…] indexed:0 read:1min 2026-09-16 ·