{"slug": "statesight-benchmarking-latent-spatial-state-reconstruction-in-vision-language", "title": "StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models", "summary": "Researchers introduced StateSight, a procedurally generated benchmark for evaluating vision-language models' latent spatial-state reconstruction, and found that OpenAI's GPT-5.5 (API model identifier gpt-5.5) achieved 59.3%, 33.3%, and 28.3% accuracy on cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting, respectively, while Claude Sonnet 5 scored 53.3%, 18.7%, and 7.3%. A 30-participant human baseline outperformed both models on all tasks with mean accuracies of 80.8%, 68.8%, and 64.3%, and the study introduced StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states, showing that format-valid responses can mask failures in spatial reasoning.", "body_md": "arXiv:2608.20414v1 Announce Type: new\nAbstract: Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain knowledge, linguistic priors, and reasoning in the same evaluation. We introduce StateSight, a procedurally generated benchmark for cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Each task family contains 300 single-image prompts with deterministic oracle labels and exact-match scoring. OpenAI GPT-5.5, using the API model identifier gpt-5.5, achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%. All final direct runs had zero format errors. A 30-participant human baseline on 60 items exceeded both models on every task, with mean accuracies of 80.8%, 68.8%, and 64.3%. Visible-derivation analysis identified recurring errors in image-state reconstruction and reasoning procedure. We also introduce StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states. The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.", "url": "https://wpnews.pro/news/statesight-benchmarking-latent-spatial-state-reconstruction-in-vision-language", "canonical_source": "https://arxiv.org/abs/2608.20414", "published_at": "2026-08-24 04:00:00+00:00", "updated_at": "2026-08-24 04:13:24.861835+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "ai-research"], "entities": ["OpenAI", "GPT-5.5", "Claude Sonnet 5", "StateSight", "StateSight-Steps"], "alternates": {"html": "https://wpnews.pro/news/statesight-benchmarking-latent-spatial-state-reconstruction-in-vision-language", "markdown": "https://wpnews.pro/news/statesight-benchmarking-latent-spatial-state-reconstruction-in-vision-language.md", "text": "https://wpnews.pro/news/statesight-benchmarking-latent-spatial-state-reconstruction-in-vision-language.txt", "jsonld": "https://wpnews.pro/news/statesight-benchmarking-latent-spatial-state-reconstruction-in-vision-language.jsonld"}}