StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models Researchers introduced StateSight, a procedurally generated benchmark for evaluating vision-language models' latent spatial-state reconstruction, and found that OpenAI's GPT-5.5 (API model identifier gpt-5.5) achieved 59.3%, 33.3%, and 28.3% accuracy on cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting, respectively, while Claude Sonnet 5 scored 53.3%, 18.7%, and 7.3%. A 30-participant human baseline outperformed both models on all tasks with mean accuracies of 80.8%, 68.8%, and 64.3%, and the study introduced StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states, showing that format-valid responses can mask failures in spatial reasoning. arXiv:2608.20414v1 Announce Type: new Abstract: Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain knowledge, linguistic priors, and reasoning in the same evaluation. We introduce StateSight, a procedurally generated benchmark for cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Each task family contains 300 single-image prompts with deterministic oracle labels and exact-match scoring. OpenAI GPT-5.5, using the API model identifier gpt-5.5, achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%. All final direct runs had zero format errors. A 30-participant human baseline on 60 items exceeded both models on every task, with mean accuracies of 80.8%, 68.8%, and 64.3%. Visible-derivation analysis identified recurring errors in image-state reconstruction and reasoning procedure. We also introduce StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states. The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.