04:00
2026-08-12
arxiv.org
artificial-intelligence
Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models
Researchers introduce Space Tokens, a lightweight, architecture-agnostic framework that equips vision-language models (VLMs) with explicit continuous spatial representations without additional inferenβ¦