arXiv:2609.36219v1 Announce Type: new Abstract: Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.
LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
A new arXiv paper (arXiv:2609.36219v1) introduces LeRF, a framework that trains Vision-Language Models to build and use explicit reference coordinate frames for perspective-taking reasoning, addressing the tendency of VLMs to default to the camera viewpoint when a query requires another viewpoint. LeRF first decides whether a coordinate frame is needed, then grounds the reference entity and predicts the frame's origin and entity-centered reference frame, with a lightweight renderer overlaying the frame onto the image so reasoning proceeds over visual cues without external perception models or explicit 3D reconstruction. Training combines supervised fine-tuning for selective tool invocation and frame prediction with reinforcement learning on spatial VQA pairs, and the authors report LeRF consistently improves over its backbone and performs strongly against existing open-source methods across diverse perspective-taking benchmarks, with gains in reference-frame grounding and orientation estimation.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.