arXiv:2609.18057v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-policy RLVR algorithms face a critical optimization bottleneck in preserving and reinforcing visually grounded reasoning behaviors: valuable visually-grounded reasoning trajectories are discarded after a single update, while uniform token advantage allocation prevents the model from reinforcing critical perception or reasoning steps. To bridge this gap, we propose PIVOT, a dual-level learning framework that anchors policy optimization around informative visual reasoning signals. Specifically, PIVOT introduces a self-calibrated experience replay mechanism, which selectively collects and replays visually-grounded historical experiences as stable reference anchors for policy optimization. Building upon this, we further design a vision-guided advantage allocation mechanism to allocate additional vision-aware advantages to tokens based on their local visual support and impact on downstream reasoning. Extensive experiments across diverse benchmarks demonstrate that PIVOT achieves highly competitive performance in enhancing the multimodal reasoning capabilities of LVLMs.
Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning
Researchers introduced PIVOT, a dual-level learning framework that anchors policy optimization around informative visual reasoning signals to improve multimodal reasoning in large vision-language models (LVLMs). PIVOT combines a self-calibrated experience replay mechanism, which collects and replays visually-grounded historical experiences as stable reference anchors, with a vision-guided advantage allocation mechanism that assigns additional vision-aware advantages to tokens based on local visual support and downstream reasoning impact. The authors report that PIVOT achieves highly competitive performance across diverse benchmarks for enhancing LVLM multimodal reasoning, addressing the bottleneck in standard on-policy RLVR algorithms that discard visually-grounded trajectories after a single update.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.