CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs A new method called CoVeR (Coverage-Based Token Pruning) reduces the number of visual tokens processed by 2D vision-language models (VLMs) when reasoning about 3D scenes from multi-view images, addressing the high computational cost of thousands of redundant tokens. The approach, detailed in a research paper, aims to improve efficiency of 3D reasoning without sacrificing accuracy. Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two fam