cd /news/computer-vision/geounipr-a-geometry-consistent-unifi… · home topics computer-vision article
[ARTICLE · art-94722] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

GeoUniPR: A Geometry-Consistent Unified Framework for Cross-Modal Place Recognition

Researchers propose GeoUniPR, a geometry-consistent unified framework for cross-modal place recognition that projects LiDAR point clouds into camera perspective to create depth image views, achieving state-of-the-art performance on KITTI and KITTI-360 datasets. The framework uses two ViT-based encoders with identical architectures and introduces Spatially-Consistent InfoNCE to suppress distance-induced false negatives, enabling strong cross-dataset generalization without auxiliary alignment modules or full backbone fine-tuning.

read1 min views1 publishedAug 13, 2026

arXiv:2608.11263v1 Announce Type: new Abstract: Cross-modal place recognition (CMPR) aims to identify the same location across heterogeneous sensing modalities, such as vision and LiDAR. Existing methods commonly bridge the modality gap using complex alignment modules, multi-stage training, or full fine-tuning of pretrained backbones. In this work, we revisit CMPR from the perspective of geometric consistency and propose GeoUniPR, a unified and concise geometry-consistent framework. GeoUniPR reduces cross-modal discrepancy at the representation level by projecting LiDAR point clouds into the camera perspective to construct Geometry-Consistent depth image views (DIV), which establish direct RGB-LiDAR correspondence. We further augment DIV with native LiDAR cues, including intensity and surface-normal information, yielding a multi-channel geometric representation that improves structural consistency. Based on this representation, GeoUniPR learns a unified embedding space using two modality-specific ViT-based encoders with identical architectures, trained through parameter-efficient adaptation without auxiliary alignment modules, multi-stage training, or full backbone fine-tuning. In addition, we introduce Spatially-Consistent InfoNCE (SC-InfoNCE), a CMPR-specific contrastive objective that suppresses distance-induced false negatives under spatial continuity. Extensive experiments on KITTI and KITTI-360 demonstrate that GeoUniPR achieves state-of-the-art (SOTA) performance in both same-modal and cross-modal place recognition, with strong cross-dataset generalization.

── more in #computer-vision 4 stories · sorted by recency
── more on @geounipr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/geounipr-a-geometry-…] indexed:0 read:1min 2026-08-13 ·