cd /news/machine-learning/cvsd-reg-cross-modal-visual-semantic… · home topics machine-learning article
[ARTICLE · art-105490] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration

Researchers propose CVSD-Reg, a global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations, achieving strict success rates of 97.7% on KITTI, 99.0% on nuScenes, and 99.3% on HeLiPR, including 97.3% on sparse 16-beam Velodyne scans. The method outperforms state-of-the-art geometric registration methods by up to 44.0 percentage points without requiring camera inputs or post-hoc ICP refinement.

read1 min views4 publishedAug 21, 2026

arXiv:2608.19536v1 Announce Type: new Abstract: Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existing methods sensitive to variations in point density, scan pattern, viewpoint, and sensor characteristics. We propose CVSD-Reg, a robust global LiDAR registration framework that distills visual semantic priors from a vision foundation model into LiDAR representations. In Stage 1, a Point Transformer V3 student learns from a frozen DINOv2 teacher through contrastive distillation and spherical-manifold alignment, which preserves the hyperspherical geometry of the teacher embedding space. Self-supervised InfoNCE consistency and soft $\mathrm{SE}(3)$ invariance further encourage viewpoint-robust descriptors. In Stage 2, the distilled representation is adapted to registration through correspondence learning, density-aware point-dropout augmentation, and end-to-end pose optimization. With a single checkpoint, CVSD-Reg generalizes to both single-sensor and zero-shot cross-sensor scenarios without sensor-specific adaptation and remains entirely camera-free at inference. On KITTI, nuScenes, and HeLiPR, CVSD-Reg achieves strict success rate (SR@0.5,m/$1^\circ$) of 97.7$%$, 99.0$%$, and 99.3$%$, respectively, including 97.3$%$ on sparse 16-beam Velodyne scans. It outperforms state-of-the-art geometric registration methods by up to 44.0 percentage points without requiring camera inputs or post-hoc ICP refinement.

── more in #machine-learning 4 stories · sorted by recency
── more on @cvsd-reg 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cvsd-reg-cross-modal…] indexed:0 read:1min 2026-08-21 ·