{"slug": "lora-enhanced-contrastive-learning-with-sas-vision-transformers", "title": "LoRA Enhanced Contrastive Learning with SAS Vision Transformers", "summary": "A three-stage parameter-efficient framework adapting DINOv3 Vision Transformer models to synthetic aperture sonar (SAS) automatic target recognition found that Low-Rank Adaptation (LoRA) alone raised area under the precision-recall curve (AUPRC) from 0.300 to 0.679 +/- 0.027 on a frozen backbone, according to an arXiv paper (2609.21061v1). Rank 4 achieved that result while training only 0.26 percent of weights, and neither refinement stage exceeded its matched control: hard-negative mining changed AUPRC by -0.0045 +/- 0.0119 versus an equal-size random curriculum, while Supervised Contrastive Learning changed AUPRC by +0.0002 +/- 0.0096 versus the preceding stage. The authors conclude one efficient adaptation stage is sufficient and stacked refinement is not.", "body_md": "arXiv:2609.21061v1 Announce Type: new \nAbstract: Automatic target recognition (ATR) with synthetic aperture sonar (SAS) supports advanced naval capabilities, but deep learning is constrained by scarce target imagery, background clutter, and human-in-the-loop assessment. We adapt DINOv3 Vision Transformer (ViT) models to underwater SAS ATR using a three-stage parameter-efficient framework. Stage 1 uses Low-Rank Adaptation (LoRA) while freezing the ViT backbone, bridging the gap between natural-image pretraining and underwater acoustic propagation. Stage 2 uses hard-negative mining to strengthen the decision boundary against acoustic mimics, including rocks and sediment formations resembling man-made targets. Stage 3 uses Supervised Contrastive Learning (SupCon) to separate target and clutter representations. We evaluate at-sea SAS data using a mission-level geographic split, compare all arms at 85 percent test recall, and repeat each comparison over three random seeds. LoRA accounts for the primary effect, increasing area under the precision-recall curve (AUPRC) from 0.300 to 0.679 +/- 0.027 using the same frozen backbone. Rank 4 achieves this result while training only 0.26 percent of weights. Neither refinement stage exceeds its matched control: hard-negative mining changes AUPRC by -0.0045 +/- 0.0119 versus an equal-size random curriculum, and SupCon changes AUPRC by +0.0002 +/- 0.0096 versus the preceding stage. These null results indicate that mining occurred on data the encoder had already fit and that supervised stages had already imposed most target-clutter geometry. One efficient adaptation stage is sufficient; stacked refinement is not.", "url": "https://wpnews.pro/news/lora-enhanced-contrastive-learning-with-sas-vision-transformers", "canonical_source": "https://arxiv.org/abs/2609.21061", "published_at": "2026-09-21 04:00:00+00:00", "updated_at": "2026-09-21 04:26:05.337006+00:00", "lang": "en", "topics": ["computer-vision", "machine-learning", "ai-research", "neural-networks"], "entities": ["DINOv3", "Vision Transformer", "Low-Rank Adaptation", "Supervised Contrastive Learning", "arXiv", "synthetic aperture sonar"], "alternates": {"html": "https://wpnews.pro/news/lora-enhanced-contrastive-learning-with-sas-vision-transformers", "markdown": "https://wpnews.pro/news/lora-enhanced-contrastive-learning-with-sas-vision-transformers.md", "text": "https://wpnews.pro/news/lora-enhanced-contrastive-learning-with-sas-vision-transformers.txt", "jsonld": "https://wpnews.pro/news/lora-enhanced-contrastive-learning-with-sas-vision-transformers.jsonld"}}