{"slug": "beyond-argmax-a-mechanistic-study-of-semantic-retention-in-frozen-foundation-for", "title": "Beyond Argmax: A Mechanistic Study of Semantic Retention in Frozen Foundation-Model Composition for Generalized Few-Shot 3D Segmentation", "summary": "A mechanistic study of frozen foundation-model composition for generalized few-shot 3D segmentation found that retaining the full distribution of semantic alternatives before fusion reaches 34.87 harmonic-mean IoU on 156 held-out ScanNet200 scenes, versus 28.47 for top-1 argmax, a gain of 6.40 points (95% CI [+5.24,+7.64]). The same-input intervention, which froze Dense RegionPLC and sparse cross-view SAM3 weights, masks, geometry, vocabularies, and fusion rules while varying only the number of retained alternatives via a matched top-k ladder, replicated on 50 ScanNet++ scenes at 26.50 versus 23.02 HM (+3.48, 95% CI [+1.64,+5.93]). The authors conclude that premature semantic collapse is a repeatable information bottleneck in heterogeneous frozen-model composition, with full-distribution HM stable across sparse-source weights 0.3–0.7 and a GroundingDINO–SAM2.1 source-replacement diagnostic showing monotonic HM growth from 14.77 to 18.75 under full retention.", "body_md": "arXiv:2609.12099v1 Announce Type: new \nAbstract: Classical classifier-combination work distinguishes score-level fusion from hard decision-level voting. We revisit this distinction where independently pretrained, frozen foundation models are composed at inference time for generalized few-shot 3D segmentation. We ask: how much useful semantic information is lost when heterogeneous sources are collapsed to a single class before they can interact?\n  We answer with a same-input semantic-retention intervention. Dense RegionPLC and sparse cross-view SAM3 evidence, model weights, masks, geometry, vocabularies, and fusion rules are frozen; only the number of semantic alternatives retained before interaction is varied via a matched top-k ladder. On 156 held-out ScanNet200 scenes, top-1 reaches 28.47 harmonic-mean (HM) IoU while full distribution fusion reaches 34.87 HM (+6.40, 95% CI [+5.24,+7.64]). The pattern replicates on 50 ScanNet++ scenes: 23.02 vs. 26.50 HM (+3.48, 95% CI [+1.64,+5.93]).\n  The conclusion is robust: full-distribution HM is stable across sparse-source weights 0.3--0.7; alternative operators (max, geometric pooling) also outperform top-1; and a GroundingDINO--SAM2.1 source-replacement diagnostic shows monotonic HM increase from 14.77 to 18.75 with full retention. Calibration diagnostics reveal opposite miscalibration of the two sources, yet correcting calibration does not eliminate the retention advantage.\n  Across datasets and source stacks, most information is recovered by retaining a compact set of plausible alternatives. The contribution is a controlled diagnosis of premature semantic collapse as a repeatable information bottleneck in heterogeneous frozen-model composition.", "url": "https://wpnews.pro/news/beyond-argmax-a-mechanistic-study-of-semantic-retention-in-frozen-foundation-for", "canonical_source": "https://arxiv.org/abs/2609.12099", "published_at": "2026-09-14 04:00:00+00:00", "updated_at": "2026-09-14 04:27:18.829451+00:00", "lang": "en", "topics": ["computer-vision", "ai-research", "machine-learning"], "entities": ["Dense RegionPLC", "SAM3", "ScanNet200", "ScanNet++", "GroundingDINO", "SAM2.1"], "alternates": {"html": "https://wpnews.pro/news/beyond-argmax-a-mechanistic-study-of-semantic-retention-in-frozen-foundation-for", "markdown": "https://wpnews.pro/news/beyond-argmax-a-mechanistic-study-of-semantic-retention-in-frozen-foundation-for.md", "text": "https://wpnews.pro/news/beyond-argmax-a-mechanistic-study-of-semantic-retention-in-frozen-foundation-for.txt", "jsonld": "https://wpnews.pro/news/beyond-argmax-a-mechanistic-study-of-semantic-retention-in-frozen-foundation-for.jsonld"}}