{"slug": "scaling-representation-diversity-modulated-attention-and-reconstructive-for", "title": "Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding", "summary": "Researchers propose a data-model co-design framework to scale representation diversity in Referring Expression Comprehension, introducing the Modulated Attention-Contrastive Head (mACH) and a text-conditioned JEPA auxiliary stream, alongside the Objects365-Caption dataset. The single-checkpoint framework achieves competitive performance on standard REC benchmarks and strong cross-dataset generalization without benchmark-specific adaptation, addressing representation degeneration as a key obstacle to unified open-vocabulary grounding.", "body_md": "arXiv:2608.12748v1 Announce Type: new\nAbstract: Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective of unified open-vocabulary grounding and identify representation degeneration as a key obstacle to scaling a single generalist model. To preserve representation diversity, we propose a holistic data-model co-design framework. Architecturally, we introduce the Modulated Attention-Contrastive Head (mACH) for efficient token-level vision-language alignment and a text-conditioned JEPA auxiliary stream that provides complementary gradient support to preserve alignment-active representations without inference overhead. On the data side, we introduce Objects365-Caption, enriching Objects365 with context-aware referring expressions for large-scale language supervision. We further provide a theoretical analysis showing that complementary gradient subspaces preserve alignment capacity and thereby scale representation diversity. Extensive experiments demonstrate that our single-checkpoint framework achieves highly competitive performance on standard REC benchmarks while exhibiting strong generalization across heterogeneous grounding datasets without benchmark-specific adaptation.", "url": "https://wpnews.pro/news/scaling-representation-diversity-modulated-attention-and-reconstructive-for", "canonical_source": "https://arxiv.org/abs/2608.12748", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:06:58.084967+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "natural-language-processing", "ai-research"], "entities": ["arXiv", "Objects365-Caption", "Modulated Attention-Contrastive Head (mACH)", "JEPA"], "alternates": {"html": "https://wpnews.pro/news/scaling-representation-diversity-modulated-attention-and-reconstructive-for", "markdown": "https://wpnews.pro/news/scaling-representation-diversity-modulated-attention-and-reconstructive-for.md", "text": "https://wpnews.pro/news/scaling-representation-diversity-modulated-attention-and-reconstructive-for.txt", "jsonld": "https://wpnews.pro/news/scaling-representation-diversity-modulated-attention-and-reconstructive-for.jsonld"}}