cd /news/artificial-intelligence/scaling-representation-diversity-mod… · home topics artificial-intelligence article
[ARTICLE · art-96259] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

Researchers propose a data-model co-design framework to scale representation diversity in Referring Expression Comprehension, introducing the Modulated Attention-Contrastive Head (mACH) and a text-conditioned JEPA auxiliary stream, alongside the Objects365-Caption dataset. The single-checkpoint framework achieves competitive performance on standard REC benchmarks and strong cross-dataset generalization without benchmark-specific adaptation, addressing representation degeneration as a key obstacle to unified open-vocabulary grounding.

read1 min views1 publishedAug 14, 2026

arXiv:2608.12748v1 Announce Type: new Abstract: Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective of unified open-vocabulary grounding and identify representation degeneration as a key obstacle to scaling a single generalist model. To preserve representation diversity, we propose a holistic data-model co-design framework. Architecturally, we introduce the Modulated Attention-Contrastive Head (mACH) for efficient token-level vision-language alignment and a text-conditioned JEPA auxiliary stream that provides complementary gradient support to preserve alignment-active representations without inference overhead. On the data side, we introduce Objects365-Caption, enriching Objects365 with context-aware referring expressions for large-scale language supervision. We further provide a theoretical analysis showing that complementary gradient subspaces preserve alignment capacity and thereby scale representation diversity. Extensive experiments demonstrate that our single-checkpoint framework achieves highly competitive performance on standard REC benchmarks while exhibiting strong generalization across heterogeneous grounding datasets without benchmark-specific adaptation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scaling-representati…] indexed:0 read:1min 2026-08-14 ·