cd /news/artificial-intelligence/what-transfers-from-text-to-vision-c… · home topics artificial-intelligence article
[ARTICLE · art-85559] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

Researchers propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts vision-language model (VLM) benchmark accuracy from textual capability scores, trained on over 150 VLMs across 34 LLMs and 7 model families. The law accurately extrapolates transfer rates from 8B to 72B-scale backbones, predicts full training trajectories, and generalizes to held-out families, while revealing that base LLMs outperform instruction-tuned counterparts as VLM backbones. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00013v1 Announce Type: new Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score $S$ extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of $S$, with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency. To fit and validate the framework, we train over 150 VLMs on 34 LLMs spanning 7 model families under a strictly controlled recipe. Evaluations on more than 200 textual and 50 multimodal benchmarks show that the law accurately extrapolates transfer rate from models up to 8B parameters to 72B-scale backbones, predicts full VLM training trajectories with high fidelity, and generalizes to entirely held-out model families. Beyond the scaling law, our analysis surfaces actionable insights: certain textual benchmarks negatively correlate with multimodal performance, exposing latent benchmark-gaming behavior; base LLMs outperform instruction-tuned counterparts as VLM backbones due to higher absorption rates and lower data-scaling decay; and different model families occupy distinct positions in the transfer--absorption space. The framework turns backbone selection from costly empirical sweeps into a principled, quantitative decision. Code and data are available at https://github.com/wangq-dev/CDMScaling.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @capability-driven multimodal scaling law 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-transfers-from-…] indexed:0 read:1min 2026-08-04 ·