04:00
2026-08-04
arxiv.org
artificial-intelligence
What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
Researchers propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts vision-language model (VLM) benchmark accuracy from textual capability scores, trained โฆ