cd /news/artificial-intelligence/vico-visual-environments-co-evolving… · home › topics › artificial-intelligence › article
[ARTICLE · art-148008] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

Researchers proposed VICO, a co-evolutionary framework that jointly trains a vision-language model actor and an Environment-as-Rewriter that edits verifiable image-side structures such as scene graphs, chart tables, and protected region masks to generate label-valid training samples calibrated to the actor's current ability. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improved over its base model by up to +5.0% on out-of-domain tasks, surpassed the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stayed comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. The work, published as arXiv:2610.10782v1, argues VLM post-training should evolve the visual environment alongside the actor rather than relying on a static training environment.

by read1 min views4 publishedOct 9, 2026

arXiv:2610.10782v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @vico 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/vico-visual-environm…] indexed:0 read:1min 2026-10-09 · —