# VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning

> Source: <https://arxiv.org/abs/2610.10782>
> Published: 2026-10-09 04:00:00+00:00

arXiv:2610.10782v1 Announce Type: new 
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
  but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many
  become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the
  visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor
  and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as
  scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty
  is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with
  actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and
  visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest
  self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized
  RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution,
  VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
