{"slug": "vico-visual-environments-co-evolving-for-vision-language-model-reasoning", "title": "VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning", "summary": "Researchers proposed VICO, a co-evolutionary framework that jointly trains a vision-language model actor and an Environment-as-Rewriter that edits verifiable image-side structures such as scene graphs, chart tables, and protected region masks to generate label-valid training samples calibrated to the actor's current ability. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improved over its base model by up to +5.0% on out-of-domain tasks, surpassed the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stayed comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. The work, published as arXiv:2610.10782v1, argues VLM post-training should evolve the visual environment alongside the actor rather than relying on a static training environment.", "body_md": "arXiv:2610.10782v1 Announce Type: new \nAbstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),\n  but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many\n  become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the\n  visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor\n  and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as\n  scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty\n  is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with\n  actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and\n  visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest\n  self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized\n  RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution,\n  VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.", "url": "https://wpnews.pro/news/vico-visual-environments-co-evolving-for-vision-language-model-reasoning", "canonical_source": "https://arxiv.org/abs/2610.10782", "published_at": "2026-10-09 04:00:00+00:00", "updated_at": "2026-10-09 04:17:02.062913+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "computer-vision", "large-language-models"], "entities": ["VICO", "EnvRewriter", "VICO-8B", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/vico-visual-environments-co-evolving-for-vision-language-model-reasoning", "markdown": "https://wpnews.pro/news/vico-visual-environments-co-evolving-for-vision-language-model-reasoning.md", "text": "https://wpnews.pro/news/vico-visual-environments-co-evolving-for-vision-language-model-reasoning.txt", "jsonld": "https://wpnews.pro/news/vico-visual-environments-co-evolving-for-vision-language-model-reasoning.jsonld"}}