{"slug": "don-t-just-look-intervene-perturbation-based-region-labeling-for-vqa-images", "title": "Don't Just Look, Intervene: Perturbation Based Region Labeling for VQA Images", "summary": "Researchers introduced Counterfactual Search for Grounding Regions (CSGR), a scalable pipeline that labels answer-critical image regions in visual question answering (VQA) data by perturbing candidate regions and measuring their effect on a model's answer distribution. CSGR annotations produced the most consistent gains over Cross Entropy-only finetuning across competing automatic region-labeling mechanisms in both in-domain and out-of-domain evaluations, according to the arXiv paper 2609.13228v1. The authors tested the labels through three existing grounding-aware training routines: attention steering, Visual CoT finetuning, and latent visual reasoning.", "body_md": "arXiv:2609.13228v1 Announce Type: new \nAbstract: Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-causal visual evidence as an annotation target, defined as the set of image regions whose counterfactual intervention changes a model's answer distribution for a given image-question pair. Based on this principle, we introduce Counterfactual Search for Grounding Regions (CSGR). CSGR is a scalable pipeline that proposes candidate regions, perturbs them, measures their effect on answer sensitivity, and aggregates this evidence across multiple judges to approximate answer-critical regions in VQA data. To assess whether CSGR annotations contain a useful supervision signal, we plug them into three existing grounding-aware training routines: attention steering, Visual CoTfinetuning, and latent visual reasoning. These experiments test whether the proposed annotation scheme can provide a useful supervision signal across multiple ways of consuming region labels, rather than introducing a new way of using them. Across competing automatic region-labeling mechanisms, CSGR annotations provide the most consistent gains over Cross Entropy-only finetuning in both in-domain and out-of-domain evaluations, indicating that the proposed labeling scheme captures useful region-level information.", "url": "https://wpnews.pro/news/don-t-just-look-intervene-perturbation-based-region-labeling-for-vqa-images", "canonical_source": "https://arxiv.org/abs/2609.13228", "published_at": "2026-09-15 04:00:00+00:00", "updated_at": "2026-09-15 04:34:45.896786+00:00", "lang": "en", "topics": ["computer-vision", "natural-language-processing", "ai-research", "machine-learning"], "entities": ["Counterfactual Search for Grounding Regions", "CSGR", "arXiv", "Visual CoT"], "alternates": {"html": "https://wpnews.pro/news/don-t-just-look-intervene-perturbation-based-region-labeling-for-vqa-images", "markdown": "https://wpnews.pro/news/don-t-just-look-intervene-perturbation-based-region-labeling-for-vqa-images.md", "text": "https://wpnews.pro/news/don-t-just-look-intervene-perturbation-based-region-labeling-for-vqa-images.txt", "jsonld": "https://wpnews.pro/news/don-t-just-look-intervene-perturbation-based-region-labeling-for-vqa-images.jsonld"}}