cd /news/computer-vision/don-t-just-look-intervene-perturbati… · home topics computer-vision article
[ARTICLE · art-129851] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

Don't Just Look, Intervene: Perturbation Based Region Labeling for VQA Images

Researchers introduced Counterfactual Search for Grounding Regions (CSGR), a scalable pipeline that labels answer-critical image regions in visual question answering (VQA) data by perturbing candidate regions and measuring their effect on a model's answer distribution. CSGR annotations produced the most consistent gains over Cross Entropy-only finetuning across competing automatic region-labeling mechanisms in both in-domain and out-of-domain evaluations, according to the arXiv paper 2609.13228v1. The authors tested the labels through three existing grounding-aware training routines: attention steering, Visual CoT finetuning, and latent visual reasoning.

by read1 min views1 publishedSep 15, 2026

arXiv:2609.13228v1 Announce Type: new Abstract: Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-causal visual evidence as an annotation target, defined as the set of image regions whose counterfactual intervention changes a model's answer distribution for a given image-question pair. Based on this principle, we introduce Counterfactual Search for Grounding Regions (CSGR). CSGR is a scalable pipeline that proposes candidate regions, perturbs them, measures their effect on answer sensitivity, and aggregates this evidence across multiple judges to approximate answer-critical regions in VQA data. To assess whether CSGR annotations contain a useful supervision signal, we plug them into three existing grounding-aware training routines: attention steering, Visual CoTfinetuning, and latent visual reasoning. These experiments test whether the proposed annotation scheme can provide a useful supervision signal across multiple ways of consuming region labels, rather than introducing a new way of using them. Across competing automatic region-labeling mechanisms, CSGR annotations provide the most consistent gains over Cross Entropy-only finetuning in both in-domain and out-of-domain evaluations, indicating that the proposed labeling scheme captures useful region-level information.

── more in #computer-vision 4 stories · sorted by recency
── more on @counterfactual search for grounding regions 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/don-t-just-look-inte…] indexed:0 read:1min 2026-09-15 ·