cd /news/computer-vision/evidence-order-calibration-for-selec… · home topics computer-vision article
[ARTICLE · art-125434] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

Evidence-Order Calibration for Selective Visual Reasoning under Progressive Loss of Question-Critical Evidence

A study of vision-language model reliability found that a frozen Qwen2.5-VL-3B-Instruct model showed an evidence monotonicity violation rate (EMVR) of 0.436, with 92.0% of 176 accepted GQA-derived trajectories containing at least one adjacent violation as question-critical regions were progressively masked across 880 masking conditions. Full critical masking cut accuracy by 28.2 percentage points versus 0.6 points for equally sized non-critical masks, a paired difference of 27.6 points (95% CI [20.0, 34.7]). Adding evidence-order supervision to binary cross-entropy reduced masking EMVR from 0.330 to 0.303 (paired difference -0.027, 95% CI [-0.044, -0.010]) and reduced EMVR from 0.449 to 0.402 on held-out question IDs under unseen local Gaussian blur, though AUROC, Brier, and AURC differences between the two learned heads were statistically inconclusive and native confidence remained stronger for selective-risk ranking.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.09184v1 Announce Type: new Abstract: Vision-language model (VLM) confidence may change in aggregate when visual evidence is degraded while remaining structurally inconsistent within individual examples. We study answer-level reliability along five-step, question-conditioned evidence-loss trajectories. Using a frozen Qwen2.5-VL-3B-Instruct model, we construct 176 accepted GQA-derived trajectories (880 masking conditions) by progressively masking scene-graph-localized question-critical regions. Native sequence confidence has an evidence monotonicity violation rate (EMVR) of 0.436, and 92.0% of trajectories contain at least one adjacent violation. A matched non-critical-region control shows that full critical masking reduces accuracy by 28.2 percentage points, compared with 0.6 points for equally sized non-critical masks; the paired difference is 27.6 points (95% CI [20.0, 34.7]). We train a lightweight post-hoc reliability head on frozen hidden states, sequence confidence, and entropy. Adding evidence-order supervision to binary cross-entropy (BCE) reduces masking EMVR from 0.330 to 0.303 (paired difference -0.027, 95% CI [-0.044, -0.010]). The same mask-trained objective reduces EMVR from 0.449 to 0.402 on held-out question IDs under unseen local Gaussian blur (difference -0.0468, 95% CI [-0.0739, -0.0199]). AUROC, Brier, and AURC differences between the two learned heads are statistically inconclusive, and native confidence remains stronger for selective-risk ranking. The results separate evidence-order consistency from conventional correctness discrimination rather than establishing generic confidence superiority.

── more in #computer-vision 4 stories · sorted by recency
── more on @qwen2.5-vl-3b-instruct 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evidence-order-calib…] indexed:0 read:1min 2026-09-10 ·