cd /news/artificial-intelligence/be-consistent-enhancing-robust-visua… · home topics artificial-intelligence article
[ARTICLE · art-74914] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints

Researchers introduce ConVBench, a benchmark for evaluating logical consistency in Large Vision-Language Models (LVLMs), pairing each image with two logically equivalent questions across six categories. They also present ConVLM, which uses Group Relative Policy Optimization (GRPO)-based reinforcement learning with a consistency reward to improve model reasoning. The framework jointly assesses correctness and consistency through two metrics: logical consistency and robust accuracy.

read1 min views1 publishedJul 27, 2026

arXiv:2607.21722v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks. Existing benchmarks largely focus on symbolic mathematical or scientific problems and simple vision-centric tasks, offering limited assessment of complex visual reasoning and logical consistency, a critical requirement for reliable reasoning systems. We introduce ConVBench, a complex vision-centric reasoning benchmark in which each image is paired with two logically equivalent questions across six categories: action and state, complex counting, spatial reasoning, causal and intent understanding, commonsense reasoning, and temporal perception. To complement this benchmark, we define two evaluation metrics, logical consistency and robust accuracy, that jointly assess both the correctness and consistency of model responses. We further present ConVLM, which improves LVLM reasoning through Group Relative Policy Optimization (GRPO)-based reinforcement learning with a novel consistency reward. This method leverages automatically generated logically equivalent question-answer pairs and a dual-reward design combining accuracy- and consistency-based signals, encouraging agreement between paired responses. The framework functions effectively with or without strict answer supervision.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @convbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/be-consistent-enhanc…] indexed:0 read:1min 2026-07-27 ·