cd /news/artificial-intelligence/missingbench-verified-probing-vision… · home topics artificial-intelligence article
[ARTICLE · art-68003] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

A new benchmark, MissingBench-Verified, reveals that ten leading vision-language models consistently fail to detect missing object parts, with failure rates persisting even when external tool evidence contradicts the models' visual perception. The study by arXiv researchers finds that existing mitigation strategies, including tool-assisted verification and fine-tuning, provide negligible improvement, highlighting a fundamental limitation requiring architectural or training-level interventions.

read1 min views1 publishedJul 22, 2026

arXiv:2607.18673v1 Announce Type: new Abstract: Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for VLMs, stemming from both real-world knowledge bias and the scarcity of such images in training data. We present MissingBench-Verified, a benchmark designed to evaluate a specific and practically relevant scenario: when vision-language models fail to recognize that an essential component of an object has been removed. Across ten leading models, we observe consistent and significant failure rates that persist even when external tool evidence explicitly contradicts the model's visual perception. We further ask whether granting models access to image processing tools (e.g., cropping, contrast adjustment) enables autonomous inspection to resolve these failures. We find that existing mitigation strategies, including tool-assisted verification, autonomous visual reasoning, longer reasoning durations, and fine-tuning on an easier dataset, provide negligible improvement, indicating that this failure mode cannot be addressed through current prompting or post-hoc correction techniques. Our findings highlight a fundamental limitation of current VLM for inspection and monitoring tasks and underscore the need for architectural or training-level interventions that enable models to override internal expectations when confronted with contradictory evidence.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/missingbench-verifie…] indexed:0 read:1min 2026-07-22 ·