cd /news/artificial-intelligence/targeting-the-attention-heads-behind… · home topics artificial-intelligence article
[ARTICLE · art-112676] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Targeting the Attention Heads Behind Object Hallucination in LLaVA

Researchers at an undisclosed institution report that targeting 32 attention heads in the LLaVA-1.5-7B vision-language model reduces object hallucination in image captions, lowering CHAIRs from 0.370 to 0.230 and CHAIRi from 0.156 to 0.096 on 400 held-out COCO images (p < 0.001, paired sign-flip tests). The combined method, using a head-sliced LoRA adapter and an inference-time grounding controller, also reduced object recall from 0.78 to 0.70, and the improvement persisted across decoding budgets (23% at 64 tokens to 58% at 128).

read1 min views1 publishedAug 27, 2026

arXiv:2608.24966v1 Announce Type: new Abstract: Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whether an interpretability diagnosis of this failure can guide a targeted fix, and we measure what that fix actually changes. We rank attention heads by how much their image attention drops around hallucinated object words, then screen the shortlist by ablating candidate heads and measuring the change in hallucination-token log probability, yielding a 32-head set. We restrict two interventions to these heads: a head-sliced LoRA adapter and an inference-time grounding controller. On 400 held-out COCO images, the combined method lowers CHAIRs (the fraction of captions with a hallucinated object) from 0.370 to 0.230 and CHAIRi (the fraction of hallucinated object mentions) from 0.156 to 0.096 (p < 0.001, paired sign-flip tests). Two controls sharpen attribution. A random-head LoRA control, matched layer-for-layer and trained identically, performs no better than the matched baseline on a separate 200-image control split, supporting the role of head selection rather than LoRA capacity. Under fixed decoding budgets, the CHAIR reduction persists and grows with budget (23% at 64 tokens to 58% at 128), arguing against a pure max-token or truncation artifact, although the method remains shorter and more conservative. The resulting behavior reduces unsupported object mentions while also lowering object recall (0.78 to 0.70). We present a diagnosis-to-intervention pipeline for object hallucination, and, more importantly, a controlled account of what acting on the diagnostic signal actually does: it localizes intervention sites with real, non-random leverage, reported as a behavioral profile rather than a single score.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @llava-1.5-7b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/targeting-the-attent…] indexed:0 read:1min 2026-08-27 ·