Targeting the Attention Heads Behind Object Hallucination in LLaVA
Researchers at an undisclosed institution report that targeting 32 attention heads in the LLaVA-1.5-7B vision-language model reduces object hallucination in image captions, lowering CHAIRs from 0.370 to 0.230 and CHAIRi …