Modifying a Qwen 2.5 Omni 3B model architecture to stop visual interference when slide context is given along with audio A developer modified the Qwen 2.5 Omni 3B model's attention mask to stop slide images from interfering with audio transcription when slide context is supplied alongside audio. The approach first ablates the visual span on the stock model to test whether hallucinated text survives, then masks transcript-rows × visual-cols to -inf, with a reverse control to check whether visual grounding helps or feeds the hallucination, and requires comparing clean-audio word error rate before and after. The change requires no retraining — only the mask tensor. Nice find — sharper aim: let the slide inform the thinking, never the writing. Their door starts closed — the image gets to see the question. Yours starts open — the transcript reading the slide. Honest note: their unlock ran during finetuning, so flipping it at inference is a weaker move. Two switches: Ablate first: zero the visual span on the stock model. If the hallucinated text survives, the doors aren’t the problem. If it clears: mask transcript-rows × visual-cols to -inf. Control: the reverse — does visual grounding help or feed it? Same slide+audio pair. Check clean-audio WER before and after: a gate that fixes the transcript but mangles audio is a trade, not a fix. No retraining — just the mask tensor.