# Modifying a Qwen 2.5 Omni 3B model architecture to stop visual interference when slide context is given along with audio

> Source: <https://discuss.huggingface.co/t/modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when-slide-context-is-given-along-with-audio/180749#post_4>
> Published: 2026-09-28 05:19:16+00:00

Nice find — sharper aim: let the slide inform the thinking, never the writing.

Their door starts closed — the image gets to see the question. Yours starts open — the transcript reading the slide. (Honest note: their unlock ran during finetuning, so flipping it at inference is a weaker move.)

Two switches:

Ablate first: zero the visual span on the stock model. If the hallucinated text survives, the doors aren’t the problem.

If it clears: mask transcript-rows × visual-cols to -inf. Control: the reverse — does visual grounding help or feed it?

Same slide+audio pair. Check clean-audio WER before and after: a gate that fixes the transcript but mangles audio is a trade, not a fix.

No retraining — just the mask tensor.
