{"slug": "modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when", "title": "Modifying a Qwen 2.5 Omni 3B model architecture to stop visual interference when slide context is given along with audio", "summary": "A developer modified the Qwen 2.5 Omni 3B model's attention mask to stop slide images from interfering with audio transcription when slide context is supplied alongside audio. The approach first ablates the visual span on the stock model to test whether hallucinated text survives, then masks transcript-rows × visual-cols to -inf, with a reverse control to check whether visual grounding helps or feeds the hallucination, and requires comparing clean-audio word error rate before and after. The change requires no retraining — only the mask tensor.", "body_md": "Nice find — sharper aim: let the slide inform the thinking, never the writing.\n\nTheir door starts closed — the image gets to see the question. Yours starts open — the transcript reading the slide. (Honest note: their unlock ran during finetuning, so flipping it at inference is a weaker move.)\n\nTwo switches:\n\nAblate first: zero the visual span on the stock model. If the hallucinated text survives, the doors aren’t the problem.\n\nIf it clears: mask transcript-rows × visual-cols to -inf. Control: the reverse — does visual grounding help or feed it?\n\nSame slide+audio pair. Check clean-audio WER before and after: a gate that fixes the transcript but mangles audio is a trade, not a fix.\n\nNo retraining — just the mask tensor.", "url": "https://wpnews.pro/news/modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when", "canonical_source": "https://discuss.huggingface.co/t/modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when-slide-context-is-given-along-with-audio/180749#post_4", "published_at": "2026-09-28 05:19:16+00:00", "updated_at": "2026-09-28 05:47:36.598762+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "machine-learning", "artificial-intelligence"], "entities": ["Qwen 2.5 Omni 3B"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when", "markdown": "https://wpnews.pro/news/modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when.md", "text": "https://wpnews.pro/news/modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when.txt", "jsonld": "https://wpnews.pro/news/modifying-a-qwen-2-5-omni-3b-model-architecture-to-stop-visual-interference-when.jsonld"}}