05:19
2026-09-28
discuss.huggingface.co
large-language-models
Modifying a Qwen 2.5 Omni 3B model architecture to stop visual interference when slide context is given along with audio
A developer modified the Qwen 2.5 Omni 3B model's attention mask to stop slide images from interfering with audio transcription when slide context is supplied alongside audio. The approach first ablatβ¦