Nice find — sharper aim: let the slide inform the thinking, never the writing.
Their door starts closed — the image gets to see the question. Yours starts open — the transcript reading the slide. (Honest note: their unlock ran during finetuning, so flipping it at inference is a weaker move.)
Two switches:
Ablate first: zero the visual span on the stock model. If the hallucinated text survives, the doors aren’t the problem.
If it clears: mask transcript-rows × visual-cols to -inf. Control: the reverse — does visual grounding help or feed it? Same slide+audio pair. Check clean-audio WER before and after: a gate that fixes the transcript but mangles audio is a trade, not a fix.
No retraining — just the mask tensor.