The Attention Triangle in Audio-Video Models A new research paper on audio-video diffusion models reveals that cross-modal attention, which coordinates text, sound, and visual content, can introduce subtle and systematic semantic leakage. The study probes the 'attention triangle' of three cross-attention edges to analyze this phenomenon. Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the attention triangle,'' comprising the three cross-attention edge