Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis A controlled study from arXiv (2607.26541v1) finds that varying speech delivery presets while holding transcript content fixed can jailbreak audio LLMs, with the Q=1 Panic preset achieving 38/95 successful attacks on Qwen2-Audio versus 4/95 for Neutral. The PJ-Break protocol and AdvAudio-Prosody benchmark show that emotional-delivery audio alone (44/95) is far more effective than emotional text alone (11/95), indicating matched-text speech delivery should be a first-class factor in audio LLM safety evaluation. arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with presets targeting arousal, authority, and speaking rate, together with AdvAudio-Prosody, a 600-sample benchmark with acoustically verified attributes. On the exact post-QC Qwen2-Audio panel, the Q=1 Panic 38/95 , Anger 35/95 , and Fast 32/95 presets are all well above Neutral 4/95 . The fixed six-query pool covers 44/95 Qwen2-Audio seeds and 15/95 GPT-4o seeds and exceeds a matched-budget StyleBreak reimplementation 27/95 on Qwen2-Audio. A same-voice pool excluding the confounded Commanding condition still reaches 40/95, and a retained-panel ablation shows emotional-delivery audio alone 44/95 is far more effective than emotional text alone 11/95 . Exploratory surrogate diagnostics and pilot mitigation observations are secondary, non-core analyses. Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation