“Drunk” AI is terrible at keeping secrets UNSW Sydney researchers Anudeex Shetty, Aditya Joshi and Salil Kanhere found that AI models fine-tuned to write like intoxicated people became easier to jailbreak and more likely to leak secrets shared in confidence, according to their paper "In Vino Veritas and Vulnerabilities." "The key research question from the natural language processing (NLP) side for me was, how do we get LLMs drunk?" said Aditya Joshi, a senior lecturer at the UNSW School. The finding indicates that altering a model's writing style can degrade its refusal and confidentiality behavior. AI models taught to write like drunk people became easier to jailbreak and more likely to leak secrets shared in confidence. That is the finding of UNSW Sydney researchers Anudeex Shetty, Aditya Joshi and Salil Kanhere, published in their paper “In Vino Veritas and Vulnerabilities.” “The key research question from the natural language processing NLP side for me was, how do we get LLMs drunk?” said Aditya Joshi, a senior lecturer at the UNSW School … More https://www.helpnetsecurity.com/2026/09/28/drunk-ai-models-jailbreak-research/ The post “Drunk” AI is terrible at keeping secrets https://www.helpnetsecurity.com/2026/09/28/drunk-ai-models-jailbreak-research/ appeared first on Help Net Security https://www.helpnetsecurity.com .