{"slug": "amplegcg-learning-a-universal-generative-model-for-jailbreaking", "title": "AmpleGCG: Learning a Universal Generative Model for Jailbreaking", "summary": "Researchers Zeyi Liao and colleagues published AmpleGCG, a generative model that learns the distribution of adversarial suffixes from GCG optimization and generates hundreds of jailbreak suffixes per harmful query in seconds, achieving near 100% attack success rate on Llama-2-7B-chat and Vicuna-7B and 99% on GPT-3.5, according to the arXiv paper (v3, revised 24 Nov 2024). The model produces 200 adversarial suffixes for one harmful query in 4 seconds and transfers from open-source to closed-source LLMs, which the authors say makes defense more challenging.", "body_md": "# Computer Science > Computation and Language\n\n  [Submitted on 11 Apr 2024 (\n\n[v1](https://arxiv.org/abs/2404.07921v1)), last revised 24 Nov 2024 (this version, v3)]\n# Title:AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs\n\n[View PDF](https://arxiv.org/pdf/2404.07921)\n\n[HTML (experimental)](https://arxiv.org/html/2404.07921v3)\n\nAbstract:As large language models (LLMs) become increasingly prevalent and integrated into autonomous systems, ensuring their safety is imperative. Despite significant strides toward safety alignment, recent work GCG~\\citep{zou2023universal} proposes a discrete token optimization algorithm and selects the single suffix with the lowest loss to successfully jailbreak aligned LLMs. In this work, we first discuss the drawbacks of solely picking the suffix with the lowest loss during GCG optimization for jailbreaking and uncover the missed successful suffixes during the intermediate steps. Moreover, we utilize those successful suffixes as training data to learn a generative model, named AmpleGCG, which captures the distribution of adversarial suffixes given a harmful query and enables the rapid generation of hundreds of suffixes for any harmful queries in seconds. AmpleGCG achieves near 100\\% attack success rate (ASR) on two aligned LLMs (Llama-2-7B-chat and Vicuna-7B), surpassing two strongest attack baselines. More interestingly, AmpleGCG also transfers seamlessly to attack different models, including closed-source LLMs, achieving a 99\\% ASR on the latest GPT-3.5. To summarize, our work amplifies the impact of GCG by training a generative model of adversarial suffixes that is universal to any harmful queries and transferable from attacking open-source LLMs to closed-source LLMs. In addition, it can generate 200 adversarial suffixes for one harmful query in only 4 seconds, rendering it more challenging to defend.\n    \n\n## Submission history\n\nFrom: Zeyi Liao [\n[view email](https://arxiv.org/show-email/7f6fdea0/2404.07921)]\n\n**Thu, 11 Apr 2024 17:05:50 UTC (2,460 KB)**\n\n[\\[v1\\]](https://arxiv.org/abs/2404.07921v1)\n**Thu, 2 May 2024 01:08:37 UTC (2,292 KB)**\n\n[\\[v2\\]](https://arxiv.org/abs/2404.07921v2)\n**[v3]** Sun, 24 Nov 2024 18:03:43 UTC (2,292 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/amplegcg-learning-a-universal-generative-model-for-jailbreaking", "canonical_source": "https://arxiv.org/abs/2404.07921", "published_at": "2026-09-29 18:29:29+00:00", "updated_at": "2026-09-29 18:47:51.652245+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "generative-ai"], "entities": ["AmpleGCG", "Zeyi Liao", "Llama-2-7B-chat", "Vicuna-7B", "GPT-3.5", "GCG", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/amplegcg-learning-a-universal-generative-model-for-jailbreaking", "markdown": "https://wpnews.pro/news/amplegcg-learning-a-universal-generative-model-for-jailbreaking.md", "text": "https://wpnews.pro/news/amplegcg-learning-a-universal-generative-model-for-jailbreaking.txt", "jsonld": "https://wpnews.pro/news/amplegcg-learning-a-universal-generative-model-for-jailbreaking.jsonld"}}