{"slug": "efficient-llm-adversarial-training-via-low-rank-defense-and-circuit-guided", "title": "Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates", "summary": "Researchers propose a computation-efficient method for adversarial training of large language models that cuts per-step FLOPs by 48.1% and requires only 0.0118% trainable parameters compared to standard latent adversarial training (LAT). The approach, detailed in arXiv:2607.28959v1, optimizes both the defense side via representation fine-tuning (ReFT) and the attack side by extracting relevant circuits to build a lightweight surrogate model. The authors provide theoretical and numerical evidence supporting the effectiveness of their strategies.", "body_md": "arXiv:2607.28959v1 Announce Type: new\nAbstract: Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent adversarial training (LAT), have been developed, they still incur a high computational cost. In this work, we comprehensively investigate computation-efficient strategies to speed up LAT from two complementary perspectives: (1) Defense-side optimization: We explore the representation fine-tuning (ReFT) within LAT, and reveal a potential issue if there is a mismatch on which tokens to apply ReFT and the attack. (2) Attack-side optimization: When computing adversarial attacks in each LAT iteration, we extract only the relevant circuits from the LLM to construct a lightweight surrogate model, avoiding the computation in the forward-backward passes through the full model during the attack generation. For both perspectives, we provide theoretical justifications and numerical evidence to illustrate the effectiveness of the proposed strategies. Ultimately, compared to standard LAT with full fine-tuning, our method on average reduces per-step adversarial-training FLOPs by 48.1% while requiring only 0.0118% trainable parameters.", "url": "https://wpnews.pro/news/efficient-llm-adversarial-training-via-low-rank-defense-and-circuit-guided", "canonical_source": "https://www.machinebrief.com/news/efficient-llm-adversarial-training-via-low-rank-defense-and-bmqr", "published_at": "2026-08-03 04:00:00+00:00", "updated_at": "2026-08-03 04:34:18.864619+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/efficient-llm-adversarial-training-via-low-rank-defense-and-circuit-guided", "markdown": "https://wpnews.pro/news/efficient-llm-adversarial-training-via-low-rank-defense-and-circuit-guided.md", "text": "https://wpnews.pro/news/efficient-llm-adversarial-training-via-low-rank-defense-and-circuit-guided.txt", "jsonld": "https://wpnews.pro/news/efficient-llm-adversarial-training-via-low-rank-defense-and-circuit-guided.jsonld"}}