cd /news/artificial-intelligence/efficient-llm-adversarial-training-v… · home topics artificial-intelligence article
[ARTICLE · art-84259] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates

Researchers propose a computation-efficient method for adversarial training of large language models that cuts per-step FLOPs by 48.1% and requires only 0.0118% trainable parameters compared to standard latent adversarial training (LAT). The approach, detailed in arXiv:2607.28959v1, optimizes both the defense side via representation fine-tuning (ReFT) and the attack side by extracting relevant circuits to build a lightweight surrogate model. The authors provide theoretical and numerical evidence supporting the effectiveness of their strategies.

read1 min views1 publishedAug 3, 2026

arXiv:2607.28959v1 Announce Type: new Abstract: Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent adversarial training (LAT), have been developed, they still incur a high computational cost. In this work, we comprehensively investigate computation-efficient strategies to speed up LAT from two complementary perspectives: (1) Defense-side optimization: We explore the representation fine-tuning (ReFT) within LAT, and reveal a potential issue if there is a mismatch on which tokens to apply ReFT and the attack. (2) Attack-side optimization: When computing adversarial attacks in each LAT iteration, we extract only the relevant circuits from the LLM to construct a lightweight surrogate model, avoiding the computation in the forward-backward passes through the full model during the attack generation. For both perspectives, we provide theoretical justifications and numerical evidence to illustrate the effectiveness of the proposed strategies. Ultimately, compared to standard LAT with full fine-tuning, our method on average reduces per-step adversarial-training FLOPs by 48.1% while requiring only 0.0118% trainable parameters.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/efficient-llm-advers…] indexed:0 read:1min 2026-08-03 ·