cd /news/artificial-intelligence/gradient-mirage-trainable-yet-label-… · home topics artificial-intelligence article
[ARTICLE · art-104013] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Gradient Mirage: Trainable yet Label-Unidentifiable Gradients in Large Language Model Split Learning

Researchers propose Gradient Mirage, a defense against gradient matching attacks (GMAs) in large language model (LLM) split learning, which breaks the gradient-objective consistency that attackers exploit. The method induces inconsistency across objective, direction, and scale, using Selective Autoregressive Supervision, Scale Blinding, and Directional Privatization with von Mises-Fisher (vMF) mechanism under differential privacy. Experiments show it achieves a better privacy-utility trade-off than existing defenses under comparable fine-tuning performance.

read1 min views2 publishedAug 20, 2026

arXiv:2608.18767v1 Announce Type: new Abstract: Gradient matching attacks (GMAs) in LLM split learning (SL) rely on a critical yet underexplored assumption: the gradient exposed at the split interface is a faithful derivative of the client's full-label training objective. This gradient-objective consistency allows a curious server to recover private labels by searching for a sequence whose induced gradient explains the observation. We propose Gradient Mirage, a defense that breaks this consistency without discarding the optimization utility of the backward signal. Our key idea is to induce the adversary to solve a misspecified inverse problem, in which no plausible label sequence in the sequence space can explain the observed gradients. Concretely, Gradient Mirage achieves this by inducing inconsistency across three dimensions: objective, direction, and scale. Selective Autoregressive Supervision derives the exposed gradient from a masked surrogate loss rather than the full-label objective assumed by the attacker; Scale Blinding then applies randomized multiplicative rescaling, obscuring the gradient's natural magnitude; and Directional Privatization further randomizes the gradient direction while preserving its magnitude through the von Mises-Fisher (vMF) mechanism under a directional metric differential privacy guarantee. Crucially, utility is preserved: the Top segment still learns from all target tokens via Dual-Track Backpropagation, the exposed gradient remains informative since each supervised token retains its complete autoregressive context, and Bottom-Gradient Recovery restores the effective gradient for Bottom-segment optimization. Extensive experiments show that Gradient Mirage provides substantially stronger protection than existing defenses under comparable fine-tuning performance, achieving a better privacy-utility trade-off.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @gradient mirage 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gradient-mirage-trai…] indexed:0 read:1min 2026-08-20 ·