{"slug": "gradient-mirage-trainable-yet-label-unidentifiable-gradients-in-large-language", "title": "Gradient Mirage: Trainable yet Label-Unidentifiable Gradients in Large Language Model Split Learning", "summary": "Researchers propose Gradient Mirage, a defense against gradient matching attacks (GMAs) in large language model (LLM) split learning, which breaks the gradient-objective consistency that attackers exploit. The method induces inconsistency across objective, direction, and scale, using Selective Autoregressive Supervision, Scale Blinding, and Directional Privatization with von Mises-Fisher (vMF) mechanism under differential privacy. Experiments show it achieves a better privacy-utility trade-off than existing defenses under comparable fine-tuning performance.", "body_md": "arXiv:2608.18767v1 Announce Type: new\nAbstract: Gradient matching attacks (GMAs) in LLM split learning (SL) rely on a critical yet underexplored assumption: the gradient exposed at the split interface is a faithful derivative of the client's full-label training objective. This gradient-objective consistency allows a curious server to recover private labels by searching for a sequence whose induced gradient explains the observation. We propose Gradient Mirage, a defense that breaks this consistency without discarding the optimization utility of the backward signal. Our key idea is to induce the adversary to solve a misspecified inverse problem, in which no plausible label sequence in the sequence space can explain the observed gradients. Concretely, Gradient Mirage achieves this by inducing inconsistency across three dimensions: objective, direction, and scale. Selective Autoregressive Supervision derives the exposed gradient from a masked surrogate loss rather than the full-label objective assumed by the attacker; Scale Blinding then applies randomized multiplicative rescaling, obscuring the gradient's natural magnitude; and Directional Privatization further randomizes the gradient direction while preserving its magnitude through the von Mises-Fisher (vMF) mechanism under a directional metric differential privacy guarantee. Crucially, utility is preserved: the Top segment still learns from all target tokens via Dual-Track Backpropagation, the exposed gradient remains informative since each supervised token retains its complete autoregressive context, and Bottom-Gradient Recovery restores the effective gradient for Bottom-segment optimization. Extensive experiments show that Gradient Mirage provides substantially stronger protection than existing defenses under comparable fine-tuning performance, achieving a better privacy-utility trade-off.", "url": "https://wpnews.pro/news/gradient-mirage-trainable-yet-label-unidentifiable-gradients-in-large-language", "canonical_source": "https://www.machinebrief.com/news/gradient-mirage-trainable-yet-label-unidentifiable-gradients-tzki", "published_at": "2026-08-20 04:00:00+00:00", "updated_at": "2026-08-20 05:14:54.283860+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-safety"], "entities": ["Gradient Mirage", "arXiv", "LLM", "von Mises-Fisher"], "alternates": {"html": "https://wpnews.pro/news/gradient-mirage-trainable-yet-label-unidentifiable-gradients-in-large-language", "markdown": "https://wpnews.pro/news/gradient-mirage-trainable-yet-label-unidentifiable-gradients-in-large-language.md", "text": "https://wpnews.pro/news/gradient-mirage-trainable-yet-label-unidentifiable-gradients-in-large-language.txt", "jsonld": "https://wpnews.pro/news/gradient-mirage-trainable-yet-label-unidentifiable-gradients-in-large-language.jsonld"}}