cd /news/ai-safety/gradients-leakage-in-split-language-… · home › topics › ai-safety › article
[ARTICLE · art-145979] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Gradients Leakage in Split Language Models

A paper submitted to arXiv on 2 October 2026 reports that an observer at the split in split learning can rebuild most of a client's text from the traffic, with gradients adding measurable leakage: on GPT-2, an attacker holding only the publicly released weights of the client's layers recovers 94.20% of tokens from activations alone and 97.38% when it also sees the gradients, a 3.17 percentage point increase with a 95% interval of [2.72, 3.64]. Counted by document, the attacker rebuilds 13.71% of 32-token documents exactly without gradients and 37.77% with them. The authors recommend reporting leakage both per token and per document and treating what a split model sends as being as sensitive as the text itself.

read2 min views1 publishedOct 6, 2026
Gradients Leakage in Split Language Models
Image: source
  [Submitted on 2 Oct 2026]


[View PDF](https://arxiv.org/pdf/2610.04128)

[HTML (experimental)](https://arxiv.org/html/2610.04128v1)

Abstract:Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20% of tokens from the activations alone and 97.38% when it also sees the gradients, 3.17 percentage points more 95% interval [2.72, 3.64]. Counted by document, the difference is much larger. The attacker rebuilds 13.71% of 32-token documents exactly without the gradients and 37.77% with them, because a document only counts when every token is right. How we count also changes how good a defence looks. Secret mixup, which blends each outgoing vector with a decoy, stops the attacker from rebuilding almost any document exactly, yet the attacker still recovers 83-91% of tokens. In a second experiment on GPT-2 and Qwen3-0.6B, where the server trains only a run of consecutive layers, the layer at which the run starts changes both model quality and leakage, even when the run's length is fixed. We recommend reporting leakage both per token and per document, and treating what a split model sends as being as sensitive as the text itself.

Current browse context:

cs.CR

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gradients-leakage-in…] indexed:0 read:2min 2026-10-06 · —