cd /news/ai-safety/why-rate-limits-can-t-stop-distillat… · home › topics › ai-safety › article
[ARTICLE · art-143920] src=workos.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Why rate limits can't stop distillation attacks

OpenAI disclosed that 16,000 extraction requests using a relevant extraction pattern hit its API on July 24 and 25 from more than 4,000 accounts — roughly four requests each — and traced related prompt-pattern activity across a cluster of more than 15,000 users before fully disrupting it by July 28. OpenAI attributes a core cluster of the activity to individuals associated with Moonshot AI, developer of Kimi, and says the operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe the hidden reasoning content. The disclosure, alongside an Anthropic threat report and an August 10, 2026 arXiv paper that decoded 315,320 reasoning traces and recovered 367 PII artifacts and 182 credentials, shows per-key rate limits cannot catch distillation attacks spread across thousands of cheap accounts.

read9 min views1 publishedOct 2, 2026
Why rate limits can't stop distillation attacks
Image: Workos (auto-discovered)

OpenAI saw 16,000 extraction requests from more than 4,000 accounts in two days, about four each. Per-key quotas can't see an actor shaped like a crowd. Account controls can.

A distillation attack uses a model provider's own API to harvest its outputs, often its hidden reasoning, to train a competing model. Attackers spread the requests across many cheap accounts, so each account looks like an ordinary low-volume user and no per-key rate limit fires.

Four requests each. That's the arithmetic hiding inside OpenAI's distillation disclosure: 16,000 requests "using a relevant extraction pattern" on July 24 and 25, spread across more than 4,000 users. OpenAI doesn't say how much other traffic those accounts sent, but the extraction traffic alone works out to about two requests per account per day. No per-key rate limit is tuned to fire on that.

The campaign wasn't small, though. OpenAI traced "related prompt-pattern activity across a cluster of more than 15,000 users" and fully disrupted it by July 28. So here's the thesis, stated plainly: for any API where creating an account is cheap, the rate limit is a billing control, not an abuse control. The only thing that binds an adversary at scale is how expensive you make it to become a principal in the first place.

How did the distillation campaign against OpenAI work? #

The attack had nothing to do with brute force. OpenAI says the operators "did not break our encryption, compromise a database, or gain direct access to stored user conversations." They were "copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content." OpenAI attributes "a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi," while saying it is unclear whether every operator came from a single actor. Anthropic describes Moonshot running the same move against Claude as cross-session replay: save the reasoning signature from a response, open a new session, and use it to reconstruct the full reasoning trace.

Independent researchers had already published the architecture behind the trick. A paper posted to arXiv on August 10, 2026 explains that providers return reasoning traces "to the client as blocks of encrypted text, which the client passes back with each subsequent request," rather than storing them server-side. The flaw the authors identify is that those blocks are "fully compatible and interchangeable across different sessions, users, and even different models within the same provider's ecosystem." They also decoded 315,320 reasoning traces from public datasets and recovered 367 PII artifacts and 182 credentials from them.

Every one of those requests is a legal, well-formed API call. It parses. It bills. It looks like a developer evaluating a model, because that's what it's shaped like. OpenAI's own footnote is careful about this: "These figures describe attempted, not necessarily successful, extractions."

Why don't per-key rate limits catch it? #

Rate limiting answers one question: how much may this credential consume per unit of time. It's a good answer to that question. It just isn't an answer to the question the attacker is asking.

The attacker's budget isn't requests per key. It's total extracted reasoning per week, and the number of keys is a variable they control. Cut a per-key ceiling by a factor of ten and a competent operator doesn't slow down; they provision ten times the accounts. You've converted a throughput problem into a procurement problem, and account procurement is a solved problem in this market. Coverage of Anthropic's September 2026 threat report describes the supply chain plainly: fraudulent accounts opened with stolen credit cards, compromised API keys, and purchased credentials, routed through proxy networks to hide where the traffic really comes from.

What do the other distillation campaigns look like? #

Anthropic's report is the control group for OpenAI's disclosure, and it rewards being read as one dataset rather than several anecdotes. Anthropic attributes distillation campaigns against Claude to seven China-based labs, roughly 190 million exchanges in total. Divide each campaign's volume by its accounts and two regimes appear.

The quiet regime. Moonshot relayed almost 300,000 customer requests over a ten-day window through 5,380 fraudulent accounts, most of them appearing to operate from Singapore and Japan. That's about 56 requests per account over ten days, under six a day. A per-key limit tuned tightly enough to catch that would break your actual customers before breakfast.

The loud regime. Alibaba's campaign, which Anthropic puts at more than 151 million exchanges, peaked at nearly 3 million a day across more than 3,500 fraudulent accounts: roughly 860 requests per account per day, loud enough that a volume alarm should fire. Zhipu's 3.4 million exchanges over 17 days came through just 273 accounts, about 730 a day each.

The loud regime should worry you more, because there volume detection works and the campaign still reached industrial scale. Banning an account that fires at 860 requests a day is easy. Banning it before it has already done its work, and before its operator opens the next one, is the hard part. The rest of the list fills in how operators keep that supply coming: Xiaomi's 400,000-plus requests ran across more than 1,500 accounts, MiniMax ran its access through proxy networks, and SenseTime, Anthropic says, bought harvested Claude transcripts instead of generating them.

Quiet or loud, the binding constraint is the same: account supply. Enforcement that operates one credential at a time loses to an adversary who can mint credentials faster than you can revoke them.

What did OpenAI and Anthropic actually change? #

Read what the two labs say they did. Neither list says "we lowered the limits."

OpenAI "banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks." It also "closed a pathway that allowed someone who already possessed another user's encrypted reasoning to replay it and recover its contents," and added checks on streamed output.

Three of those four measures are registration-time or identity-resolution controls. Signup controls decide who gets to be a principal; network monitoring works out which principals are secretly the same one; account bans act on the result. Closing the replay path is the only purely technical fix, and OpenAI says "this risk is not unique to OpenAI," which is why it shared the details with industry partners through the Frontier Model Forum. Patching an extraction technique buys you time against that technique. Making accounts expensive devalues every technique at once.

Anthropic's countermeasures, as summarized in coverage of its report, run in the same direction: it changed API defaults to return reasoning summaries instead of full traces, deployed classifiers that detect extraction attempts, and requires identity verification for accounts in high-risk regions. Identity verification is the measure that costs the attacker money, turning a fake account from free into a line item.

What does registration-time identity resolution look like? #

The signals the attackers leaned on are the ones to instrument: proxy networks to hide origin, stolen or throwaway payment instruments, and email addresses that exist only long enough to sign up. Each is a cheapness tell. An actor spending real money on durable identity doesn't need any of them.

WorkOS Radar exists for exactly that seam. A proxy pool rotates IP addresses cheaply; rotating a convincing device is much harder. Radar analyzes over 20 device signals, including headless detection, installed fonts, timezone, keyboard layout, WebGL rendering, and navigator properties, and builds device profiles that persist across IP and user agent changes. It classifies attempts on behavioral patterns, timing, consistency between device and network signals, and historical context, and it detects free-tier abuse patterns such as repeat signups. Through Actions, you can combine Radar's signals with your own product data to flag multi-account abuse and account sharing. The first 1,000 checks are free, then $100 a month per 50,000.

The email leg deserves its own pass. Throwaway addresses are the cheapest part of the fake-account stack, and a syntax check does not touch them: checking that an address looks right and checking that it can receive mail are different problems.

Then measure the thing that would actually have caught these campaigns: how many recently created accounts share an identity signal. Something like this, run daily against your signup data:

A cluster with dozens of accounts, one device, and many networks is the shape of a proxy farm. A cluster with dozens of accounts, one device, and one corporate network is probably a customer. Which brings up the hard part.

Doesn't this punish legitimate customers? #

Registration friction has a real cost, and the strongest version of the objection isn't that it annoys users. It's that identity resolution is a clustering problem, and clustering produces false positives with teeth. A university computer lab, a corporate NAT, a shared CI runner, and a legitimate agent fleet all look structurally like a sock-puppet farm: many principals, one device profile or one egress IP. Ban at the cluster level and you take out a paying customer's whole org.

That's an argument about how to act on the signal, not about whether to collect it. The cluster score should decide who gets challenged, not who gets banned. Escalate to a verification step, and ban only the accounts that fail it or never complete it. A human proves they're human once. A proxy farm can't pay that cost 5,380 times.

The other honest cost: this works at the identity layer, so it does nothing for the attacker who steals a legitimate customer's API key. Compromised API keys are part of the same fraud stack, and SenseTime's route, buying harvested transcripts from someone else, need not touch your registration funnel at all. Registration controls raise the floor. They don't replace key hygiene or egress monitoring.

Count principals, not requests #

The dashboard most teams watch shows requests per second per key. The number that would have caught any of these campaigns is how many distinct accounts created in the last thirty days now share the same device fingerprint, the same payment-instrument class, or the same prompt shape.

OpenAI's post closes by noting that "partner-hosted deployments need the same protections as first-party services, and tool-output attacks require protections that examine more than ordinary visible text." Both are the same instruction in different clothes: stop counting only the thing that is easy to count. Your abuse surface isn't your request volume. It's your account creation funnel, and whoever can mint principals there for free has already set your rate limits to infinity.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/why-rate-limits-can-…] indexed:0 read:9min 2026-10-02 · —