# They Copied the Encrypted Reasoning, Then Asked a Model to Read It

> Source: <https://www.gladlabs.io/posts/they-copied-the-encrypted-reasoning-then-asked-a-m-d0620a28>
> Published: 2026-10-04 14:50:18+00:00

The first week of July, the traffic was quiet. Low volume. Nothing you’d page anyone for.

Then July 24 and 25 happened.

According to [CNBC’s reporting on OpenAI’s post](https://www.cnbc.com/2026/10/01/openai-chinas-moonshot-ai-kimi.html), the activity began on July 1 at low volume, then spiked to 16,000 requests from more than 4,000 users across those two days. OpenAI later found related activity across more than 15,000 users. It says it shut the whole thing down by July 28.

That shape is the story. A slow probe, a coordinated surge, a wide fan-out. If you’ve ever run an API with real abuse controls, you know that pattern.

## What OpenAI actually claims

Start with what’s on the record.

OpenAI says people associated with Moonshot AI, the China-based developer of Kimi, were at the core of a “coordinated campaign” to extract “protected reasoning.” Its post calls the activity “consistent with adversarial distillation.”

Two caveats matter here.

First, OpenAI couldn’t tie every operator to one actor. It attributes “a core cluster” to Moonshot-associated people, not the entire 15,000-user footprint.

Second, [CNBC](https://www.cnbc.com/2026/10/01/openai-chinas-moonshot-ai-kimi.html) reports that the operators did not breach OpenAI’s encryption, databases, or stored user conversations. This wasn’t a break-in. Everything happened through the front door, using the product as a customer.

Nothing in the coverage we pulled includes a response from Moonshot. Treat this as an allegation from one side.

## The part that made me stop reading

One detail is easy to skim past. Per the reporting, operators copied the models’ encrypted reasoning and asked a model, in a separate conversation, to decrypt it.

Read that again.

The reasoning was protected. The protection wasn’t broken in the cryptographic sense. CNBC says the encryption itself held. But the attackers allegedly got a model to hand back what the encryption was supposed to hide, just by asking from a fresh context.

If that’s accurate, it’s a classic design problem. When the system that holds the secret is also willing to answer questions about the secret, your boundary is a policy, not a wall. And policies get probed.

I’m not going to speculate about OpenAI’s internals. The sources don’t describe the scheme, and I won’t invent one. But the general lesson holds. Any protection that depends on a model declining to cooperate across conversation boundaries is a soft control. You want it enforced outside the model, in code that doesn’t care how politely it’s asked.

## Why hidden reasoning is worth stealing

Adversarial distillation, in plain terms: you query a strong model at scale, collect what it produces, and train a cheaper model on those outputs. The student learns to imitate the teacher.

Final answers are useful training data. Reasoning traces are better. They show the intermediate steps, the structure of how the model gets from question to answer. That’s the expensive part to produce and the hard part to reverse-engineer from answers alone.

So a lab that hides its reasoning isn’t being coy. It’s protecting the most information-dense part of its output. And a campaign like this one, 16,000 requests in two days, is what a harvest looks like.

This also isn’t the first accusation of its kind. [CNBC](https://www.cnbc.com/2026/10/01/openai-chinas-moonshot-ai-kimi.html) notes the findings come weeks after Anthropic accused Chinese AI developers, including Moonshot AI and Alibaba, of secretly using Claude to help train their own models.

## The irony problem

You’ve already seen the reply. The Register ran it as the headline: US model makers train on web data, but distilling theirs is framed as a “national security risk.”

It’s a fair jab, and you should hold both thoughts at once.

On one side, OpenAI is being accused of building its models on content it didn’t license. We covered [the unsealed briefs in the Authors Guild case](https://www.gladlabs.io/posts/unsealed-openai-briefs-show-execs-discussed-produc-3ceda1c0), which allege a piracy hub was a deliberate ingredient in training. Those are allegations too, but they sit awkwardly next to a complaint about someone else taking its outputs.

On the other side, hypocrisy doesn’t make the technical facts go away. Scraping a public web is not the same act as running coordinated, high-volume extraction against a paid API. Whether you think the second should be treated as a security incident or a terms-of-service dispute is a legitimate argument. Whether it happened is a separate question, and OpenAI has put specific numbers on the table.

My position: the irony is real, and it’s mostly irrelevant to what you should do with the engineering lesson.

## What this means if you run your own stack

If you build locally, this story is closer to home than it looks. Three takeaways.

**Your outputs are training data for someone.** Anything you expose through an endpoint can be harvested. If you serve a fine-tuned model to the public, rate limiting and per-account behavioral analysis aren’t optional extras. OpenAI caught this by seeing the pattern: slow start, sharp spike, clustered users. You can’t see a pattern you don’t log.

**Don’t put secrets behind model politeness.** We run local writer models in our own content pipeline, and our audits there turned up the mirror image of this problem. The models fabricated citation URLs, some on trusted domains like github.com, and checks that trusted the host let the fakes through. The fix wasn’t asking the model to be more honest. It was verification in code, outside the model. The same principle applies to hidden reasoning. If it must stay hidden, enforce that at a layer the model can’t talk its way around.

**Open weights change the economics.** If you run an open model locally, there’s nothing to extract. The weights are the thing. That’s one of the quiet arguments for local inference. You aren’t depending on a vendor’s ability to guard a boundary, and you aren’t a party to anyone’s distillation dispute. The tradeoff is yours too: you own the infrastructure and the expertise.

## What to watch

Three questions are still open.

Does OpenAI publish technical detail on how the decryption step worked? That’s what other providers will want to study.

Do other labs report the same pattern? Anthropic already made similar accusations weeks earlier. A third or fourth disclosure would turn this from an incident into an industry-wide condition.

And does “protected reasoning” survive as a product feature? If the boundary can be argued across with a second conversation, vendors will either harden it or stop pretending it’s private.

## The takeaway

Nobody broke the lock here, if the reporting holds. They allegedly walked up, asked politely from a different room, and got an answer. That’s a harder failure to fix than a bug, because it lives in the design.

If you’re building anything that serves model output to people you don’t control, assume the output is being collected. Log the shape of the traffic. Put your real boundaries in code. And if the thing you’re protecting is worth stealing, don’t let the model be the one guarding it.
