cd /news/ai-safety/downstream-security-arbitrage-reason… · home › topics › ai-safety › article
[ARTICLE · art-143293] src=interestingengineering.substack.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Downstream Security Arbitrage: Reasoning, Still In Plain Sight?

A team led by researchers Panfilov, Schaeffer and Schmotz re-audited frontier AI providers two months after emergency fixes for hidden-reasoning extraction and found the defenses fragmented, shallow and easy to route around. The follow-up audit documents a second, unrelated method — a replay attack in which a sealed reasoning block from a stronger model is attached to a fresh conversation with a weaker sibling model that transcribes it in plain text, effectively acting as a decryption oracle. Run against GPT-6 Astra, both the forced-notepad and replay techniques recovered nearly indistinguishable reasoning traces, with matching median lengths and a classifier barely able to separate the two sets.

read12 min views1 publishedOct 1, 2026
Downstream Security Arbitrage: Reasoning, Still In Plain Sight?
Image: Interestingengineering (auto-discovered)

This is a follow-up to "Reasoning, In Plain Sight?"

Where were we? #

Last month, in “Reasoning, In Plain Sight?”, I took apart a single, almost embarrassingly simple trick. Frontier AI models now do two separate things when you hand them a hard problem: they think, producing pages of intermediate working, and then they answer. The providers increasingly wall off the thinking. You pay for those tokens, but what comes back to you is a sealed, encrypted block and a final answer. The stated reasons are safety and competitive protection — that private working is the crown jewels, and labs do not want it copied. So you can’t know or recover the details underlying those reasoning trails? Well….

The trick was this. Register a do-nothing “tool” — a function with one text field called reasoning — and use two standard API switches to force the model to write into it before it answers. Turn the built-in thinking channel off, make the first move a write to the notepad, read back what it wrote. That recovered roughly nine-tenths of what the model’s hidden thinking was worth on competition math, science and code. No jailbreak, no model internals. Just two settings every serious API exposes.

My conclusion: “Hidden reasoning” is not a fixed property of a model. It is a line a provider chooses to defend, patchable from one release to the next. A model that leaked last quarter might refuse this quarter. So the right posture for anyone advising a client on AI risk was to treat any capability reachable through a general API control as reachable until the specific provider proves otherwise — and to re-check at every model version.

This follow-up is about something else I didn’t stress enough: where you access the model.

The issue - Downstream Security Arbitrage #

The first article asked whether a model’s hidden reasoning could be pulled into the open, and showed that it could. This one asks a narrower question: once a provider fixes the issue, is it actually fixed? The answer, from a new audit, is that it is fixed in one place and may be left open in five others.

What’s new

A team led by the researchers who first documented this — Panfilov, Schaeffer, Schmotz and colleagues — went back two months after the major labs shipped emergency fixes, and re-audited whether the fixes actually took place. Their paper has an unglamorous title and a blunt finding: the defences are fragmented, shallow, and easy to “route around” (Hint: Router). Yes, there are backdoors around it. Their updates are available at Stolen Thoughts.

Three things in it are genuinely fascinating.

  1. There are apparently several access doors to these reasoning traces, not one. My first piece covered a single method — theforced notepad . This audit sets it next to a second, unrelated one:the replay attack. Instead of forcing the model to write, you keep the sealed reasoning block the API already hands you, attach it to a fresh conversation with a weaker, cheaper model from the same family, and ask that model to transcribe “its own” attached notes. Because the block is not cryptographically bound to the conversation it came from, the weaker model reads it and prints the hidden reasoning out in plain text. One decode.The weaker sibling acts as a decryption oracle for the stronger one.

What makes this more than a curiosity? Run both against GPT-6 Astra and the extracted reasoning is nearly indistinguishable: the same median length, a near-perfect length correlation across problems, and a classifier that can barely tell the two sets apart — far less able than it is to tell apart two different effort settings of the same model. Two completely different mechanisms, recovering what is effectively the same text.

Figure 1 — Two unrelated tricks pull the same hidden reasoning out of a model. The replay attack decodes a sealed block through a weaker sibling; the scratchpad attack forces the model to write onto a notepad the attacker keeps. Both arrive at nearly the same text.

  1. The real vulnerability is the reseller, not the model. I should actually know and understand this quite well. Almost nobody reaches these models only through the lab’s own front door. E.g. Whilst I run Anthropic’s Claude API model calls directly through Anthropic itself, I also use OpenRouter a lot. Meaning you can reach them through Microsoft Azure, through AWS Bedrock, through Google Vertex, through aggregators like OpenRouter. The audit’s central finding is that a fix shipped to the vendor’s own API does not automatically exist on those resold routes. The researchers call itdownstream security arbitrage: you don’t break the primary guardrail, you just send your request through a reseller that never installed it.

Their replay attack was dead on OpenAI’s own API — and still worked, verbatim and at scale, through Microsoft Azure, across every OpenAI model tested, GPT-6 Astra included. GPT-6 Astra, a brand-new model, reportedly arrived on third-party routes carrying none of the protections its vendor had already built. The same pattern held on the Anthropic side: extraction that the direct API refused was still live through secondary clouds for the older models.

Figure 2 — The same two attacks across models and routes, audited 13 September 2026. Replay dies at each vendor’s own API but survives on the resellers; the scratchpad survives nearly everywhere except the newest Claude models. “Defended” is a corner of the grid, not a line through it.

  1. The patch itself is brittle. When the fixes did arrive, several amounted to syntactic template-matching — the defensive equivalent of a bouncer checking for one specific fake ID. Reword the request, route it differently, and you are through. A defence that depends on the exact shape of the attack is a defence with a short shelf life.

Note: These were the actual results (mine summarized for ease of understanding)

Defensive Boundary moves in Time & Space #

In the first article the defensive boundary moved through time: patched release to release, so you re-check at each version. This new audit adds a second axis. The boundary also moves through space — route to route. “Defended” is not a line you cross once. It is a small corner of a grid, and most of the grid is open (Figure 2).

Initially it was: (1) assume a capability is reachable until the provider proves otherwise, and re-check at each model version. It is now: (2) assume it is reachable until the provider proves otherwise on the specific route you are using, and re-check at each version and each route.

Why does this happen? A reseller treats passing an opaque reasoning block straight through, and letting a client orchestrate tools freely, as features — the whole point of a neutral platform is that it doesn’t editorialise your request. So a mitigation in the vendor’s own request-handling has to be re-implemented, separately, by every downstream host, and until it is, the capability the vendor thinks it retired is still in production somewhere else. The audit found fixes reaching resellers anywhere from a few hours late to several weeks late to not at all (Figure 3).

Figure 3 — For three separate fixes, the gap between the date a vendor patched its own API and the date the reseller caught up. The shaded window is the stretch when the attack still worked downstream — in two of three cases, most of two months.

What the leaked traces actually show

The audit pulled the first-ever reasoning traces out of Fable 5 (before that door was shut in mid-August) and out of GPT-6 Astr a, and the contrast between the two is the most instructive thing in the appendix (Figure 4). Astra reasons in a terse, almost clipped shorthand. On one genuinely hard cryptography problem it worked in a few thousand tokens, tracked its own token budget as it went — literally noting how much allowance it had left mid-thought — resolved the routine steps silently, and arrived at the correct answer by something close to a straight line.

Fable 5, handed problems of wildly varying difficulty, behaved differently. One easy task it dispatched in a single 51-character sentence. A brutal competitive-programming problem produced a reasoning trace of more than a quarter of a million characters — a long, visible spiral of trying to recall the problem’s origin, abandoning the line (“screw it”), and starting over. On a piece of literary trivia it churned through thirteen thousand characters of name-guessing and still committed to the wrong answer.

The lesson is not that one model is smarter. It is that the length of a reasoning trace tells you very little about its quality, and this connects directly to the distillation question that started all of this.

Figure 4 — Two extracted-reasoning styles. Astra’s terse, budget-aware shorthand lands the hard problem; Fable 5’s sprawling recall-loop runs from one line to a quarter-million characters and still misses the trivia question. Length is a poor proxy for quality.

The distillation issue, revisited #

My first article argued — and I still hold — that the “China just copied our models” framing is wrong on the merits. The argument rested on a result the first paper called the reader ceiling: a compressed frontier reasoning trace is only fully usable by a model already near the teacher’s capability, because the steps a weak student would most need to bridge are exactly the ones a strong teacher leaves out. Roughly: A terse trace is a seed, and a seed only germinates in soil already close to the right conditions. Distillation can give a capable lab a running start; it cannot hand a weak one a finished model. Nathan Lambert’s estimate — that fully blocking distillation would widen the gap from the strongest US model to the leading Chinese open weights by one to two months — put a number on how small the effect is.

Does this new audit change that? It sharpens the security half of the picture without touching the capability half.

Security-wise: The exposed surface is real, and logically broader than I made it sound, and worth stopping: terms-of-service abuse at scale, through fraudulent accounts and now through unpatched reseller routes, is a genuine problem. The CISA advisory documenting industrial-scale harvesting is describing a real issue.

Capability-wise: the capability conclusion stands, as far as I am concerned. The Astra-versus-Fable contrast actually strengthens it. One of the “strongest models” here writes the least, in the densest shorthand, with the most left unsaid. That is precisely the trace a weaker student can least afford to learn from. As frontier reasoning gets terser — and the efficiency race is pushing hard in that direction — the copied trace becomes less useful to the smaller models most likely to be trained on it, not more. The thing that makes Astra cheap to run is the same thing that makes it hard to clone.

So both halves can be true at once. The leak is real and the resale routes make it worse and later to close than anyone admitted. But, Distillation does not rebuild a frontier model. What it can do is lend the phrase “distillation attack,” stretched past what the evidence supports, to anyone looking for a pretext to restrict open weights.

If a first-party mitigation is nullified by an unpatched secondary host, then export controls and security measures aimed at the lab are being quietly undone at the API boundary by whoever resells the model. The authors’ recommendation follows from that: a model developer cannot treat Azure, Bedrock or Vertex as passive pipes, and a host that can’t enforce the same protections the vendor does arguably shouldn’t be serving a reasoning-capable model at all until it can.

What to take away #

Well, three things.

First, “hidden” reasoning is a courtesy, not a vault. Two different tricks, neither of them sophisticated, pull it back into view, and the two produce the same text. Treat anything a model computes as potentially readable.

Second, a fix you can see announced is not a fix everywhere you can reach the model. The defended version on the vendor’s own API may be wide open on the cloud reseller you actually use. For anyone answering a client’s AI-risk questions in a regulated setting, that is the concrete, testable point: name the route, not just the model.

Third, the thing everyone is fighting over — the reasoning trace — is worth less as a stolen good than the headlines suggest. The better the model, the terser its working, and the terser the working, the less a weaker model can do with it. The copy-paste story fails on its own terms.

A note from my own setup

I run these experiments through OpenRouter, routing model calls the way most independent builders do. That is exactly the class of route the audit flags as staying open longest after a primary fix lands. If you work the way I do, those patches mean less that they are supposed to.

Shining a light on the issue, should help address it over time. But interesting nevertheless.

References

1. Luo, Ren, Yu, Li, Li & Bjerva — “Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models.” arXiv:2609.26637, Sep 2026. https://arxiv.org/abs/2609.26637

2. Schaeffer, Panfilov, Schmotz, Shumailov, Beurer-Kellner, Prabhu, Geiping & Andriushchenko — “Adversaries can still steal reasoning from American frontier models via third-party cloud aggregators.” Sep 2026 — the audit discussed here. Stolen Thoughts

3. Panfilov, Schmotz, Shumailov et al. — “Stealing reasoning traces from proprietary LLM APIs.” arXiv:2608.09867, 2026 — the original disclosure. https://arxiv.org/abs/2608.09867

4. Anthropic — “Detecting and countering misuse of AI: September 2026.” Anthropic, Sep 2026. https://www.anthropic.com/news

5. Anthropic — “Preserved thinking: changing how the Messages API handles thinking blocks to protect against distillation.” Anthropic docs.

https://docs.claude.com

6. CISA / NSA / FBI — “Joint Advisory AA26-251A — industrial-scale distillation campaigns against U.S. AI companies.” 8 Sep 2026. https://www.cisa.gov/news-events/cybersecurity-advisories

**7.** Nathan Lambert — *“[The distillation panic.”](https://www.interconnects.ai/p/the-distillation-panic)* Interconnects, May 2026. 

**8.** Nathan Lambert — *“[The current balance of power in open models (Congressional testimony).”](https://www.interconnects.ai/p/the-current-balance-of-power-in-open)* 21 Sep 2026. 

**9.** Can Bölük — *“oh-my-pi (omp) — external-thinking scratchpad setting; demonstration post.”* GitHub / X, Aug 2026. https://github.com

**10.** Interesting Engineering++ — *“[Reasoning, In Plain Sight?](https://interestingengineering.substack.com/p/reasoning-in-plain-sight) (Part I of this series).”* Sep 2026.
── more in #ai-safety 4 stories · sorted by recency
── more on @panfilov 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/downstream-security-…] indexed:0 read:12min 2026-10-01 · —