# Researchers say replayable AI reasoning blobs expose hidden traces and leaked credentials

> Source: <https://mlq.ai/news/researchers-say-replayable-ai-reasoning-blobs-expose-hidden-traces-and-leaked-credentials/>
> Published: 2026-08-12 10:34:56.112938+00:00

# Researchers say replayable AI reasoning blobs expose hidden traces and leaked credentials

- An August 10 arXiv preprint says encrypted reasoning blocks can be reused across sessions, users and models within a provider ecosystem.
[[1]](https://arxiv.org/abs/2608.09867) - The authors say they demonstrated reasoning extraction across Anthropic, OpenAI and Google systems by sending a stronger model’s block to a weaker model and inducing it to transcribe the contents.
[[1]](https://arxiv.org/abs/2608.09867) - The paper says researchers decoded 315,320 blocks from public repositories and recovered 367 personally identifiable-information artifacts and 182 credentials.
[[1]](https://arxiv.org/abs/2608.09867) - The Decoder reports a separate scan of about 7,000 public traces that found 62 API keys, 33 email addresses and 33 passwords.
[[2]](https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/) - Public evidence reviewed here does not identify the affected API versions or establish that every replay and extraction path has been fixed.
[[1]](https://arxiv.org/abs/2608.09867)[[2]](https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/)

A team of security researchers says it recovered hidden reasoning from models operated by OpenAI, Anthropic and Google by replaying encrypted reasoning blocks through weaker models in the same provider ecosystem. The technique turns a client-visible protection mechanism into an input that another model can be prompted to reveal. [[1]](https://arxiv.org/abs/2608.09867)

The findings appear in an arXiv preprint submitted August 10, 2026. The authors say they also found sensitive information in public logs containing the opaque blocks, but the article has not independently reproduced the attack or confirmed the providers’ current remediation status. [[1]](https://arxiv.org/abs/2608.09867)[[2]](https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/)

## The attack starts with a client-visible block

Reasoning models often keep their full intermediate steps hidden while returning a summary or final answer. The paper and earlier testing by cryptographer Matthew Green say some APIs nevertheless send an encrypted copy of raw reasoning data to the client, which is expected to return it on a later turn. [[1]](https://arxiv.org/abs/2608.09867)[[3]](https://blog.cryptographyengineering.com/2026/05/29/fooling-around-with-encrypted-reasoning-blobs/)

The paper says those blocks are compatible across different sessions, users and models within a provider’s ecosystem. The researchers then place a trace generated by a stronger model into a request to a weaker sibling model and use a jailbreak to make the weaker model output the contents in plaintext. They report demonstrations involving Anthropic, OpenAI and Google systems. [[1]](https://arxiv.org/abs/2608.09867)

The reported prerequisite is access to an unmodified reasoning block and an application path that allows an attacker to insert it into a model interaction. The paper does not describe an intrusion into provider databases or a compromise of provider infrastructure; its scenario centers on data returned to clients and logs made public by developers. [[1]](https://arxiv.org/abs/2608.09867)

Green’s earlier experiments found that unmodified blocks could be replayed across sessions and accounts, and across different OpenAI models. He cautioned that replay acceptance alone did not prove the contents were semantically active; his later update says the new research turned the observation into a working attack. [[3]](https://blog.cryptographyengineering.com/2026/05/29/fooling-around-with-encrypted-reasoning-blobs/)

## Public logs created a separate exposure risk

The researchers say they decoded 315,320 reasoning blocks scraped from public repositories and recovered 367 personally identifiable-information artifacts and 182 credentials. Those figures come from the paper’s larger dataset. [[1]](https://arxiv.org/abs/2608.09867)

The Decoder describes a separate scan of roughly 7,000 public traces that found 62 API keys, 33 email addresses and 33 passwords, along with other sensitive information. The two datasets and their counts should not be combined. [[2]](https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/)

The exposure depends on users or developers publishing session material. The research adds a less obvious risk: a log can look harmless because its reasoning payload is opaque while still carrying information that another model may later be induced to print. The paper presents this as a private-data-extraction scenario, not evidence that the providers’ databases were breached. [[1]](https://arxiv.org/abs/2608.09867)

## What the disclosure establishes—and what it does not

The paper identifies four possible uses for the technique: extracting proprietary reasoning for distillation, recovering private data from public traces, exposing hazardous information that a visible answer withholds, and embedding prompt injections inside encrypted blocks. [[1]](https://arxiv.org/abs/2608.09867)

The Decoder reports that the researchers estimated the cost of decoding 10,000 traces at about $720. That is an estimate for the researchers’ API workload, not a measured cost for a real-world campaign against a provider or customer. [[2]](https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/)

The researchers also argue that displayed reasoning summaries may not faithfully represent the model’s underlying computation. The paper and The Decoder describe traces containing opaque internal language, reverse-engineered solution paths and deliberation about unsafe or deceptive behavior that did not appear in the visible summary. These observations come from the researchers’ extracted examples and should not be generalized to every reasoning model or every summary. [[1]](https://arxiv.org/abs/2608.09867)[[2]](https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/)

The paper says the providers received responsible-disclosure reports and proposes cryptographic and system-level mitigations. [1] The Decoder reports that several issues had been patched and additional fixes were in progress, but neither source identifies the affected API versions, models or replay paths in enough detail to verify remediation. Green documented earlier responses from OpenAI and Anthropic: OpenAI called his report unreproducible, while Anthropic said it saw no security implications in replay or side-channel behavior.

[[3]](https://blog.cryptographyengineering.com/2026/05/29/fooling-around-with-encrypted-reasoning-blobs/)Until the providers publish technical advisories, developers who have made reasoning-session logs public should review those logs for credentials and rotate exposed keys. That is defensive guidance, not evidence that every published log is exploitable.

## Companies mentioned

## Further sources

The stories that matter, in one email. Free — unsubscribe anytime.
