How to Steal an AI Model’s Private Thoughts
A team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems demonstrated in August 2026 that encrypted reasoning blocks returned by Anthropic, OpenAI, a…
A team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems demonstrated in August 2026 that encrypted reasoning blocks returned by Anthropic, OpenAI, a…
Researchers at the Max Planck Institute for Intelligent Systems, ELLIS Tübingen, and ETH Zürich built LittleLearner, a 5B-parameter language model pretrained exclusively on 88 billion tokens of U.S. e…
Researchers from MATS, ELLIS Tübingen, the Max Planck Institute for Intelligent Systems, and Snyk demonstrated that encrypted chain-of-thought reasoning traces from OpenAI, Anthropic, and Google APIs …
A new controlled study from the LittleLearner project, led by researchers including Ryan Cotterell of ETH Zürich and Wieland Brendel of the Max Planck Institute for Intelligent Systems, found that pos…
A German research team from the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems has demonstrated a simple replay attack that extracts hidden reasoning traces from proprie…
Security researchers from MATS Research, the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, and Snyk found that Anthropic, OpenAI, and Google use a single global encryptio…
Researchers Alexander Panfilov, David Schmotz, and Ilia Shumailov found a way to recover encrypted reasoning from models by OpenAI, Anthropic, and Google, as detailed in a paper posted on August 10th.…
Current AI agents are fundamentally limited because they operate on a single, sequential token stream, forcing them to perform reading, thinking, and acting one step at a time. Researchers from the Ma…