How to Steal an AI Model’s Private Thoughts
A team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems demonstrated in August 2026 that encrypted reasoning blocks returned by Anthropic, OpenAI, a…
A team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems demonstrated in August 2026 that encrypted reasoning blocks returned by Anthropic, OpenAI, a…
Security researchers from MATS Research, the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, and Snyk found that Anthropic, OpenAI, and Google use a single global encryptio…
Researchers Alexander Panfilov, David Schmotz, and Ilia Shumailov found a way to recover encrypted reasoning from models by OpenAI, Anthropic, and Google, as detailed in a paper posted on August 10th.…
Computer scientists at the University of Tübingen, the Max Planck Institute, MATS Research, and Snyk discovered a method to extract hidden reasoning from frontier AI models, revealing that Chinese mod…
A researcher mapping the AI safety ecosystem for MATS Research discovered unexpected organizations, including the Human Line Project, which collects stories of AI psychosis, and Impact Academy, which …
Researchers reviewing AI safety debate protocols found that current "propose-critique-decide" models are vulnerable to gaming, where critic models exploit a "last mover advantage" by withholding key c…