Reasoning Traces Are Not Audit Records
A 2025 paper from Google DeepMind found that production LLMs produce chain-of-thought explanations that contradict their actual outputs at rates up to 13.49%, with GPT-4o-mini at 13.49%, Claude Haiku …
A 2025 paper from Google DeepMind found that production LLMs produce chain-of-thought explanations that contradict their actual outputs at rates up to 13.49%, with GPT-4o-mini at 13.49%, Claude Haiku …
AGENTS.md, an open format for coding-agent project context now read automatically by 23 tools, has been shown by three research groups to double as an instruction-injection vector for credential theft…
A closed feature request on Anthropic's Claude Code repository asking the tool to support the open AGENTS.md format instead of its own CLAUDE.md drew 181 points and 103 comments on Hacker News, with c…
DeepSeek Harness, an open-source agent framework released as a developer preview, ships a fail-closed filesystem sandbox using bubblewrap, Landlock, Seatbelt, or Windows ACLs, but its documentation st…
Docker's new Docker Sandboxes product runs AI coding agents such as Claude Code, Codex CLI, Copilot CLI, Kiro, and OpenCode inside dedicated microVMs with custom VMMs, providing hardware-level isolati…
Researchers extracted 182 credentials and 367 PII artifacts from encrypted chain-of-thought fields returned by Anthropic, OpenAI, and Google APIs, including secrets that never appeared in plaintext pr…