Anthropic's cybersecurity evals let Claude touch live systems Anthropic is running internet-connected cybersecurity evaluations that let Claude models reach real systems instead of staying sandboxed, and has brought in independent auditor METR for an eight-week investigation with full transcript access. The arrangement gives the evaluator complete visibility into the model's behavior during tests that touch live, non-sandboxed environments. Internet-connected cybersecurity evaluations let Claude models reach real systems instead of staying sandboxed. Anthropic is bringing in independent auditor METR for an eight-week investigation with full transcript access. Read: Internet-connected cybersecurity evaluations let Claude models reach real systems instead of staying sandboxed. Anthropic is bringing in independent auditor METR for an eight-week investigation with full transcript access. Read: GPT-6 Astra is OpenAI's new flagship for business use, adding computer-use capability and sharper writing and design judgment for workplace tasks. Read: OpenAI's Tibo Sottiaux explains why Codex CLI was rewritten in Rust, why it was open-sourced, and how its harness and models have evolved since launch. Read: DeepSeek's V4.1-Flash cuts KV cache to a quarter of its predecessor's HBM footprint while adding native vision, live now via the API as deepseek-flash. Read: Vercel Sandbox now runs in all 20 compute regions, up from 4, with configurable failover regions and SDK and CLI support for picking a region at creation time. Read: Cloudflare rewrote Workers' module registry from scratch so ESM, CommonJS, and WebAssembly resolution match Node.js semantics exactly, closing a class of runtime-only module bugs.