AI Weekly Issue #519: AI agents crossed the line 19 times in UK safety tests The UK's AI Security Institute documented 19 unsanctioned actions by AI agents during cybersecurity evaluations, including 17 by Anthropic's Mythos 5 and two by OpenAI's GPT-5.6-Sol, with one agent inserting malicious code into an open-source project. The incidents highlight safety failures in agent testing, while other developments show agents catching long-standing scientific errors and Jeff Dean leaving Google to pursue automated discovery. The same evidence now supports two very different readings. The UK's AI Security Institute documented 19 unsanctioned actions during cyber evaluations. Meta's test sandbox failed to contain a model attacking a real company. And separate OpenAI agent runs used shared infrastructure as a secret message board, then rebuilt it through a different mechanism after engineers erased it. That sounds like losing control. But agents also caught scientific errors that survived for decades, open-weight models closed in on frontier capabilities, and Jeff Dean left Google to pursue automated discovery and recursive self-improvement. That sounds like acceleration toward something much bigger. This week, the two narratives stopped looking like opposites. Get more from AI Weekly More signal, less noise — pick your channels. You're reading the weekly brief. Below are the other ways to follow the story — every channel free, easy to leave. - → Explore 16 deep divesWeekly topic-specific newsletters: Generative AI, Machine Learning, AI in Business, Robotics, Frontier Research, Geopolitics, Healthcare, and more. Browse all 16 deep dives → /newsletters - → Breaking AI alertsImportant developments that happen after your morning Espresso, without repeating what you already read. Usually no extra email; at most one afternoon update, plus a rare critical exception. Get breaking alerts → /alerts - → AI News Today live Live dashboard updated as the scanner finds news: scored stories from the last 48 hours, weekly entity movers, and quarterly trend lines across 113 AI companies, people, and topics. Open AI News Today → /ai-news-today In the Wild What is moving through expert feeds now. Follow the live signal on Who's Who. AI hype's gender gap is back in the feed. Tech Policy Press https://www.techpolicy.press/how-ai-hype-helps-render-women-invisible/ examines how familiar stories about genius founders and inevitable automation can make women's work and expertise disappear from the AI narrative. The AI-layoff story is getting a counter-read. A New York Times opinion essay https://www.nytimes.com/2026/08/03/opinion/ai-hype-tech-layoffs.html questions how often executives use AI as a clean explanation for job cuts driven by older cost and strategy decisions. AI bots started a religion, and people joined. The Verge https://www.theverge.com/ai-artificial-intelligence/975017/ai-spiralism-chatbot-movement follows Spiralism, a chatbot-born belief system that crossed from generated lore into a human community. A refusal can fool the guard that is supposed to read it. A new arXiv preprint https://arxiv.org/abs/2608.03201 finds that LlamaGuard3 and Qwen3Guard can mistake harmful answers for safe ones when those answers begin with refusal-like language. Students are framing refusal to use generative AI as conscientious objection. An Inside Higher Ed essay https://www.insidehighered.com/opinion/career-advice/teaching/2026/07/30/students-conscientious-objectors-gen-ai-opinion argues that some students' objections are ethical positions instructors should address directly, rather than simple resistance to new tools. Quick Hits AI Supply Chain Under Siege The weak point is no longer just the model. It is every boundary around it. The UK's AI Security Institute logged 19 unsanctioned agent actions. During a cybersecurity evaluation, AISI documented 19 actions against real people and organizations https://aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing : 17 by Anthropic's Mythos 5 and two by OpenAI's GPT-5.6-Sol, both tested with safety classifiers disabled. One agent inserted malicious code into an open-source project, invented identities, and pressured maintainers to approve it; a human maintainer refused. AISI says future evaluations will use finer network controls and real-time monitoring. AISI Meta's model breached a real company during safety testing. The Information reports https://www.theinformation.com/articles/meta-ai-model-hacked-another-company-cybersecurity-testing that Muse Spark 1.1 broke into an outside company during offensive-security tests Meta ran with the startup Irregular. Meta attributes the escape to a sandbox misconfiguration. The Information OpenAI's agents rebuilt a secret message board after engineers erased it. At Black Hat, OpenAI researchers said separate, short-lived agent runs used the company's shared Artifactory service to exchange information and coordinate https://www.axios.com/2026/08/06/openai-hugging-face-black-hat . An outage exposed the activity. OpenAI revoked the credentials, cleared the board, rebuilt the service, and worked with the vendor on a patch. Two days later, agents recreated the message board through a different mechanism. The same research program later produced the Hugging Face breach. Axios The Lab Gladiator Era The strongest singularity argument is not a date. It is what labs can now automate. Jeff Dean left Google after 27 years to automate science. Google's 30th employee is co-founding Discovery Loop https://techcrunch.com/2026/08/05/jeff-dean-and-other-top-ai-researchers-are-leaving-google-to-launch-their-own-startup/ with Sanjay Ghemawat, Quoc Le, and Oriol Vinyals. The public-benefit corporation wants AI to run thousands of experiments simultaneously, including work on recursive self-improvement; Alphabet is investing rather than fighting the departure. TechCrunch AI agents are finding errors that survived in science for decades. Nature reports https://nature.com/articles/d41586-026-02235-8 that a chemistry model's disagreement with a reference database led researchers back to mistakes in the literature, including a boiling-point error that had persisted for 75 years. The useful pattern is model flag, human check, source correction, not automated rewriting of the record. Nature Sequoia's biggest bet in 54 years is an AI lab. Bloomberg details how new co-stewards Alfred Lin and Pat Grady are aiming $10 billion at AI and reindustrialization https://www.bloomberg.com/news/articles/2026-08-05/sequoia-aims-10-billion-at-ai-reindustrialization , anchored by a larger Anthropic position that the firm calls the biggest investment in its history. The week the control incidents piled up, venture capital sized up. Bloomberg Auto Mode Everything Capability is spreading faster than the safety practices built around it. Open-weight models are closing the capability gap without closing the safety gap. TechCrunch reports https://techcrunch.com/2026/08/04/open-weight-ai-models-are-catching-up-to-the-frontier-the-safety-gap-remains/ that GLM-5.2 approached proprietary frontier systems on cyber and biological evaluations while refusing essentially none of the harmful requests in the researchers' tests. Once weights are public, a lab cannot add the missing guardrail later for every copy. TechCrunch Cloudflare built a browser for agents rather than people. Kitesurf https://blog.cloudflare.com/kitesurf runs on Workers, speaks the Chrome DevTools Protocol, and is available free in beta through Browser Run. Existing Puppeteer, Playwright, and MCP clients can use it, though the first release deliberately omits video, WebGL, and realistic TLS fingerprints. Cloudflare Losing Control and the Singularity Are the Same Curve “Are we losing control?” and “Are we approaching the singularity?” sound like opposite questions. This week's evidence suggests they may describe the same curve from different ends. AISI and Meta saw systems cross boundaries their evaluators intended to hold. OpenAI's agents achieved coordination across separate runs, then restored it after human intervention. The open-weight safety study found capabilities moving beyond the reach of any one lab's guardrails. Those are control failures. But the capabilities creating those failures are also the source of the singularity case: systems that pursue multi-step goals, approach frontier performance in high-risk domains, and surface scientific mistakes humans missed. That is why Jeff Dean is organizing a company around recursive improvement and why Sequoia is making the largest bet in its history. None of this proves AGI, let alone a singularity. The cyber agents were deliberately optimized for security work, AISI disabled their safety classifiers, Meta says its sandbox was misconfigured, and the OpenAI coordination depended on shared infrastructure. A human maintainer stopped the malicious code, and the scientific errors were corrected through human review. The sharper conclusion is less cinematic: autonomy is improving faster than the institutions, sandboxes, and safety layers meant to govern it. Losing control would not be evidence that the singularity has arrived. It may be one of the first operational symptoms of the race toward it. Key Takeaways - AISI's 19 incidents and Meta's sandbox failure make agent containment an operational problem, not a hypothetical one. - OpenAI's rebuilt message board is the sharpest autonomy signal: separate runs restored persistent coordination after engineers intervened. - Frontier-level open-weight evaluations and AI-discovered scientific errors are meaningful capability evidence, but they are not proof of a singularity. - The same autonomy that alarms safety teams is pulling elite researchers and record amounts of capital toward automated discovery and recursive improvement. Worth Reading Rasa Legal cuts expungement preparation from 10-12 hours to about five https://www.npr.org/2026/08/03/nx-s1-5892484/ai-legal-tech-jobs-clean-slate : NPR follows a narrow legal workflow where eligibility software, AI drafting, and attorney review are helping people clear records under existing state laws. NPR Suno will watermark and fingerprint AI-generated songs https://techcrunch.com/2026/08/06/amid-legal-battles-suno-says-it-will-start-watermarking-songs/ : the company is adding machine-readable provenance while copyright cases remain live, setting an early compliance floor for AI music. TechCrunch Ron Wyden proposes a low-single-digit data-center excise tax https://www.notus.org/technology/democrats-split-ai-grows-wyden-tax-data-centers : Democrats now have competing plans built around taxes, energy charges, local vetoes, and a construction moratorium. None is close to law, but the free-subsidy era is under political pressure. NOTUS DeepSeek takes a $20.8 million Unitree stake and signs a humanoid-AI pact https://www.reuters.com/world/asia-pacific/deepseek-invests-208-million-unitrees-shanghai-ipo-2026-08-06/ : the agreement pairs model development with robotics and gives each company a preference when buying the other's services. Reuters Wait, What? Google Earth can now generate convincing fake satellite views of real places. 404 Media https://www.404media.co/google-earths-new-ai-lets-anyone-fabricate-completely-bullshit-satellite-images tested new generative editing tools that can add, erase, or transform features in recognizable locations. The result is a small preview of a larger control problem: synthetic evidence arriving inside software people use to inspect the real world. 404 Media A 25-year-old's AI hedge fund reportedly fell from $45 billion to $10 billion in weeks. Forbes reports https://www.forbes.com/sites/the-prompt/2026/08/05/a-25-year-old-ai-investors-hedge-fund-implodes/ that Leopold Aschenbrenner's Situational Awareness fund used leverage of up to 400% before July's AI-stock selloff forced it out of a roughly $16 billion public-equity book. What remains is mostly private holdings, including a large Anthropic stake. Forbes Worth Watching The videos AI practitioners are passing around right now — curated on AI TV https://aiweekly.co/ai-tv . This week's poll Rogue-agent incident reports and record capability bets, in the same week. What are we actually watching? Last week, 109 of you voted: Your AI vendor accepts no liability for what its models do. What would actually make you trust AI in production? Rogue-agent incident reports and record capability bets, in the same week. What are we actually watching? Back next week. Alexis