Harness Engineering Is Cybernetics
OpenAI's harness engineering approach, where engineers design environments and feedback loops instead of writing code, mirrors the cybernetic patterns of James Watt's centrifugal governor and Kubernet…
OpenAI's harness engineering approach, where engineers design environments and feedback loops instead of writing code, mirrors the cybernetic patterns of James Watt's centrifugal governor and Kubernet…
Researchers Milad Nasr and Nicholas Carlini, sponsored by Anthropic, used an AI model to improve a meet-in-the-middle attack on seven-round AES-128, reducing the complexity from 2^105 to 2^104.5 chose…
Anthropic's Claude Mythos Preview discovered CVE-2026-5194 in wolfSSL, a TLS library used in over 5 billion devices, allowing certificate forgery; the flaw was patched in version 5.9.1 on April 8, 202…
Security expert Thomas Ptacek argues that LLM agents are uniquely suited for vulnerability research, and Anthropic researcher Nicholas Carlini demonstrated that a simple prompt to Claude Opus can iden…
Developer Maneshwar is building git-lrc, a free and source-available Micro AI code reviewer that runs on every commit. The project highlights the critical need for data privacy in AI systems, warning …
Anthropic's Mythos AI model identified vulnerabilities in highly sensitive U.S. government computer systems during a testing exercise, with Senator Mark Warner claiming it broke into almost all classi…
The US government issued an export control directive on June 12 suspending access to Anthropic's Mythos 5 and Fable 5 models for all foreign nationals, prompting Anthropic to disable the models for al…
Anthropic researcher Nicholas Carlini, who initially warned colleagues in March that the company's next-generation AI model Mythos was too capable to release, has become central to Anthropic's argumen…
Anthropic sent hacker Nicholas Carlini to reassure government officials about AI safety concerns, leveraging his expertise in adversarial machine learning to demonstrate the company's commitment to re…
Anthropic sent security researcher Nicholas Carlini to Washington to address US government concerns about AI safety after export control directives forced the company to temporarily suspend global acc…
Anthropic researchers found that its Claude Mythos Preview model can identify complex software vulnerabilities and combine them into complete end-to-end attack chains, marking a significant advancemen…
Large language models (LLMs) may be transformative but pose serious risks and are already causing harm, prompting the author to question whether they are "worth it." The author, Nicholas Carlini, work…
Nicholas Carlini argues that advanced AI systems pose significant societal risks precisely because they are designed for "ruthless efficiency." He outlines a spectrum of concerns, from current harms l…
In a 2025 blog post, researcher Nicholas Carlini argues that the future of large language models is highly uncertain, with two plausible but opposing outcomes. He suggests that within three to five ye…
In his article, Nicholas Carlini explains that his research on "memorization" in machine learning models, which demonstrates that models can sometimes output verbatim training data, is often cited in …
Nicholas Carlini announced he is leaving Google DeepMind after seven years to join Anthropic for one year, citing disagreements with DeepMind leadership over its support for high-impact security and p…
In a 2025 article, Nicholas Carlini asked readers to make 30 forecasts about AI in 2027 and 2030, requiring them to give 90% confidence intervals rather than point estimates. Analyzing the responses, …
Nicholas Carlini describes a project where he uses a different large language model (LLM) each day for twelve days to completely rewrite his personal website homepage and bio. He prompts each model to…
Many people hold overly confident but vague predictions about AI's future, which are often proven wrong. To address this, author Nicholas Carlini presents a set of about 30 specific, refutable questio…
According to Nicholas Carlini's 2023 article, while computers have surpassed humans in chess for decades using specialized game-playing models, OpenAI's GPT-3.5-turbo-instruct—a language model designe…