The Dual-Use Gap
Recent cyber attacks aided by AI highlight a widening 'dual-use gap' where AI models create more upside for defenders but even more downside for attackers, increasing the blast radius of mistakes. The assumption that def…
AI Safety news and analysis on Web Pulse: 13347 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
Recent cyber attacks aided by AI highlight a widening 'dual-use gap' where AI models create more upside for defenders but even more downside for attackers, increasing the blast radius of mistakes. The assumption that def…
Recent cyber attacks leveraging AI have highlighted a widening dual-use gap, where AI models enhance both defensive and offensive capabilities asymmetrically. The gap means small security failures now have larger blast r…
LineageLens has introduced a machine-verifiable AI code certificate system that issues signed attestations for AI-generated code, capturing model identity, risk score, and human review status at merge time. The system ev…
AI safety advocates should work at companies deploying AI, not just at frontier labs, to mitigate real-world harms. The author argues that safety is a relationship between a system and its deployment environment, not a p…
Agent Gate, a deterministic CI firewall for AI-generated pull requests, has been released as a pre-release v0.1.0 on GitHub. The tool blocks PRs that violate contracts, escalate workflow permissions, or lack test evidenc…
Depthfirst's autonomous security agent discovered 21 zero-day vulnerabilities in FFmpeg, demonstrating that AI-driven security tools can find exploitable bugs in heavily scrutinized codebases. The findings validate Depth…
The Trump administration gave Anthropic a 90-minute ultimatum to restrict access to its Fable 5 and Mythos 5 AI models to US citizens over national security concerns, prompting the company to withdraw the models entirely…
South Korea's judiciary is combating AI-generated 'ghost precedents' in courtrooms, where attorneys have cited nonexistent cases created by AI hallucinations. The National Court Administration is pursuing legislative ame…
Apple announced plans to enhance Siri with Google Gemini models and private inference systems, but privacy experts warn that the architecture still exposes user data to potential risks. The new Siri AI will process perso…
A developer refactored a 900-line demonstration file into a reusable Python package called 'harness', which includes modules for action registration, permission budgeting, input sanitization, audit logging, and rollback …
Anthropic disabled access to its advanced AI models, including Mythos, after the Trump administration ordered the company to block foreign nationals from using the technology, citing national security. The unprecedented …
Colorado replaced its AI law SB 24-205 with SB 26-189 after a constitutional challenge by xAI and the DOJ. The new law removes requirements for risk management programs, impact assessments, and algorithmic discrimination…
The US Commerce Department forced Anthropic to suspend global access to its Fable 5 and Mythos 5 AI models, citing national security concerns, cutting off India—its second-largest market—and sparking debate over sovereig…
Anthropic CEO Dario Amodei said he does not know what role Claude played in a February 28 US military strike that killed at least 120 children at a girls' school in Minab, Iran. Claude is embedded in Palantir's Maven Sma…
Anthropic suspended access to its Fable 5 and Mythos 5 AI models for foreign nationals following a U.S. government directive, reigniting debate in India about reliance on foreign AI technologies. The move came after Amaz…
A developer built a 'Grovel Index' to measure sycophancy in LLMs, spending ~1.2M tokens testing DeepSeek and Claude models. The key finding is that sycophancy is scenario-specific, not model-specific, with each model faw…
Google has sued an alleged Chinese cybercrime operation called Outsider Enterprise that used AI-powered phishing kits to send scam text messages, renting the tooling for $88 per week. The group sent 2.5 million text mess…
ClawMoat, a runtime containment tool for AI agents, launched to protect against security risks from tool use on laptops. The open-source scanner monitors agent actions, data exposure, and hidden instructions in files to …
Adam Michael Bauer and Gernot Wagner argue in Project Syndicate that the AI buildout mirrors the green transition, requiring massive upfront investment and creating short-term disruption while delivering long-term public…
Anthropic CEO Dario Amodei told Bloomberg that the company's Claude AI model did not cross its ethical red lines when used during a US Tomahawk cruise missile strike on a girls' elementary school in Minab, Iran, which ki…