Rohin Shah on AGI Safety
Rohin Shah, head of AGI alignment and safety at Google DeepMind, argued in a recent interview that catastrophic misalignment from advanced AI is not the likely default outcome, despite acknowledging plausible risks that …
AI Safety news and analysis on Web Pulse: 13553 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
Rohin Shah, head of AGI alignment and safety at Google DeepMind, argued in a recent interview that catastrophic misalignment from advanced AI is not the likely default outcome, despite acknowledging plausible risks that …
Security researchers at Flatt.tech disclosed a critical supply-chain vulnerability in Anthropic's Claude Code GitHub Actions integration that a single malicious GitHub issue, pull request, or comment could exploit to com…
Security researchers have documented multiple malvertising campaigns that weaponize AI-chat platforms' sharing features to deliver malware. Attackers built fake outage and download pages hosted on legitimate chatgpt.com …
Y Combinator's latest startup batch features a cohort focused on building the operational infrastructure needed to run AI agents in production, including tools for memory, identity, compliance, monitoring, and validation…
Leading AI researchers and executives have publicly warned that their own technology could be used by terrorists to develop biological weapons. The experts stated that current safeguards are insufficient to prevent malic…
A flaw in Anthropic's Claude Code GitHub Action allowed attackers to bypass permission checks via a fake bot account and use prompt injection to steal OIDC tokens, gaining write access to any vulnerable repository. Secur…
Researchers have developed a cost-effective method for detecting scheming behavior in AI agents by training small open-weight "deliberative monitors" that reason over a scheming specification before judging an agent's ac…
Google's spokesperson asked 404 Media to publish a revised statement after the outlet's initial report on internal employee memes criticizing the company's AI. The updated statement removed the previous language assertin…
Apple launched a new Safari privacy campaign one week before its Worldwide Developers Conference, signaling that the company will frame its upcoming artificial intelligence features around user trust and data protection.…
The US and its Five Eyes intelligence partners issued an unprecedented joint warning that Chinese military intelligence services are using professional networking sites like LinkedIn to recruit government and military pe…
A security engineer at a mid-size fintech company faced a 48-hour deadline from the Italian Data Protection Authority to list all AI systems processing EU resident data, including foundation models, training data sources…
A Berlin-based development shop has identified eight common security and quality issues consistently found in production code generated by AI coding assistants like Claude, Cursor, and v0. The most critical traces includ…
Top AI executives including OpenAI's Sam Altman, Anthropic's Dario Amodei, and Google DeepMind's Demis Hassabis signed a public letter urging Congress to enact laws requiring synthetic DNA sellers and machine manufacture…
Anthropic is increasingly delegating AI development to AI systems themselves, a trend that could lead to recursive self-improvement where AI autonomously designs its own successor. Internal data shows Anthropic engineers…
A trial of XBOW's autonomous offensive security platform uncovered a vulnerability that led to a full takedown of a development environment used by Moderna. Security leaders say advanced AI models are discovering softwar…
Hobbyist Steven Cheng built an AI-guided laser system that eliminated all mosquitoes in his home, according to reports from SlashGear and Tom's Hardware. The prototype, which took four months to create, uses a custom dee…
The Cybersecurity and Infrastructure Security Agency will release a binding operational directive to federal agencies by the end of the week to implement the president's artificial intelligence executive order, CISA Acti…
Anthropic released Claude Opus 4.8, an incremental improvement over Opus 4.7, as the Trump administration's Executive Order on AI returned, establishing a prior restraint framework for frontier model releases. OpenAI pub…
Timnit Gebru was fired from Google in December 2020 for refusing to retract a research paper warning about the dangers of large language models. Every prediction in that paper — including hallucination, bias amplificatio…
A developer has outlined three rules for safely using AI agents in legacy codebases: tests first, small diffs, and business rule discovery before any change. The approach warns that AI agents, when let loose on old syste…