Judging AGI Output (2020)
A new AI safety researcher has raised the question of how humans can judge the output of an Artificial General Intelligence, particularly in ethical and philosophical disputes. The researcher warns that AGI decisions may…
AI Safety news and analysis on Web Pulse: 13553 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
A new AI safety researcher has raised the question of how humans can judge the output of an Artificial General Intelligence, particularly in ethical and philosophical disputes. The researcher warns that AGI decisions may…
Pope Leo XIV published an encyclical titled "Magnifica Humanitas" on May 25, 2026, calling for artificial intelligence to be "disarmed" and freed from logics of domination, exclusion, and death. The document, signed May …
Chinese autonomous driving company Pony.ai stated it remains unaffected by a U.S. safety review of self-driving vehicles. The company confirmed its operations and regulatory compliance continue as normal despite the prob…
The US military has already integrated AI into warfare, with systems like Project Maven using machine learning to analyze drone surveillance footage since 2017. Despite Anthropic's recent efforts to limit its AI technolo…
The AI Resist List, a project by We & AI, DAIR, the Refugee Law Lab, and Karen Hao, launches to track resistance to artificial intelligence. Pope Leo XIV issues an encyclical on AI calling for community-led moderation, w…
LLMs can theoretically generate millions of correct lines of code, but any important code still requires hand-auditing every line—a process harder than writing it from scratch. This creates a fundamental limit on the spe…
A new report from Omdia, commissioned by Google, found that 92% of organizations now permit employees to use public generative AI applications, with most activity occurring in the browser. The survey of 400 IT and cybers…
Pope Leo XIV on Monday released a long-awaited encyclical titled "Magnifica Humanitas" that calls for robust regulation of artificial intelligence and urges developers to work for the common good rather than profit. The …
Anthropic plans to release its Mythos-class models to the public while keeping its specialized AI flaw-finder tool restricted to a limited set of users, according to reporting by The Register. The company is developing a…
Security researchers at EclecticIQ identified an SEO poisoning campaign in early March 2026 that used typosquatted domains to impersonate Gemini CLI and Claude Code installation pages. Victims who copied a PowerShell com…
A new review framework based on Pope Leo XIV's encyclical *Magnifica Humanitas* provides tools and skills for both humans and AI agents to evaluate whether projects preserve judgment, conscience, relationships, and respo…
Nearly half of Americans now use AI to find information and generate ideas, but a professional fact-checker at WIRED reports the technology is wrong more often than people realize. The fact-checker found Google's AI Over…
Major AI labs including Google DeepMind and Anthropic have hired teams of in-house philosophers to address ethical questions around artificial intelligence, with at least 14 philosophers now employed across the two compa…
A major U.S. corporation has mandated that all employees complete an artificial intelligence training module by the end of the week or face disciplinary action, including potential termination. The policy, announced in a…
Federal law enforcement agencies, including the Department of Homeland Security and FBI, have begun surveilling a newly designated category of "anti-tech extremists" following attacks on AI executives, protests at data c…
Developer workstations have become the primary target for software supply chain attacks, with attackers exploiting unmonitored endpoints to steal credentials and install malicious packages. Security tools like EDR fail t…
Claude Code's `--dangerously-skip-permissions` flag is actually safer than the default permission-requesting mode, according to founding engineer Jim Fisher. The default mode creates "approval fatigue" that leads users t…
A critical vulnerability tracked as CVE-2026-48710 (BadHost) in Starlette versions prior to 1.0.1 allows attackers to bypass path-based authentication middleware by injecting a crafted Host header that manipulates the `r…
Most AI agents that perform well in controlled demonstrations fail in production due to messy real-world data conditions, including stale information, conflicting facts, and changing system states. The core problem is th…
Google is previewing a new Content Detection API for its SynthID AI watermarking technology, allowing businesses to identify AI-generated images from Google and other popular models. The API, available on Google Cloud’s …