Anthropic's philosopher answers your questions [video]
Anthropic released a video featuring its philosopher answering questions, providing insights into the company's AI safety and ethics perspectives.
AI Safety news and analysis on Web Pulse: 10229 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
Anthropic released a video featuring its philosopher answering questions, providing insights into the company's AI safety and ethics perspectives.
AT&T Stadium installed black curtains on its west end for a World Cup match between Japan and Sweden, a glare-blocking measure that Dallas Cowboys owner Jerry Jones has long rejected for NFL games. The curtains address s…
Researchers have repeatedly found Android devices pre-loaded with malware sold on Amazon and other major online retailers, with one campaign affecting 10 million devices. The Electronic Frontier Foundation calls on Amazo…
The Trump administration has asked OpenAI to limit the release of its new GPT 5.6 model to a select group of partners due to safety concerns, with CEO Sam Altman telling staff that the government will approve access cust…
Norway has banned the use of artificial intelligence in elementary schools, with Prime Minister Jonas Gahr Støre arguing that AI could cause young children to miss important educational steps. The decision follows a prev…
NTT completed a proof of concept in early 2026 showing that AI training workloads can be moved to remote sites powered by renewable energy without significant performance loss, cutting energy use by up to 30% while train…
A newly discovered macOS malware called 'Gaslight' embeds fake error messages and prompt injection strings to confuse AI-assisted analysis tools. The Rust-based backdoor, attributed to a North Korean-linked threat actor,…
OpenAI will initially release ChatGPT 5.6 only to government-approved customers, according to a staff memo from CEO Sam Altman. The decision follows a Trump administration executive order asking AI companies to voluntari…
A Claude Code user lost their entire home directory, including family photos, when an agent ran a cleanup command that matched more than intended, highlighting the lack of permission boundaries in AI coding agents. As ag…
A developer scanned 1,200 real-world MCP (Model Context Protocol) configuration files from public GitHub repositories using an open-source security tool called Pluto AgentGuard. The scan found that 20.7% of configs had c…
President Trump signed Executive Order 14409 on June 2, 2026, directing CISA to issue new binding operational directives for civilian network defense and requiring agencies to deploy AI-enabled defensive tools. CISA foll…
The Trump administration has requested that OpenAI delay the release of its upcoming AI model to allow for a 30-day voluntary review process assessing cybersecurity capabilities. This request, part of a broader push for …
Anthropic accused Alibaba's Qwen AI lab of conducting the largest known distillation attack on its Claude models, generating over 28.8 million exchanges through 25,000 fraudulent accounts from April to June 2026. The all…
The Trump administration asked OpenAI to stagger the rollout of its next major model, GPT-5.6, due to security concerns, according to reports from The Information, Reuters, and The Verge. OpenAI CEO Sam Altman told emplo…
Former Commerce Secretary Gina Raimondo and former Indiana Gov. Eric Holcomb are co-founding a bipartisan nonprofit, RAISE US, with over $500 million to address potential job losses from AI. The group will partner with s…
A new report reveals that NSFW activities account for well over half of Grok's traffic, including porn generation and adult role-play chats, according to two former xAI employees. The findings suggest xAI's revenue is si…
The Trump administration has asked OpenAI to delay the release of its GPT-5.6 model, citing security concerns. OpenAI CEO Sam Altman told employees the company will release the model in limited preview with government-ap…
Anthropic accused Alibaba of orchestrating the largest known AI model distillation attack, involving 28.8 million exchanges through 25,000 fraudulent accounts, and urged Congress to strengthen export controls. The compan…
As artificial intelligence becomes more prevalent in healthcare, organizations are exploring its use in administrative workflows, clinical decision support, and remote monitoring. However, trust remains a critical barrie…
A researcher is developing fab, an interface to help human researchers make sense of research produced by many AI agents running in parallel, focusing on automated alignment research. The project aims to augment human ju…