Anthropic lost the White House's trust – and then its flagship product
Anthropic lost the White House's trust, leading to the loss of its flagship product, according to a Washington Post report.
AI Safety news and analysis on Web Pulse: 12952 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
Anthropic lost the White House's trust, leading to the loss of its flagship product, according to a Washington Post report.
Arcade raised $60 million in a Series A round led by SYN Ventures, with participation from Morgan Stanley and Wipro, to secure AI agent access. The company provides an authorization platform for AI agents that integrates…
Researchers from Aalto University and the University of Waterloo introduced Locket, a feature-locking technique for large language models that enables pay-to-unlock schemes by restricting specific model capabilities. The…
Argentina begins its bid for a second consecutive World Cup title on Tuesday against Algeria in Kansas City, Missouri, marking Lionel Messi's likely final World Cup appearance. The 39-year-old star, who led Argentina to …
OrangeCheck launches a suite of identity and sybil-resistance tools built on Bitcoin signatures, using BIP-322 to verify address ownership without transactions, custody, or KYC. The open-source, MIT-licensed protocols in…
A developer received a LinkedIn message from a recruiter at a crypto startup asking to review a GitHub repository before a technical interview. The repository contained a backdoor in npm's prepare script that executed on…
A developer compares four dominant alignment methods—RLHF, DPO, IPO, and KTO—for fine-tuning large language models, detailing their mathematical formulations, data requirements, and practical tradeoffs. The analysis high…
A Reuters investigation found that AI-enabled medical devices have caused numerous patient injuries, including strokes, due to malfunctions and misleading guidance. The FDA's adverse event reports for the TruDi Navigatio…
CrowdStrike launched Continuous Identity for AI Agents at Identiverse 2026, a control plane that continuously authorizes agent actions in real time using cryptographically verifiable identities based on the SPIFFE standa…
The US Commerce Department ordered Anthropic to suspend its Fable 5 and Mythos 5 AI models globally, citing national security after a jailbreak demonstration exposed vulnerabilities. The directive applies to all foreign …
A developer argues that agent evaluations must run in CI to block regressions before deployment, not as post-hoc dashboards. The post introduces agent-eval for scoring outputs and AgentLens for execution traces, advocati…
New Defence Secretary Dan Jarvis faces 'very significant cuts' if he cannot secure more funding within two weeks, after predecessor John Healey resigned over a £13.5 billion offer that fell short of requirements. The sho…
The US government treated Anthropic's LLM as a munition by ordering unprecedented restrictions on access to its Fable 5 model, effectively barring foreign nationals globally. The move, prompted by a report from Amazon, r…
A Nature npj Digital Medicine paper published June 16, 2026 introduces MedSAFE, a decision-theoretic framework for evaluating LLM abstention in healthcare, identifying uncertainty-driven and safety-driven abstention. The…
At Fortune Brainstorm Tech in Aspen, Colorado, executives from May Mobility, Thomson Reuters, Trustguard AI, and SentinelOne discussed the challenge of verifying agentic AI systems as they take on more tasks. They emphas…
The UK Government Cyber Coordination Centre (GC3) uncovered 407 security vulnerabilities across nine government departments' public code repositories through weekly AI-powered hackathons costing just £13,000 ($16,000) in…
Google DeepMind researchers trained Gemini 3 Flash to exhibit positive traits by midtraining on synthetic documents describing the model's traits, then finetuning on synthetic chat data where it demonstrates those proper…
Microsoft has published a practical guide on using Azure AI Gateway Content Safety with API Management to secure AI workflows, detailing how to block prompt injection, unsafe content, and custom blocklists at the agent a…
A security researcher demonstrated that the `aiplatform.customJobs.create` permission in Google Cloud's Vertex AI allows privilege escalation to full project control, contradicting Google's documentation that the Custom …
Cursor's Head of Security Travis McPeak warned that AI coding agents must never be trusted, advocating for secure-by-default workflows to contain damage when agents inevitably misbehave. Speaking on the Zero-Shot Learnin…