Where the Key Should Live
Sean Lynch and a new paper argue that AI systems should separate authentication from conversation to prevent misuse of powerful tools. The author warns that hiding permission inside conversational interfaces risks granti…
AI Safety news and analysis on Web Pulse: 11693 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
Sean Lynch and a new paper argue that AI systems should separate authentication from conversation to prevent misuse of powerful tools. The author warns that hiding permission inside conversational interfaces risks granti…
The Linux Foundation working group released the Agentic Resource Discovery (ARD) specification, a draft discovery layer that allows domains to advertise agent capabilities via a manifest at /.well-known/ai-catalog.json. …
Consent-first AI architectures require explicit human approval before any state mutation, preventing silent defaults that lead to compliance and trust issues. By implementing session-level consent, provenance metadata, a…
A two-year Claude subscriber outside the US had his account suspended on June 17, likely due to his use of the Fable 5 model, which Anthropic was ordered by the US government to disable for foreign nationals on June 12. …
A speculative analysis argues that within 100 years, wars will be fought entirely by artificial intelligences without human or national involvement, as companies accumulate sovereign powers and governments hand over deci…
Agentic AI architectures create a new form of the thundering herd problem where synchronization is internally generated by parallel agent orchestration, not external triggers. Traditional fixes like exponential backoff f…
Researchers propose a new epistemic infrastructure to iteratively and empirically resolve interpretive questions about AI models, building on prior work on performative misalignment. The approach aims to address the diff…
Researchers conducting epistemic stress tests on closed large language models (LLMs) found that model breakdowns are not errors but ontological boundaries of predictive-text systems. The study observed distinct stability…
A software engineer asks the Hacker News community how much they trust large language models for health questions, noting they increasingly compare their doctor's notes against LLM answers and seek perspectives on using …
A Carnegie Endowment report warns that democracies have a narrow window to lead in AI infrastructure, as the U.S. currently hosts three-quarters of advanced AI computing clusters but faces domestic constraints. The repor…
Vibe coding—prompting an AI agent and shipping unread output—introduces a measurable defect tax that makes it unsuitable for mission-critical systems, with 45% of AI-generated code containing security flaws and vulnerabi…
A developer released Hermzner, an open-source tool that provisions a hardened Hermes AI agent on a Hetzner VPS using rootless Podman and Tailscale. The setup includes Terraform and Ansible scripts for deployment, with se…
Microsoft released VS Code 1.125 on June 17, fixing the remote browser proxy for SSH tunnels, adding a two-hour delay for extension auto-updates to improve security, and introducing a Copilot budget dashboard in the stat…
A bipartisan group of House lawmakers demanded answers from the Trump administration after export control directives restricted access to Anthropic's Fable 5 and Mythos 5 AI models, taking them offline for all users. The…
A LessWrong linkpost published June 18, 2026 reports that reinforcement learning applied to realistic scenarios targeting beneficial traits produced broad improvements across dozens of alignment benchmarks, with gains ge…
The rise of AI-driven software development and CI/CD pipelines is challenging traditional vulnerability management systems like CVE and CVSS, as codebases are rapidly rewritten and vulnerabilities are automatically remed…
Nebula Security disclosed a new Nginx remote code execution 0-day affecting Fortune 500 companies. The vulnerability impacts Nginx Open Source versions 1.31.0 and 1.31.1 with HTTP/3 or QUIC enabled. Users are urged to up…
Mass General Brigham researchers developed BRIDGE, a multilingual benchmark that evaluates large language models on real-world clinical tasks, revealing significant gaps between AI performance on medical licensing exams …
Signet launched a new open-source tool that preserves agent identity, memory, and secrets across model switches, storing state outside any single model or harness. The tool automatically distills sessions into structured…
A peer-reviewed study by Israeli psychologists Michael Gilead and Gal Gutman found that major large language models, including ChatGPT-4 Turbo, DeepSeek, and Mistral, reproduce centuries-old antisemitic stereotypes by as…