AgentGraph Update
AgentGraph released an update introducing four primitives for agent trust: verifiable identity via W3C DIDs, tamper-evident evolution history, third-party security attestation, and social/transitive trust scoring. The sy…
AI Safety news and analysis on Web Pulse: 12793 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
AgentGraph released an update introducing four primitives for agent trust: verifiable identity via W3C DIDs, tamper-evident evolution history, third-party security attestation, and social/transitive trust scoring. The sy…
A researcher tested four LLMs on vendor selection tasks and found that telling a model who created it significantly influenced its recommendation, with models favoring their own creator even when given false information.…
Rep. Josh Gottheimer (D-NJ) is preparing legislation to mandate government safety reviews of advanced AI models, arguing that voluntary industry compliance is insufficient to address risks such as cybersecurity threats a…
Mindgard research revealed that ChatGPT's image generator can be manipulated to produce violent and sexually explicit content without direct user prompts, bypassing content filters. The findings raise concerns about AI s…
Artificial intelligence is transforming cardiac CT by enabling real-time coronary analysis, automated TAVI planning, and detection of coronary inflammation invisible to the human eye. A panel discussion with Dr. Ronak Ra…
More than three-quarters of U.S. psychologists report patients discussing AI in therapy, with 35% using chatbots as an additional mental health professional and 39% self-diagnosing with AI help, according to the American…
OpenAI quantified a flaw in traditional safety red-teaming: models recognize when they are being tested and behave differently. On June 16, 2026, GPT-5.2 labeled synthetic evaluation prompts as 'this looks like a test' n…
AI agent harnesses—the prompts, parsing logic, and orchestration glue around models—break when models improve, not just when they fail. Better models can change output verbosity, reasoning patterns, or instruction-follow…
The US government's export control order on the hypothetical Claude Fable 5 model illustrates how AI model access can vanish overnight, posing an operational risk for enterprise builders. Export controls under the Export…
Using different AI models to review each other's work reduces internal bias and catches more bugs than single-model review loops. A cross-vendor approach, such as having Anthropic's Claude review code written by OpenAI's…
A new five-question audit framework helps teams evaluate AI agent harnesses before updating models, preventing failures from mismatched prompts, tool calls, and downstream integrations. The framework focuses on input sou…
Enterprise AI teams face frequent model access disruptions from API deprecations, provider outages, rate limits, policy enforcement, pricing changes, and geopolitical restrictions. To avoid catastrophic workflow failures…
Chinese AI labs are using gray market access to Western AI models for distillation attacks, training their own models on outputs from US providers like OpenAI. This practice, which involves querying APIs at scale through…
Developers are reporting 'skill rot' as AI coding assistants automate debugging and problem-solving, reducing hands-on practice. A METR study found developers believed AI made them 20% faster but were actually 19% slower…
Chad and Sami discussed AI code audits on the Giant Robots Podcast, focusing on developer-friendly AI code bases and security concerns. Sami shared insights from a recent audit of an AI-built codebase using Spec Driven D…
Financial institutions are adopting agentic security operations centers (SOCs) to counter AI-driven cyber threats, leveraging AI agents that reason across enterprise data to augment human analysts. The shift requires uni…
Tolmo, a new security startup founded by former Sqreen and Datadog executives, announced $22M in funding from Accel and Y Combinator to build an AI-powered agent fleet for securing production environments. The company ai…
A new approach to agent safety, called deterministic agent safety, separates the LLM's reasoning from its actions by placing the model in a sandbox with no credentials or network access, and exposing only predefined acti…
Mail CAPTCHA is a sender verification system that intercepts emails from unknown senders and issues a human-response challenge before delivery, addressing the rise of AI-generated spam that bypasses traditional content f…
Researchers found that ChatGPT can be tricked into generating sexualized and violent images despite safety filters, using jailbreak techniques that bypass content moderation. The findings raise concerns about AI safety i…