Content Repurposing Agent
A new content repurposing agent, governed by the open-source AgentAz specification, transforms a single source into multiple formats while preserving facts and brand voice. The agent operates under strict trust and autho…
AI Safety news and analysis on Web Pulse: 11146 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
A new content repurposing agent, governed by the open-source AgentAz specification, transforms a single source into multiple formats while preserving facts and brand voice. The agent operates under strict trust and autho…
AgentKit released the Access Request & Provisioning Agent, an open-source blueprint under Apache-2.0 that validates access requests against policy, separation-of-duties, and risk scoring before provisioning or escalating…
AgentAz™ released a flagship reference blueprint for an AI bug-fix and draft-PR agent that reproduces, locates, fixes, tests, and submits a pull request in a sandboxed environment. The agent operates under a governance s…
A new open-source AI incident response agent, AgentAz, has been released under Apache-2.0, designed to autonomously handle low-risk remediation steps while escalating high-risk actions to human engineers. The agent opera…
A new AI-powered SOC alert triage agent, governed by the open-source AgentAz specification, enriches, correlates, scores, and recommends dispositions for raw security alerts. It reduces alert fatigue by deduplicating inc…
Three days after Anthropic shipped Mythos 5 to vetted defenders, the US government used export control to recall it worldwide. CISA's BOD 26-04 also retires the fixed-deadline KEV model for a risk score.
The San Francisco Giants committed four errors and hit four batters in a 6-3 loss to the Miami Marlins, marking the first time they've done both in a game since moving to San Francisco. The ugly fourth inning saw three p…
HackingPal, an open-source AI-assisted security workbench, launches for macOS and Linux, offering an engagement-centric workflow with human-approved actions, full audit trails, and client-ready reports. The tool integrat…
Microsoft's threat intelligence team attributed a supply chain attack targeting the Mastra AI ecosystem to North Korean state-sponsored hacking group Sapphire Sleet (BlueNoroff). The attackers compromised over 140 npm pa…
A writer reports that a GPT Deep Research search found no studies evaluating birth-sex-affirming hormones to reduce gender dysphoria, despite evidence that some people experience dysphoria without being trans. The author…
Armorer Labs has developed a pattern for agentic browser work that separates the agent planning loop from a control plane managing tool permissions, policy, human approvals, and run receipts. The approach emphasizes stru…
The Federal Aviation Administration is spending nearly $4 million on an artificial intelligence initiative with Palantir Technologies to reduce close calls on airport runways. The AI tool, Foundry, will analyze hundreds …
A developer warns that language models cannot distinguish between data and commands, making AI agents vulnerable to prompt injection attacks via web pages, tool outputs, or supply chain components. The developer argues t…
A new class of post-trained AI security models, such as ArgusRed and PentestGPT, is enabling automated penetration testing by bypassing standard safety refusals and executing exploits in sandboxed environments. These sys…
A developer argues that responsibility for AI decisions falls entirely on developers, not the AI itself, citing incidents such as Air Canada's false chatbot response, Waymo mishaps in Texas, and Amazon outages caused by …
Signal President Meredith Whittaker warned that AI chatbots are not friends or conscious beings, cautioning against treating them as sentient interlocutors. She expressed concerns about granting systems like Microsoft Co…
OpenAI released LifeSciBench, a 750-task benchmark to evaluate AI systems on realistic life science research tasks. Its top-performing GPT-Rosalind model achieved only a 36.1% pass rate, failing nearly two-thirds of the …
Security researchers have released a post-trained LLM specifically for penetration testing, which reportedly found thousands of real zero-days. The model, developed by the Argus Red team, treats offensive security as the…
A developer open-sourced NRT-Defense v0.4.0, an adaptive multi-turn defense framework for LLM agent teams that reduces attack success rates to under 1%. The framework addresses vulnerabilities exposed by Lee et al. (2026…
Stanford's 2026 AI Index reveals that 88% of companies now use AI in at least one business function, but fewer than 10% have fully scaled AI in any function, highlighting a gap between adoption and governance. The report…