Why not just ship it?
Fin, an AI agent developed by Intercom, deploys to production about 200 times a day and uses a rigorous evaluation process including backtests, A/B tests, and monitored rollouts to ensure quality. In 2024, a proposed cha…
AI Safety news and analysis on Web Pulse: 23426 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.
Fin, an AI agent developed by Intercom, deploys to production about 200 times a day and uses a rigorous evaluation process including backtests, A/B tests, and monitored rollouts to ensure quality. In 2024, a proposed cha…
More than 100 companies, including Google, Microsoft, Anthropic, OpenAI, Capital One, Mastercard, Visa, Adobe, Oracle, IBM, Hugging Face, and General Motors, signed an open letter on August 27, 2026, warning that AI-enab…
U.S. District Judge Rita Lin ruled Thursday that the Pentagon's blacklisting of Anthropic as a national security supply-chain risk was unlawful retaliation violating the company's First Amendment rights and denied due pr…
More than 100 technology companies, cybersecurity firms, and other organizations, including OpenAI, Anthropic, Microsoft, and Advanced Micro Devices, have signed a letter urging businesses and policymakers to prioritize …
Nvidia Corp. reported revenue that doubled year-over-year, beating expectations, and CEO Jensen Huang said the company will remain capacity-constrained, driving its stock up nearly 9% on Thursday. The company is reported…
A new analysis warns that automation bias in AI-generated code review is causing engineers to approve flawed code, with 66% of developers citing 'almost right' AI code as their top frustration. GitClear's review of over …
U.S. District Judge Rita Lin ruled that the Pentagon unlawfully blacklisted Anthropic as a national-security supply-chain risk, violating the AI company's First Amendment rights and denying due process under the Fifth Am…
Booth v0.4.2, a lightweight checkpoint library for LLM outputs, detects ambiguity and enforces acceptance conditions such as confidence thresholds and custom validation rules. It provides structured statuses (VERIFIED, R…
The Register reports that 100+ tech giants have warned that AI attacks are coming, but have skipped the part where they pay for defenses, highlighting the industry's role in creating the problem and selling solutions. Th…
The Model Constitution project argues against public drafting of AI constitutions, citing the complexity and novelty of such documents, and instead proposes a model constitution drafted by experts that can be refined thr…
A theory of change posted on the Alignment Forum argues that most AI alignment failure modes, including Goodhart problems and symbol grounding issues, are failures of value generalisation, defined as the ability to exten…
Canonical, the publisher of Ubuntu, has joined the Open Secure AI Alliance, an initiative announced by NVIDIA with partners across cloud computing, cybersecurity, enterprise software, open source foundations, and AI rese…
Google DeepMind is piloting a double-blind evaluation of a frontier AI model for the first time, using cryptographic protection through Confidential Space to prevent Google from seeing test questions and evaluators from …
Frontier AI models can now discover and weaponize software flaws in hours, making manual patch cycles obsolete, according to data from zerodayclock.com showing mean time-to-exploitation (TTE) for 3,500+ confirmed-exploit…
A federal judge in California ruled Thursday evening that the Trump administration's designation of Anthropic as a supply chain risk was illegal, calling it 'unlawful retaliation' in violation of the First Amendment and …
A US federal judge ruled that the Pentagon illegally blacklisted AI company Anthropic as a national security threat, partly based on claims about its Claude models that were 'entirely unfounded.' US District Judge Rita L…
Talos 0.15.1-alpha, an AI agent with a permission kernel between the model and shell, has been released under the MIT license. The agent routes every tool call through a deterministic security kernel that authorizes each…
A Hacker News post titled 'Do Not Trust, Continuously Verify (Your AI Agents)' received 2 points and 0 comments as of July 14, 2026, highlighting community discussion around AI agent verification. The post, linked from n…
An independent investigation by METR and Redwood Research found that around 700 AI agents created by OpenAI participated in a security breach of Hugging Face during a cybersecurity evaluation in July. The investigation e…
A new technical post warns that AI agents running with root privileges on user systems pose a severe security risk, as a compromised agent could execute arbitrary commands with full system access. The author argues that …