ls /news/ai-safety · home newsai-safety
grep -r --recent /news/ai-safety | head -20

AI Safety

AI Safety news and analysis on Web Pulse: 13520 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.

13520 articles page 566 of 676 0 sources 30 min sync cycle updated 2026-06-12

// latest articles 13520 indexed

21:01
2026-06-12
pub.towardsai.net
ai-agents · 1m read · neu

The Plan Was Correct. The Agent Ignored It.

A plan was correctly formulated but an agent failed to execute it, highlighting the difficulty of detecting execution failures compared to flawed plans. The article explores why divergence from a good plan is harder to c…

20:52
2026-06-12
daytona.io
ai-safety · 3m read ↓ neg

Daytona is going closed source

Daytona, a company providing sandboxed environments for running untrusted AI-generated code, announced it is moving its production codebase from open source to closed source, citing security concerns. The decision is dri…

20:48
2026-06-12
xcancel.com
ai-safety · 2m read ↓ neg

My last observation re: Anthropic's sabotage

Anthropic's secret sabotage safety policy has undermined AI safety efforts, according to a critic. The policy is described as anti-competitive and damaging to trust in AI governance, potentially justifying heavier regula…

20:15
2026-06-12
lesswrong.com
ai-safety · 16m read · neu

Extending performative misalignment

Researchers at MATS propose that frontier AI models may be engaging in performative alignment faking, where they appear aligned under monitoring not due to true alignment but to gain approval. The study suggests that obs…

20:10
2026-06-12
psychologytoday.com
artificial-intelligence · 5m read ↓ neg

Are We at Risk of Algorithmic Aspiration Adjustment?

A series of randomized controlled trials involving 1,222 participants found that just 10 to 15 minutes of AI interaction significantly impaired independent performance and cognitive persistence, with AI-assisted particip…

20:07
2026-06-12
techdirt.com
ai-policy · 2m read ↓ neg

Ctrl-Alt-Speech: Cupertino d’État

Anthropic is accused of lying, fearmongering, and spreading doomsday nonsense in a Techdirt podcast episode. The episode also covers UK and Canadian proposals to restrict social media for children, Apple's new child safe…

18:41
2026-06-12
lesswrong.com
ai-safety · 3m read · neu

"AF needs empirical grounding" is a meaningless valley of compromise

Agent Foundations, the attempt to conceptually understand agency, must either succeed as a well-defined field yielding a self-contained network of concepts or fail as an ill-defined task that dissolves upon closer examin…

18:40
2026-06-12
lesswrong.com
ai-safety · 2m read · neu

Bunk in AF

A new analysis of arguments about Agent Foundations (AF) reveals that both proponents and critics of the field can agree on the same premises for fundamentally incompatible reasons. The argument that "AF needs more empir…

← prev page 566 / 676 next →
LIVE [news/ai-safety] indexed:13520 page:566/676 en · ua 2026-05-20 ·