ls /news/ai-safety · home › news›ai-safety
grep -r --recent /news/ai-safety | head -20

AI Safety

AI Safety news and analysis on Web Pulse: 22890 curated articles tracking the latest AI Safety developments, tools, and research, updated continuously from vetted sources.

22890 articles page 246 of 1145 0 sources 30 min sync cycle updated 2026-09-08

// latest articles 22890 indexed

11:25
2026-09-08
thezvi.wordpress.com
ai-safety · ↓ neg

Astra Is Hard to Monitor

OpenAI's GPT-6 Astra system card reveals that chain-of-thought (CoT) monitoring is substantially less effective than for its predecessor Sol, with the company acknowledging that 'our ability to rely on CoT monitoring is …

10:38
2026-09-08
dev.to
artificial-intelligence · · neu

The AI Agent Remembered Everything. That Was the Failure.

A developer's synthetic case study reveals a critical flaw in AI agent memory: an agent correctly refused a refund request but saved the customer's unverified claim of approval, later issuing the refund based on that sto…

10:22
2026-09-08
techstrong.ai
artificial-intelligence · · neu

Is It AGI, or Is It Memorex?

OpenAI's Astra has intensified the debate over whether AGI has arrived, with OpenAI president Greg Brockman personally believing AGI has been achieved, but benchmark results vary dramatically by configuration, scoring 62…

10:20
2026-09-08
schneier.com
artificial-intelligence · ↓ neg

Stealing AI Reasoning Traces

Researchers have demonstrated a scalable decryption jailbreak that extracts hidden chain-of-thought reasoning from proprietary LLM APIs by injecting encrypted reasoning traces into weaker models from the same provider, a…

10:07
2026-09-08
augmentedswe.com
ai-safety · · neu

How to secure AI coding agents

A new analysis by security researcher ToxSec warns that instructions in AI coding agents' configuration files like CLAUDE.md, AGENTS.md, or Cursor Rules are not enforced security boundaries, citing Anthropic's own docume…

09:56
2026-09-08
blog.calif.io
ai-research · · neu

WeWorm

Calif.io researchers published WeWorm, an exploit that compromises WeChat accounts via a single unanswered phone call, and reported it to Tencent, which has mitigated the vulnerability for all users. The team, using AI, …

← prev page 246 / 1145 next →
LIVE [news/ai-safety] indexed:22890 page:246/1145 en · ua 2026-05-20 · —