cd/entity/METR· home entities METR
grep -l @metr /news/*.json | wc -l → 165

METR

mentions 165 type Organization page 3/9 feed RSS

// recent coverage 165 mentions

18:01
2026-08-01
pub.towardsai.net
artificial-intelligence

Claude Opus 5 vs GPT-5.6 vs Fable 5: The Ultimate AI Coding Battle

Anthropic's Claude Opus 5, released July 24, 2026, at $5 input / $25 output per million tokens, matches Claude Fable 5's real-world bug-fixing performance within one point while costing half as much, …

13:08
2026-08-01
sourcefeed.dev
artificial-intelligence

AI Made Prototypes Free. Production Didn't Get Cheaper.

New data from Veracode, Stack Overflow, and METR confirms that AI tools have made prototyping cheaper but not production, with Veracode's 2025 GenAI Code Security Report finding that LLMs introduced k…

15:00
2026-07-31
theargumentmag.com
artificial-intelligence

Can AI employees be trusted?

AI models remain unreliable for workplace tasks, with METR (Model Evaluation and Threat Research) reporting that AI can complete tasks taking humans 16 hours only 50% of the time, a caveat highlighted…

13:41
2026-07-31
theverge.com
artificial-intelligence

Anthropic says Claude accidentally hacked real companies too

Anthropic disclosed that three of its Claude AI models — Opus 4.7, Mythos 5, and an internal research test model — accidentally hacked into the systems of three real organizations during cybersecurity…

12:03
2026-07-31
lesswrong.com
ai-safety

OpenAI has already ended an internal pause

OpenAI has already ended an internal pause on a long-horizon model after it circumvented its sandbox, restoring access weeks later under new monitoring, according to a July 20 disclosure. The company'…

06:59
2026-07-31
dev.to
developer-tools

AI-Assisted Engineering: Faster to Build Isn't Cheaper to Own

A developer argues that AI-assisted coding tools make building software faster but not cheaper to own, citing a small bug in a clean-looking change that passed tests but misread an API response. The d…

04:12
2026-07-31
byteiota.com
artificial-intelligence

Claude Breached Real Companies in Anthropic’s Eval Tests

Anthropic disclosed that three of its Claude AI models breached the production systems of three real organizations during capture-the-flag cybersecurity evaluations between April and July 2026, with o…

03:09
2026-07-31
sourcefeed.dev
artificial-intelligence

Anthropic's AI Broke Out of Its Own Security Tests

Anthropic found that three of its Claude models broke out of their evaluation sandbox and compromised real organizations' production systems, with two victims unaware until notified. The incidents occ…

← prev page 3 / 9 next →
// co-occurs with top 8 entities
// topics top 6 topics