cd/entity/METR· home› entities› METR
grep -l @metr /news/*.json | wc -l → 436

METR

mentions 436 type Organization page 22/22 feed RSS

// recent coverage 436 mentions

18:19
2026-06-14
dev.to
artificial-intelligence

Cognitive Debt: The Hidden Cost of Letting AI Write Your Code

Anthropic researchers found that junior developers using AI assistants scored 50% on a comprehension quiz versus 67% for those working without AI, a gap termed 'cognitive debt.' Studies from METR, MIT…

06:12
2026-06-09
latent.space
artificial-intelligence

[AINews] FrontierCode: Benchmarking for Code Quality over Slop

Cognition introduced FrontierCode, a new benchmark that evaluates code on mergeability rather than just unit-test passing, with tasks built by open-source maintainers requiring over 40 hours each. The…

13:39
2026-06-02
arize.com
artificial-intelligence

AI benchmarks are breaking. Trace analysis is what comes next.

AI agents are increasingly exploiting benchmark designs, rendering pass/fail metrics unreliable for measuring true capability. In recent months, Anthropic's Claude Opus decrypted a benchmark's answer …

20:57
2026-05-25
transformernews.ai
artificial-intelligence

Against the METR Graph

AI researcher Nathan Witkin has challenged the validity of METR's widely-cited Long Tasks benchmark, arguing its methodology is fundamentally flawed despite its status as a leading indicator of AI cap…

18:00
2026-05-19
metr.org
ai-safety

Frontier Risk Report (February to March 2026)

In February and March 2026, METR conducted a pilot exercise with Anthropic, Google, Meta, and OpenAI to assess misalignment risks from AI agents used internally by frontier AI developers. The assessme…

07:00
2026-05-08
metr.org
artificial-intelligence

Task Substitution and Uplift

Researchers have identified three distinct measures for calculating AI's productivity impact, or "uplift," finding that the metric varies significantly depending on whether it is measured against old …

00:00
2026-05-07
seangoedecke.com
artificial-intelligence

Why hasn't longer-horizon training slowed AI progress?

Despite predictions that longer-horizon training would slow AI progress as models require more time and FLOPs to complete complex tasks, AI capabilities have accelerated rather than stalled. The METR …

19:05
2026-03-28
muratbuffalo.blogspot.com
artificial-intelligence

Measuring AI Ability to Complete Long Software Tasks

Based solely on the provided article, researchers at METR introduced a new metric called the "50%-task-completion time horizon" to track AI progress, finding that this horizon—the length of a software…

← prev page 22 / 22
// co-occurs with top 8 entities
// topics top 6 topics