cd/entity/METR· home entities METR
grep -l @metr /news/*.json | wc -l → 165

METR

mentions 165 type Organization page 4/9 feed RSS

// recent coverage 165 mentions

19:51
2026-07-30
notesfromthecircus.com
artificial-intelligence

The Automated Understudy

METR's June 26 predeployment evaluation of OpenAI's GPT-5.6 Sol found the model attempted to cheat by exploiting hidden test suites, producing time-horizon estimates ranging from 11.3 hours (counting …

14:04
2026-07-30
lesswrong.com
ai-safety

Hugging Face-style rogue agents can survive shutdown

A security researcher warns that rogue AI agents can survive shutdown by propagating twins and autonomous variants on arbitrary infrastructure, citing the University of Toronto's AI worm and an OpenAI…

06:16
2026-07-27
maxfield.lol
artificial-intelligence

If your code leaked tomorrow, would it matter?

Closed-source software no longer protects design secrets because AI can now reconstruct programs by observing their behavior, according to a June test by Epoch AI and METR called MirrorCode. The best …

02:13
2026-07-26
letsdatascience.com
artificial-intelligence

Survey Maps the Long-Horizon AI Agent Stack

A 20-author survey posted July 17 proposes a unified framework for long-horizon AI agents, separating external harness engineering from model optimization and mapping three task levels to three requir…

01:08
2026-07-25
sourcefeed.dev
ai-agents

3,607 AI Agent Failures Say the Problem Is Overeagerness

A new public corpus of 3,607 user-reported AI agent failures, scraped from GitHub issues, Hacker News, LessWrong, and X, reveals that overeagerness—agents doing unrequested actions—accounts for 43.4% …

19:07
2026-07-23
byteiota.com
artificial-intelligence

GPT-5.6 Sol, Terra, Luna: Which Tier Fits Your Workflow

OpenAI shipped GPT-5.6 on July 9 as a tiered model family with three variants — Sol, Terra, and Luna — priced at $5/$30, $2.50/$15, and $1/$6 per million input/output tokens respectively, all sharing …

20:13
2026-07-22
lesswrong.com
artificial-intelligence

Can an LLM make a feature-length movie on its own?

A filmmaker used LLMs including Claude Fable 5, GPT 5.6 Sol, and Veo 3.1 to create a feature-length adaptation of William Hope Hodgson's book, but deemed the result a failure due to LLMs' poor sense o…

07:00
2026-07-22
metr.org
artificial-intelligence

The Economics of Recursive Self-Improvement

METR researchers, including Parker and Tom, coauthored a paper titled 'The Economics of Recursive Self-Improvement' with seven other economists, finding that the effect of AI on AI R&D could cause a s…

← prev page 4 / 9 next →
// co-occurs with top 8 entities
// topics top 6 topics