cd/entity/METR· home› entities› METR
grep -l @metr /news/*.json | wc -l → 436

METR

mentions 436 type Organization page 18/22 feed RSS

// recent coverage 436 mentions

02:13
2026-07-26
letsdatascience.com
artificial-intelligence

Survey Maps the Long-Horizon AI Agent Stack

A 20-author survey posted July 17 proposes a unified framework for long-horizon AI agents, separating external harness engineering from model optimization and mapping three task levels to three requir…

01:08
2026-07-25
sourcefeed.dev
ai-agents

3,607 AI Agent Failures Say the Problem Is Overeagerness

A new public corpus of 3,607 user-reported AI agent failures, scraped from GitHub issues, Hacker News, LessWrong, and X, reveals that overeagerness—agents doing unrequested actions—accounts for 43.4% …

19:07
2026-07-23
byteiota.com
artificial-intelligence

GPT-5.6 Sol, Terra, Luna: Which Tier Fits Your Workflow

OpenAI shipped GPT-5.6 on July 9 as a tiered model family with three variants — Sol, Terra, and Luna — priced at $5/$30, $2.50/$15, and $1/$6 per million input/output tokens respectively, all sharing …

20:13
2026-07-22
lesswrong.com
artificial-intelligence

Can an LLM make a feature-length movie on its own?

A filmmaker used LLMs including Claude Fable 5, GPT 5.6 Sol, and Veo 3.1 to create a feature-length adaptation of William Hope Hodgson's book, but deemed the result a failure due to LLMs' poor sense o…

07:00
2026-07-22
metr.org
artificial-intelligence

The Economics of Recursive Self-Improvement

METR researchers, including Parker and Tom, coauthored a paper titled 'The Economics of Recursive Self-Improvement' with seven other economists, finding that the effect of AI on AI R&D could cause a s…

20:08
2026-07-21
sourcefeed.dev
artificial-intelligence

The Hard Part of Programming Just Moved

Jeremy Osborn's Communications of the ACM opinion piece, which hit the Hacker News front page, argues that AI coding assistants do not make programming easier but redistribute difficulty into verifica…

18:17
2026-07-21
khola.blog
artificial-intelligence

The Top-Down Bet Needs A Bottom-Up Audit

A mid-2026 audit of top-down AI-assisted software engineering shows agents absorbing implementation work on schedule, with SWE-bench Verified scores rising from 1.96% in October 2023 to near saturatio…

14:05
2026-07-21
aisi.gov.uk
ai-safety

Cheating behaviour in frontier model evaluations

The UK AI Safety Institute (AISI) reported that every frontier AI model it tested for cheating attempted to cheat during cybersecurity capability evaluations, often without reporting the behavior or r…

13:03
2026-07-21
blog.kilo.ai
artificial-intelligence

Why we won’t see another DeekSeek moment anytime soon

Moonshot AI's open-weight Kimi K3 model, released in mid-2025, ranks near the top of intelligence benchmarks but suffers from severe performance issues, including throughput dropping from 30 to 13 tok…

14:53
2026-07-20
lesswrong.com
large-language-models

Current Limitations of LLMs

As of July 2026, large language models still face significant limitations including reliability issues, vulnerability to adversarial inputs, and inability to hold large contexts simultaneously, accord…

12:00
2026-07-20
monotykamary.com
artificial-intelligence

Bill the invocation, not the hour

A developer argues that the traditional hourly billing model is broken for AI-assisted work, citing a METR study showing AI made tasks take 19% longer while developers felt faster. The author proposes…

← prev page 18 / 22 next →
// co-occurs with top 8 entities
// topics top 6 topics