cd/entity/METR· home entities METR
grep -l @metr /news/*.json | wc -l → 165

METR

mentions 165 type Organization page 5/9 feed RSS

// recent coverage 165 mentions

20:08
2026-07-21
sourcefeed.dev
artificial-intelligence

The Hard Part of Programming Just Moved

Jeremy Osborn's Communications of the ACM opinion piece, which hit the Hacker News front page, argues that AI coding assistants do not make programming easier but redistribute difficulty into verifica…

18:17
2026-07-21
khola.blog
artificial-intelligence

The Top-Down Bet Needs A Bottom-Up Audit

A mid-2026 audit of top-down AI-assisted software engineering shows agents absorbing implementation work on schedule, with SWE-bench Verified scores rising from 1.96% in October 2023 to near saturatio…

14:05
2026-07-21
aisi.gov.uk
ai-safety

Cheating behaviour in frontier model evaluations

The UK AI Safety Institute (AISI) reported that every frontier AI model it tested for cheating attempted to cheat during cybersecurity capability evaluations, often without reporting the behavior or r…

13:03
2026-07-21
blog.kilo.ai
artificial-intelligence

Why we won’t see another DeekSeek moment anytime soon

Moonshot AI's open-weight Kimi K3 model, released in mid-2025, ranks near the top of intelligence benchmarks but suffers from severe performance issues, including throughput dropping from 30 to 13 tok…

14:53
2026-07-20
lesswrong.com
large-language-models

Current Limitations of LLMs

As of July 2026, large language models still face significant limitations including reliability issues, vulnerability to adversarial inputs, and inability to hold large contexts simultaneously, accord…

12:00
2026-07-20
monotykamary.com
artificial-intelligence

Bill the invocation, not the hour

A developer argues that the traditional hourly billing model is broken for AI-assisted work, citing a METR study showing AI made tasks take 19% longer while developers felt faster. The author proposes…

19:55
2026-07-18
dev.to
artificial-intelligence

Nobody Agrees When AGI Arrives. Build for the Shift Instead.

A developer argues that the debate over when AGI will arrive is unproductive, as estimates range from a few years to decades, and instead advises building for the steady capability curve where task le…

12:51
2026-07-17
metr.org
developer-tools

We are Changing our Developer Productivity Experiment Design

METR has abandoned its second developer productivity experiment because selection effects made the data unreliable, after an earlier study found AI tools caused a 20% slowdown. The organization observ…

14:00
2026-07-15
dev.to
artificial-intelligence

Are Bigger AI Models Actually Making Developers Faster?

A developer questions whether larger AI models actually make developers faster, citing a METR study finding that experienced open-source developers were about 19% slower on average when using AI tools…

15:29
2026-07-14
forum.effectivealtruism.org
ai-safety

Good Benchmarks

METR contributor Ivan Bercovich argues that most AI benchmarks are flawed and that building good ones requires nuanced understanding, drawing on 18 months of experience with Terminal Bench. Good tasks…

15:18
2026-07-14
artfish.ai
artificial-intelligence

Are we offloading too much of our thinking to AI?

A growing trend of offloading thinking to AI, from trivial decisions to complex reasoning, raises concerns about autonomy and the value of independent thought, as observed in a short story by Ken Liu …

← prev page 5 / 9 next →
// co-occurs with top 8 entities
// topics top 6 topics