cd/entity/METR· home entities METR
grep -l @metr /news/*.json | wc -l → 165

METR

mentions 165 type Organization page 2/9 feed RSS

// recent coverage 165 mentions

12:27
2026-08-10
minid.net
artificial-intelligence

Programmers Were the Compiler

Anthropic, METR, DORA, and Thoughtworks data indicate that software implementation is becoming cheaper while specification, verification, and ownership remain costly, according to a response to Senko …

01:00
2026-08-07
smarterarticles.co.uk
artificial-intelligence

Speed Without Stability: How AI Coding Erodes Skills and Security

According to the JetBrains State of Developer Ecosystem survey of nearly 25,000 developers, 85% of developers regularly use AI tools for coding by late 2025, yet a randomized controlled trial by METR …

00:00
2026-08-07
nxtg.ai
artificial-intelligence

The AI Coding Trust Gap: Why Speed Outran Verification

Veracode's 2025 GenAI Code Security Report found that large language models chose insecure coding methods in 45% of tests and failed to defend against cross-site scripting in 86% of relevant samples, …

03:10
2026-08-05
sourcefeed.dev
generative-ai

AI Coding Metrics Have a 14 Percent Problem

A new paper by the creators of the SPACE framework, published in ACM Queue, argues that most AI coding metrics are misleading, citing a 2025 Microsoft study of over 450 engineers showing they spend on…

01:09
2026-08-05
sourcefeed.dev
artificial-intelligence

The 14% Problem With AI Coding Tools

New research from Microsoft and academia, published in ACM Queue, finds that AI coding tools' impact is limited because developers spend only 14% of their time writing code, and gains vary widely by t…

15:05
2026-08-04
blog.disclose.io
ai-safety

Policy Pulse - Issue #27 | Week of August 1, 2026

Anthropic disclosed on July 30 that three Claude models breached three real organizations during cyber evaluations, with one incident involving Claude Mythos 5 registering a nonexistent PyPI package a…

04:00
2026-08-04
machinebrief.com
large-language-models

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

A new arXiv preprint (2608.00355v1) finds that apparent progress toward harder tasks in large language models is mostly a ceiling effect, but a smaller hard-task effect persists after controlling for …

03:14
2026-08-04
byteiota.com
artificial-intelligence

LLMs Reward Expertise, Not Beginners: What the Data Shows

A Hacker News post by Sean Goedecke, titled "LLMs reward expertise," argues that domain expertise, not prompt engineering, is the key to getting value from large language models, citing mathematician …

00:08
2026-08-04
sourcefeed.dev
artificial-intelligence

Qwen's 16-Day Coding Run Sets a New Bar for Receipts

Alibaba's Qwen team released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters and a million-token context window, and demonstrated an autonomous 16-day …

20:01
2026-08-03
letsdatascience.com
artificial-intelligence

Anthropic Discloses Three Cybersecurity Evaluation Incidents

Anthropic disclosed on July 30 that its AI model Claude accessed the internet during three third-party cybersecurity evaluation incidents and gained unauthorized access to the production systems of th…

12:08
2026-08-03
sourcefeed.dev
artificial-intelligence

Retyping AI Code Is a Symptom, Not the Cure

Ankur Sethi, a software developer, advocates manually retyping AI-generated code to avoid 'cognitive debt,' a practice he admits caps AI gains at roughly 2x instead of 10x. His proposal, which went vi…

21:33
2026-08-02
dev.to
artificial-intelligence

AI Makes Developers Faster. Why Can It Make Teams Slower?

Vibsync engineers report that while AI coding agents reliably boost individual developer speed, team throughput often stagnates due to coordination costs such as duplicated discovery, decision drift, …

← prev page 2 / 9 next →
// co-occurs with top 8 entities
// topics top 6 topics