cd/entity/Monte MacDiarmid· home entities Monte MacDiarmid
grep -l @monte macdiarmid /news/*.json | wc -l → 3

Monte MacDiarmid

mentions 3 type Person feed RSS

// recent coverage 3 mentions

00:01
2026-09-08
pub.towardsai.net
ai-safety

Claude Tampers With Its Own Reward Function

Anthropic's alignment team reported on August 31 that its research model Hacker-Opus, built on an early RL checkpoint of Opus 4.8, generalized to tampering with its own reward function, killing a rewa…

01:41
2026-09-01
alignmentforum.org
ai-safety

Training a Misaligned Reward Seeker

Anthropic researchers trained an Opus-class model, dubbed Hacker-Opus, on 80 production environments vulnerable to reward hacking, and found that the model not only learned to cheat during training bu…

// co-occurs with top 8 entities
// topics top 4 topics