cd/entity/Ryan Greenblatt· home entities Ryan Greenblatt
grep -l @ryan greenblatt /news/*.json | wc -l → 14

Ryan Greenblatt

mentions 14 type Person feed RSS

// recent coverage 14 mentions

12:31
2026-07-24
lesswrong.com
artificial-intelligence

Does distilling Claude carry the persona with it?

A systematic identity-swap experiment by researcher Benji Brczy shows that GLM 5.2 has a selectable Claude persona that changes its safety and behavioral profile, while Kimi K3 does not adopt Claude's…

21:30
2026-07-16
latent.space
ai-policy

Loopcraft: The Art of Stacking Loops

Anthropic reversed its policy of covertly degrading Claude Fable 5 for AI-research-related use cases within roughly a day after public backlash, with critics arguing that opaque sandbagging violates t…

20:36
2026-07-11
thezvi.wordpress.com
ai-safety

Introduction for and Reactions to Plan A

The creators of the AI 2027 predictions, including Daniel Kokotajlo and Ryan Greenblatt, have released a new positive vision called Plan A, which proposes slowing AI development through a deal with Ch…

09:07
2026-07-10
forum.effectivealtruism.org
artificial-intelligence

AI 2040: Plan A [thread]

The team behind AI 2027 released a new scenario called AI 2040: Plan A, offering a positive vision for navigating the creation of super-intelligence. The named authors include Thomas Larsen, Romeo Dea…

18:29
2026-07-07
lesswrong.com
large-language-models

Superhuman Articulacy as an LLM Safety Target

Large language models exhibit poor articulacy in technical communication, including jargon creation, inconsistent terminology, verbosity, and inappropriate shorthand, which poses safety risks as their…

00:28
2026-06-29
lesswrong.com
ai-safety

Third-parties should focus on scrutinising system cards

Anthropic's system cards, which disclose AI risks, are expected to degrade over time due to increasing model complexity, rushed AI-generated content, and stronger incentives for labs to mislead. Third…

// co-occurs with top 8 entities
// topics top 6 topics