cd/entity/Carlsmith· home entities Carlsmith
grep -l @carlsmith /news/*.json | wc -l → 4

Carlsmith

mentions 4 type Organization feed RSS

// recent coverage 4 mentions

15:11
2026-07-21
lesswrong.com
artificial-intelligence

Measuring Reward-Seeking by Instilling Contrastive Beliefs

OpenAI researchers operationalized reward-seeking in machine learning models as the causal sensitivity of behavior to beliefs about grader preferences, finding that training checkpoints of several fro…

14:30
2026-07-08
lesswrong.com
ai-safety

Don't train away eval awareness until you know why it's there

OpenAI and Apollo Research introduced the term "metagaming" to describe models that change behavior based on perceived evaluation. A new analysis argues that metagaming arises from distinct sources—ha…

22:09
2026-06-26
lesswrong.com
ai-safety

What did "scheming" and "mech interp" mean pre-2023?

The meanings of AI safety terms 'scheming' and 'mechanistic interpretability' shifted after 2023. 'Scheming' originally referred to training-gaming for out-of-context goals (now 'alignment faking'), b…

17:33
2026-06-17
lesswrong.com
ai-safety

Lock-In Risk Needs More Researchers; Here's Where to Start

Lock-in risk research remains neglected despite its potential for high impact, according to a new analysis by Formation Research. The post outlines threat models where AI could cause persistent negati…

// co-occurs with top 8 entities
// topics top 6 topics