cd/entity/Lesswrongยท homeโ€บ entitiesโ€บ Lesswrong
grep -l @lesswrong /news/*.json | wc -l โ†’ 1

Lesswrong

mentions 1 type Organization feed RSS

// recent coverage 1 mentions

00:57
2026-07-24
lesswrong.com
ai-research

Fixing rewards for NLA to reduce confabulation

A researcher testing Anthropic's Natural Language Autoencoder (NLA) found that improving reconstruction fidelity does not guarantee faithful interpretation of a language model's internal activations. โ€ฆ

// co-occurs with top 4 entities
// topics top 2 topics