cd/entity/Natural Language Autoencoders· home entities Natural Language Autoencoders
grep -l @natural language autoencoders /news/*.json | wc -l → 3

Natural Language Autoencoders

mentions 3 type Person feed RSS

// recent coverage 3 mentions

07:04
2026-07-09
lesswrong.com
artificial-intelligence

Interpretability is becoming increasingly uninterpretable

Anthropic's Natural Language Autoencoders (NLAs), a new interpretability method for large language models, use an activation verbalizer and reconstructor to convert activations into natural language a…

00:47
2026-07-09
lesswrong.com
artificial-intelligence

NLAs read thoughts beyond the J-space

Anthropic's research reveals that language models can only verbally report about 10% of their internal activations, confined to a 'J-space' mental workspace. A new experiment using Natural Language Au…

// co-occurs with top 8 entities
// topics top 6 topics