00:57
2026-07-24
lesswrong.com
ai-research
Fixing rewards for NLA to reduce confabulation
A researcher testing Anthropic's Natural Language Autoencoder (NLA) found that improving reconstruction fidelity does not guarantee faithful interpretation of a language model's internal activations. โฆ