23:42
2026-06-25
lesswrong.com
large-language-models
Exploring Generalization in NLA's
A researcher reproduced Anthropic's paper on natural language activations (NLAs), training models to generate textual descriptions of neural network activations. The study found that a single model tr…