02:40
2026-07-20
lesswrong.com
large-language-models
Is there even a ground-truth for LLMsโ internal representations?
A new blog post argues that existing methods for decoding internal representations in large language models (LLMs), such as Logit Lens, Tuned Lens, Patchscopes, and Anthropic's Jacobian Lens, are all โฆ