20:08
2026-07-24
lesswrong.com
ai-safety
Intent Is All You Need.
An anonymous researcher claims to have developed a full-stack interpretability suite for large language models, including a replication of the Arditi et al refusal direction research, using only a freβ¦