14:22
2026-07-20
lesswrong.com
large-language-models
Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit
A new approach to mechanistic interpretability traces causal structure in LLM-generated text rather than internal activations, producing attribution graphs that resemble reasoning trajectories. The meβ¦