Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit
A new approach to mechanistic interpretability traces causal structure in LLM-generated text rather than internal activations, producing attribution graphs that resemble reasoning trajectories. The me…