Where Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch A new arXiv paper (2609.17571v1) reports that grokking in Transformers is a spectral recoding of an existing distributed circuit rather than a switch between modules, based on behavior-aligned exact activation games the authors call Transition Games. The study found distributed utility gain with a prospective block-0 attention bias, with selected degree-two modes accounting for 67–92% of its addition contrast across replacement games, and a disjoint exact path study confirming block-1 MLP mediates more of the effect than all other tested downstream paths in 12/12 pairs. The sharper "MLP memorizes, attention generalizes" prediction reversed (-.331 bits/example at the memory anchor; 0/12 in the predicted direction), while routing onset, global rank collapse, and a prime-invariant architecture ridge also failed. arXiv:2609.17571v1 Announce Type: new Abstract: Where in a Transformer is the change from memorization to generalization functionally expressed? We introduce Transition Games--behavior-aligned exact activation games with paired non-generalizing controls--and find distributed utility gain with a prospective block-0 attention bias; selected degree-two modes account for 67--92% of its addition contrast across replacement games, and a disjoint exact path study confirms that block-1 MLP mediates more of their effect than all other tested downstream paths in 12/12 pairs. The sharper "MLP memorizes, attention generalizes" prediction instead reverses -.331 bits/example at the memory anchor; 0/12 in the predicted direction , while routing onset, global rank collapse, and a prime-invariant architecture ridge also fail, identifying grokking here as spectral recoding of an existing distributed circuit rather than a module switch.