Do Transformer representations progressively structure across depth and time? Results from 8 open models
A new preprint analyzing 8 open Transformer models finds that internal representations show statistically significant structured progression across depth (p = 0.00019996, 8/8 leave-one-model-out check…