{"slug": "do-transformer-representations-progressively-structure-across-depth-and-time-8", "title": "Do Transformer representations progressively structure across depth and time? Results from 8 open models", "summary": "A new preprint analyzing 8 open Transformer models finds that internal representations show statistically significant structured progression across depth (p = 0.00019996, 8/8 leave-one-model-out checks) and a coherent cross-model depth profile (mean r = 0.789), but no universal temporal trajectory, as a common temporal pattern failed to replicate (0/8 checks). The study, 'Progressive Representational Structuring in Small Language Models: Functionally Labelled Trajectories Across Depth and Time' (DOI: 10.5281/zenodo.22116637), also found no significant structural similarity among models with similar functional outcomes (p = 0.408) or same architecture family (p = 0.771).", "body_md": "Hi everyone,\n\nI’ve just published a new preprint that brings together several months of experiments on hidden-state dynamics in small open Transformer models.\n\nThe question is fairly simple:\n\nDuring inference, do internal representations simply change from layer to layer, or is there evidence of a more structured progression across depth and generation time?\n\nI tried to study this without assuming that hidden-state dynamics are equivalent to “reasoning”.\n\nThe working framework is:\n\ntokens → embeddings → contextualisation → relational structuring → functional structuring → decision formation → projection\n\nThis is a descriptive hypothesis about representation dynamics, not a claim that these stages correspond to a universal reasoning mechanism.\n\nThe expanded study uses 8 locally instrumented open models, with synchronized hidden-state and output observations and explicit separation between:\n\ndepth — what changes as information passes through Transformer layers\n\ntime — what changes as autoregressive generation progresses\n\nA few results were particularly interesting.\n\nFirst, local ordering across model depth survived expansion.\n\nThe observed ordering was significantly more structured than random layer permutations (`p = 0.00019996`\n\n) and remained supported when each model was removed from the panel one at a time (**8/8 leave-one-model-out checks**).\n\nSecond, cross-model depth profiles remained surprisingly coherent.\n\nThe mean correlation across normalized depth profiles was approximately r = 0.789.\n\nThis does not mean that all models follow the same trajectory. Rather, it suggests that some aspects of where changes occur along depth may be more shared than I initially expected.\n\nThird, functionally labelled events were not uniformly distributed across depth.\n\nEvent type showed a statistically supported association with normalized layer depth (`p = 0.0024`\n\n).\n\nI’m deliberately calling this an association, not evidence of a causal mechanism.\n\nBut one of the most useful results was actually a failure to replicate.\n\nIn an earlier smaller panel, a common temporal pattern in local trajectory instability looked promising. After expanding the panel, that common temporal mode disappeared — it survived 0/8 leave-one-model-out checks.\n\nTwo other intuitive hypotheses also failed:\n\nmodels with similar observed functional outcomes were not significantly more structurally similar (`p = 0.408`\n\n), and models from the same architecture family were not significantly more similar either (`p = 0.771`\n\n).\n\nTo me, this is probably the most important part of the result.\n\nThe data do not support a simple story where architecture determines one characteristic trajectory or where one universal temporal dynamic explains inference.\n\nWhat remains is a narrower hypothesis:\n\nTransformer inference may contain reproducible structure along depth while remaining highly conditional in time and behavior.\n\nI refer to this as Progressive Representational Structuring.\n\nThe framework is summarized by:\n\nRepresentation ≠ Function ≠ Behavior\n\nA representation can contain information without that information yet serving the same function, and a functional transition does not guarantee a particular final behavior.\n\nI would be especially interested in feedback from people working on:\n\nmechanistic interpretability, activation patching, probing, hidden-state geometry, steering, representation engineering, or larger open models.\n\nIn particular, I’m curious whether others observe similar **ordered depth structure without a universal temporal trajectory.\n\nPreprint:\n\n**Progressive Representational Structuring in Small Language Models: Functionally Labelled Trajectories Across Depth and Time**\n\nDOI: 10.5281/zenodo.22116637\n\nThis is still descriptive work. Causal intervention and structural-transfer experiments are separate next steps rather than claims of this paper. [Progressive Representational Structuring in Small Language Models: Functionally Labelled Trajectories Across Depth and Time | Zenodo](https://doi.org/10.5281/zenodo.22116637)", "url": "https://wpnews.pro/news/do-transformer-representations-progressively-structure-across-depth-and-time-8", "canonical_source": "https://discuss.huggingface.co/t/do-transformer-representations-progressively-structure-across-depth-and-time-results-from-8-open-models/179289#post_1", "published_at": "2026-08-26 18:52:16+00:00", "updated_at": "2026-08-26 19:13:21.502504+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["Zenodo"], "alternates": {"html": "https://wpnews.pro/news/do-transformer-representations-progressively-structure-across-depth-and-time-8", "markdown": "https://wpnews.pro/news/do-transformer-representations-progressively-structure-across-depth-and-time-8.md", "text": "https://wpnews.pro/news/do-transformer-representations-progressively-structure-across-depth-and-time-8.txt", "jsonld": "https://wpnews.pro/news/do-transformer-representations-progressively-structure-across-depth-and-time-8.jsonld"}}