Introducing PIRAMID: Physics-Informed Research for Ambitious Mechanistic Interpretability Principles of Intelligence (PrincInt, formerly PIBBSS) is launching PIRAMID, an internal research division that uses statistical physics to build scientific foundations for ambitious mechanistic interpretability. The division comprises three teams led by Dmitry Vaintrob, Andrew Mack, and Ari Brill, focusing on learning theory, interpretability applications, and data models. PIRAMID aims to develop interpretability tools grounded in scientific principles to enable scalable alignment of AI systems. Principles of Intelligence PrincInt, formerly PIBBSS is launching PIRAMID https://princint.ai/piramid-project/ , an internal research division using the tools and techniques of statistical physics to build scientific foundations for ambitious mechanistic interpretability. PIRAMID’s central premise is that scalable alignment will require more than persuasive ad-hoc explanations of model behavior. It will require interpretability tools that develop alongside a scientific understanding of the structure of data, learning, and representations. To reflect this, we divide our attention across three synergistic research teams: Advancements in Learning Theory led by Dmitry Vaintrob , Interpretability Applications led by Andrew Mack , and Data Models and Validation Methods led by Ari Brill . Together, they form a loop: theory predicts how structure can be learned and organized in networks, interpretability tools built on these principles help us recover and intervene on that structure, and synthetic datasets with built-in ground truth provide settings in which both theory and tools can be validated. We can think of this as loosely mirroring physics' methodological division of labor, with each group prioritizing theory, empirics, and phenomenology, respectively. This methodological coverage helps to build up a scientific understanding of real-world neural networks that narrows the theory-practice gap. PIRAMID is part of PrincInt’s larger field-building efforts. Over the past year and a half, we hired a cohort of affiliate researchers to test candidate directions several of whom went on to form the PIRAMID leadership team , started a series of workshops https://www.lesswrong.com/posts/53yqropw9DCtd9e7s/statistical-physics-for-ambitious-interpretability-a connecting statistical physics with AI interpretability, and began incubating new academic research groups as part of PIAMI Physics-Informed Ambitious Mechanistic Interpretability , a coordinated research program of which PIRAMID is one working group. PIAMI is how we plan to stay connected to the communities of expertise this work draws on -- across physics, learning theory, and interpretability -- which we see as a deep well of ideas, heuristics, and talent for building scientific foundations in AI safety. We intend to share our thinking with those communities early and often in posts like this one, and are particularly excited by potential synergies with other theory-informed research agendas e.g., Simplex, Timaeus/Resolution, ARC and academic groups e.g., learning mechanics https://learningmechanics.pub/ . We'll have more to say about this program soon. PrincInt’s goal is to develop the foundations to support scalable alignment. While ambitious interpretability https://www.alignmentforum.org/posts/Hy6PX43HGgmfiTaKu/an-ambitious-vision-for-interpretability – fully reverse engineering an AI system – is neither necessary nor sufficient for this, we see it as a proxy for the kind of faithful mechanistic transparency that would make scalable alignment more feasible. A faithful explanation must track the mechanism the model actually learns and uses, not just correlate with behavior. Future systems may differ from today's models in terms of architecture, continual learning, or memory. A method tied only to today's empirical probes may fail when the model generalizes in a new way, undergoes fine-tuning, learns hidden strategies, or moves into a regime where our probes no longer behave as expected. However, one grounded in broader principles governing learning and computation has a better chance of transferring. In pursuit of that goal, PIRAMID’s operative target is to make AI systems sufficiently transparent to support high-confidence, faithful statements about their internal computations. Physics-informed methods can guide us toward principled definitions of faithfulness and a structural understanding of what a network learns that is grounded in the data, training dynamics, and what is learned e.g., representational geometry 1 . Concretely, we think progress on ambitious interpretability requires answering the three questions central to PIRAMID’s research program that are often studied separately: The presence of structurally relevant randomness – for which statistical physics is the canonical framework – is a common thread across PIRAMID’s research groups. Neural networks are stochastic objects, with fluctuations that come from, for example, initialization, data randomness, the choice of optimizer, and the measurement tools themselves. Every attempt to interpret a neural network implicitly treats some piece of its structure as signal rather than noise; analyses and tools that cannot make principled, supported statements accounting for this randomness are unlikely to be faithful, whatever their benchmark scores. Statistical physics gives us the language to formalize – and the tools to track – the mechanistic role of this randomness. Instead of accounting for every microscopic detail, physicists often identify the variables that matter most at a particular scale . These are Advancements in Learning Theory. Our core hypothesis – implicit in other statistical approaches to learning theory – is that the microscopic state of a trained network is too irregular to reason about directly, while coarse loss or benchmark performance measures obscure the structure we care about. Instead, we consider aggregate statistical measures of neural network distributions as the correct level of abstraction: coarse enough to theoretically describe regular structure, fine enough to uncover mechanistic detail that can guide interpretability work. Statistical and geometric properties of learning we consider may include training time order parameters, error correlations, the dynamics of feature subspaces, and other geometric properties across data or in weight space. Concretely, we aim to produce an end-to-end toy-to-real case study in which a theoretically motivated statistic predicts and explains generalization or capability-relevant change within the next few months. Developing a comprehensive statistical theory of feature learning in neural networks is hindered, in large part, by the gap between tractable idealizations and the messy reality of learned representations. This group spends a significant amount of time thinking about how to build theories that are useful . Theory helps identify which structures matter for interpretability, and a driving goal of the group is to define a model-natural, statistical notion of "feature" that can anchor both theory and tools. The aim is not to declare one current object — neurons, SAE features, kernels, or circuits — to be the fundamental unit of neural computation, but to identify where existing theories break, construct cleaner toy settings that expose those breaks, and use those failures to discover better theoretical objects\footnote{This approach has often proven useful. For example, the inability to probe polysemantic feature structure led to the development of SAEs and compressed sensing methods, and the inability of older large-N limit methods e.g., Roberts et al. to explain deep compositional behaviors and generalization on complex tasks led to more sophisticated frameworks including mean field and dynamical mean field methods. }. Interpretability Applications. The goal of this group is to develop the tools – interpretable-from-scratch architectures, principled feature discovery methods, and mechanistic techniques for eliciting and steering model behaviors – that recover mechanistic structure that reflects the causal and hierarchical nature of learning and computation. Though many post-hoc interpretability methods assume some neural representation hypothesis is true, there are few examples that are derived or tested against a principled theory of data, learning, or computation. Our mission for the next year is to develop a suite of interpretability tools and methods that are empirically and theoretically grounded, computationally tractable at scale, and make meaningful gains in alignment-relevant applications. These tools are physics-informed in the sense of being grounded in theoretical hypotheses about how networks learn structure across scales: which features, directions, circuits, or basins are relevant at a given level of description, and how those levels interact. To ensure we’re actually making progress, we will validate progress on pragmatic downstream tasks e.g., data attribution, backdoor detection, sandbagging, alignment faking as well as intermediate measures of faithfulness defined by our physics-informed approaches. Interpretability applications turn theoretical hypotheses into tools for analyzing and steering real systems. Tool failures, in turn, reveal where theories or data models are incomplete. For example, if a method only recovers nonlinear or distributed structure when we expected clean hierarchical features, that mismatch becomes evidence about what the network actually learned and what our validation setup failed to capture. Data Models and Validation Methods. Mechanistic interpretability currently lacks rigorous benchmarks for validating interpretability tools, in part because of the large gap between tractable toy setups that model data and its illegible ground-truth structure. We aim to address this gap by i constructing analytically tractable, physics-inspired data models that capture key aspects of natural data and ii quantifying the relationship between this data structure and learned features. Our current considerations for a theory of natural data include hierarchy, sparsity, criticality, and power-law statistics. Our core hypothesis is that synthetic datasets generated by models that qualitatively and quantitatively capture a semantically relevant, ground-truth feature hierarchy of natural data can be used to validate theories of feature learning and interpretability tools. Because the latent structure is known by construction, we can test whether theory predicts, or tools recover, the structure the model actually learns and at what scale , rather than a plausible but unfaithful explanation. Conversely, empirical anomalies will provide phenomenological signals about how to reason about realistic data structure. Within the next year, we aim to turn these datasets into a public benchmark with stronger faithfulness guarantees than current evaluation practice. Recent Work for the theory team includes posts distilling mean-field theory for a broader audience https://www.lesswrong.com/posts/rduzFkTKx5pGKWKcL/mean-field-sequence-an-introduction , a preprint https://arxiv.org/abs/2607.05735 on learnability in mean-field Bayesian networks, and a post connecting mean field theory with computation in superposition https://www.lesswrong.com/posts/siu22scEfuKxpSgfK/a-tale-of-three-theories-sparsity-frustration-and . Work from the tools team includes feature identification with the empirical NTK https://openreview.net/forum?id=XVOMFslJOK , an update to MELBO https://arxiv.org/abs/2606.29604 , and a preprint detailing an interpretable-by-design architecture https://arxiv.org/abs/2607.20652 built on approximately orthogonal hashed feature vectors. Work from the data models team includes two https://openreview.net/forum?id=9mAX9GZK5e papers https://arxiv.org/abs/2606.20347 on critical percolation as a synthetic data model for interpretability and accompanying code to generate datasets https://github.com/aribrill/percolation-synthetic-data and train models https://github.com/tomingebretsencarlson/mechinterp-percolation/tree/main . PIRAMID's work has been supported by grants from the UK AI Security Institute, the Long-Term Future Fund, and Coefficient Giving. We have plans to grow — if you'd like to be informed of future events on this topic or support our work as we expand, please get in touch. We’ll also be hiring – keep an eye out for our open roles here https://princint.ai/about/careers/ . We are not claiming novelty in making this statement, and are excited to add our perspectives to other groups who build on this claim, either implicitly or explicitly e.g., Simplex, Timaeus, and many, many academic groups .