{"slug": "neuron-statistics-notes-on-the-tensor-programs-master-theorem", "title": "Neuron Statistics: Notes on the", "summary": "Greg Yang's Tensor Programs III master theorem proves that in wide neural networks, averages over neurons converge to expectations in a simpler scalar random process, even when random weight matrices are reused, as in weight sharing and backpropagation. The proof conditions on earlier uses of a matrix to separate the next output into a deterministic correction and a fresh Gaussian part, rigorously extending infinite-width analysis to Gaussian-process limits and neural tangent kernels. The theorem, stated for polynomially bounded nonlinearities, establishes convergence of empirical coordinate laws without asserting independence of finite-width coordinates.", "body_md": "**TL;DR.** Tensor programs are a mathematical language for describing\ncomputations in wide neural networks. Their master theorem says that, as the\nwidth grows, averages over the neurons become predictable: they converge to\nexpectations in a much simpler scalar random process. This turns the analysis\nof a high-dimensional random network into a tractable probability calculation\nand provides a rigorous foundation for studying its infinite-width behavior,\nincluding Gaussian-process limits\n[Lee et al. (2018)](https://arxiv.org/abs/1711.00165) and neural tangent\nkernels [Jacot et al. (2018)](https://arxiv.org/abs/1806.07572).\n\nThis note explains the proof for computations that may reuse a random weight matrix and its transpose, as happens with weight sharing and backpropagation. Such reuse creates correlations, so each multiplication cannot simply be replaced by independent Gaussian noise. The proof conditions on all earlier uses of the matrix and separates the next output into a correction determined by those uses and a genuinely fresh Gaussian part. An induction then shows that the scalar process tracks the full computation, including cases in which some intermediate vectors are linearly dependent.\n\nLet and let , independently of . Conditioned on , each coordinate of is Gaussian with variance . This may suggest replacing every matrix product by fresh Gaussian noise, but the next use of the same matrix disproves that idea:\n\nThe first term converges to . If , the second term, conditioned on , is centered Gaussian with variance\n\nand is independent of . Hence\n\nThus in fixed-coordinate distribution, not in vector norm. The -term is precisely the dependence that a fresh-noise approximation misses.\n\nThe program language, scalar -calculus, and master theorem in this section\nare from Section 2 of Yang's *Tensor Programs III*\n[[Yang 2020](https://arxiv.org/abs/2009.10685)]. A\nprogram starts\nfrom random vectors and random matrices\n, and repeatedly applies\n\nHere and every nonlinearity are fixed as . The matrices are independent, with\n\nand the coordinate slices are iid copies of a fixed Gaussian vector .\n\nThe limiting scalar attached to each program vector is defined recursively.\n\nFor , use the corresponding component of , and set\n\nFor a coordinatewise operation,\n\nWrite\n\nFor a fixed oriented matrix , the variables are jointly centered Gaussian, with\n\nFresh Gaussian families belonging to different matrices or orientations (in particular, and ) are independent. Reuse through the opposite orientation is restored by\n\nHere the sum is over earlier products on whose fresh Gaussian variables depends. The derivative is symbolic: first express as a deterministic function of all preceding 's, then differentiate that expression.\n\nFor nonsmooth functions, the expected derivatives can equivalently be defined by Gaussian regression. If are the relevant earlier products, set\n\nThen define the vector of to be . For differentiable functions this agrees with Stein's identity\n\nand its multivariate version.\n\nA function is polynomially bounded if for some . If all program nonlinearities are polynomially bounded, then for any program vectors and any polynomially bounded test function ,\n\nThe theorem concerns empirical coordinate laws; it does not assert that the finite-width coordinates are independent.\n\nThis is the conditioning step used in the proof of the master theorem\n(Appendices K and L of\n[Yang 2020](https://arxiv.org/abs/2009.10685)).\nOrder the initial and matrix-generated vectors (the G-vars) as\n. Suppose the next one is\n\nwhere . Collect all earlier uses of in both orientations:\n\nand form\n\nEmpty matrices are allowed when or .\n\nLet . Conditioned on , all the displayed vectors are fixed and is subject to\n\nGaussian conditioning under these linear constraints gives the exact law\n\nwhere is an independent copy of , , , and\n\nMultiplying by yields\n\nwith independent of and\n\nThis separates the part forced by previous uses of from the remaining Gaussian randomness.\n\nIntroduce the finite Gram data\n\nMoment convergence for the preceding G-vars gives deterministic limits, denoted with bars; for example,\n\nRank stability, proved in the next section, implies and almost surely. Therefore\n\nIndeed, the expression in parentheses is the squared -distance from to .\n\nUsing Equation (5), define\n\nThen\n\nConsequently, with\n\nwe have\n\nSet . By Equation (1),\n\nGaussian regression therefore gives, conditionally on all earlier -variables,\n\nThe mean and variance here are exactly and Equation (8).\n\nApplying the ZDot rule, Equation (2), with Stein's identity or its pseudoinverse form, gives the matching drift identity\n\nAdding Equation (11) and Equation (12) proves the key scalar conditional law\n\nconditionally on the earlier -variables. For , this reduces to , recovering the opening example.\n\nThe core-set and rank-stability mechanism below follows Appendix L of\n[Yang (2020)](https://arxiv.org/abs/2009.10685). The proof simultaneously\nmaintains two induction statements.\n\nFor every polynomially bounded ,\n\nThere is such that:\n\nThe last property upgrades almost-everywhere limiting identities to exact finite-width identities at every coordinate.\n\nTo see the consequence, let be polynomially bounded functions of the preceding G-vars and let\n\nIf , then\n\nThus the linear combination vanishes almost everywhere under the core law. The density and null-avoidance properties make it vanish at every finite-width coordinate for all large , so . Ordinary lower semicontinuity of rank gives the reverse inclusion. Therefore, eventually,\n\nThe pseudoinverse is continuous on a fixed-rank stratum, proving the pseudoinverse convergence used above.\n\nReturn to . By Equation (8),\n\nThere are two cases.\n\nIf , fixed coefficients satisfy\n\nExpressing and each as coordinatewise functions of the old core turns this into an almost-everywhere functional identity. Core density and null avoidance make it exact at finite width:\n\nalmost surely for all sufficiently large . Hence is already in the old span and the core need not change.\n\nIf , Equation (13) has the form\n\nConditioned on the old core, has an everywhere-positive Gaussian density. Fubini's theorem then shows that adjoining preserves the two-way null-set property, so the new core is .\n\nAt finite width, . Moreover, rank stability and old-core null avoidance imply\n\nThus every , conditionally on the past, retains positive Gaussian variance. Applying Fubini to each prescribed null set and then old-core null avoidance proves the new null-avoidance property. The same fresh Gaussian component prevents from lying in the old span. This completes the core-set update in both cases.\n\nThe final concentration argument combines the projection estimates and\nweakly correlated Gaussian strong law from Appendices K and L of\n[Yang (2020)](https://arxiv.org/abs/2009.10685). It remains to prove\n. When , the exact relation above\nmakes a fixed polynomially bounded function of earlier G-vars, so\napplies directly. Assume henceforth that\n.\n\nFrom Equation (6),\n\nLet\n\nand define the one-dimensional Gaussian average\n\nThe difference in the [moment invariant](https://www.lesswrong.com/feed.xml#moment-invariant) is bounded by\n, where\n\nConditioned on , the second term inside braces is the expectation of the first. Although the coordinates of are correlated, the dependence has fixed rank. Indeed, if , then the normalized off-diagonal correlations of satisfy\n\nA strong law for such weakly correlated Gaussian arrays applies. Its required high moments follow from polynomial boundedness, , the convergence of , and . Therefore\n\nThe projection diagonal satisfies\n\nHence at most coordinates have ; their contribution vanishes by Hölder's inequality and the same polynomial moment bounds.\n\nOn the remaining coordinates, both variances stay bounded away from zero because . Gaussian convolution smooths even a merely measurable polynomially bounded , with\n\nThese derivatives are bounded by polynomial weights in the preceding\ncoordinates and . Meanwhile, Equation (10), Equation (8), and\n[Equation (17)](https://www.lesswrong.com/feed.xml#term-b-replace-finite-parameters-by-their-limits) imply\n\nThe mean-value bound followed by Cauchy--Schwarz now gives\n\nEvery and is a polynomially bounded coordinatewise function of earlier G-vars. Consequently\n\nis polynomially bounded. Applying , then using the scalar identification in Equation (13), gives\n\nThus , proving .\n\nFor the initial Gaussian vectors, the ordinary strong law proves the moment invariant. Choose a maximal linearly independent subfamily of their limiting Gaussian variables as the first core. Its covariance is positive definite, so its law has an everywhere-positive density. Coordinatewise linear relations and countable null-set avoidance prove the other core properties.\n\nThe preceding sections establish the simultaneous step\n\nInduction therefore proves the moment statement for all G-vars.\n\nFinally, every program vector is a fixed polynomially bounded coordinatewise function of the complete list of G-vars:\n\nThe composition of polynomially bounded functions is polynomially bounded. Applying the G-var moment result to gives Equation (3), completing the proof.", "url": "https://wpnews.pro/news/neuron-statistics-notes-on-the-tensor-programs-master-theorem", "canonical_source": "https://www.lesswrong.com/posts/u9fC4DyYiptdk6yCb/neuron-statistics-notes-on-the-tensor-programs-master", "published_at": "2026-08-02 22:42:30+00:00", "updated_at": "2026-08-02 22:59:01.995094+00:00", "lang": "en", "topics": ["machine-learning", "neural-networks", "ai-research"], "entities": ["Greg Yang", "Tensor Programs III", "Lee et al. (2018)", "Jacot et al. (2018)"], "alternates": {"html": "https://wpnews.pro/news/neuron-statistics-notes-on-the-tensor-programs-master-theorem", "markdown": "https://wpnews.pro/news/neuron-statistics-notes-on-the-tensor-programs-master-theorem.md", "text": "https://wpnews.pro/news/neuron-statistics-notes-on-the-tensor-programs-master-theorem.txt", "jsonld": "https://wpnews.pro/news/neuron-statistics-notes-on-the-tensor-programs-master-theorem.jsonld"}}