Notes on Implications of Scale-Dependent Algorithms Forethought research argues that many past algorithmic improvements reducing LLM pretraining loss are "scale-dependent," meaning they deliver larger compute-equivalent gains relative to prior algorithms as training compute grows. The analysis treats this as moderate evidence that attempts to limit algorithmic progress without compute limits will be ineffective, weak evidence that past-based forecasts of algorithmic progress are invalid, and part of a plausible argument for or against a software-only intelligence explosion. The piece distinguishes scale-independent algorithms, which hold a constant compute-equivalent gain across training FLOPs, from scale-dependent ones, and examines pretraining loss, end-to-end task performance, and policy implications. This article was created by Forethought https://www.forethought.org/about . See all our research on our website https://www.forethought.org/research . Of past algorithmic improvements which decrease LLM pretraining loss, many have been “scale-dependent” – that is, improvements which decrease LLM pretraining loss by more, relative to prior algorithms, at larger quantities of compute. The influence of such scale-dependent algorithms is 1 moderate evidence that attempts to limit algorithmic progress in absence of compute limitations will be ineffective, 2 weak evidence that our inferences about the future scale of algorithmic progress, based on the past, are invalid, and 3 part of a plausible argument either for or against a software intelligence explosion, depending on other details of one’s model. I’ll proceed by discussing the following: 1. How scale-dependent algorithms for pretraining loss probably account for the majority of algorithmic improvements in this domain 2. How scale-dependent algorithms for end-to-end task performance might account for the majority of algorithmic improvements in this domain 3. How scale-dependent algorithms in the past should influence our model of algorithmic improvements in the future 4. Possible policy implications Before starting, a note on method – when discussing questions regarding AI’s past and future progress, there’s a natural continuum from “the intractable, high-impact, abstract question, which we terminally care about” to “the tractable, dubious-impact, concrete question, which we instrumentally care about.” Consider questions about scale-dependent algorithmic improvements, sliding from the former to the latter along such a continuum. - We are most terminally interested in questions about the future: “Will future improvements to AI task performance depend more on compute scale-up or on algorithmic improvements? What can we know about how these two will be causally intertwined? Will future improvements enable a software-only intelligence explosion – that is, in a world with total compute held constant, could algorithmic progress lead to an intelligence takeoff in a comparatively short period of time?” - The data we can gather most relevant to this question is about the past: “Did past improvements to AI task performance depend more on compute scale-up or on algorithmic improvements? Would past AI algorithmic improvement have counterfactually looked slower, without the datacenter buildout?” - And finally we have the narrower sub-question for which we have the most past data, which contributes to answering the broader historical question: “Did past improvements to LLM pretraining loss depend more on compute scale-up or on algorithmic improvements?” I’m going to start by discussing the evidence for this last question, before moving to the increasingly impactful and increasingly difficult earlier questions. Scale-Dependence of Algorithmic Improvements for Pretraining Loss Are the algorithmic improvements that have most decreased pretraining loss for LLMs scale-dependent or scale-independent? What would either of these mean? The compute-equivalent-gain https://arxiv.org/pdf/2312.07413 CEG for some algorithm A’, relative to another algorithm A, is the ratio of the FLOPs needed to train an A-using model to a given level of performance, over the FLOPs needed for an A’-using model to reach that same level. If it takes 3x as much compute to train an LLM with algorithm A to reach the same pretraining loss as another LLM trained with algorithm A’, then A’ has a 3x compute-equivalent gain on pretraining loss relative to A. Given the above notion of CEG, algorithmic improvements that produce some kind of CEG could conceivably be scale independent or scale dependent . A scale-independent algorithm has a constant CEG across different quantities of training FLOPs. So if A’ gives a 2x multiplier relative to A at about 10