Why didn't we get GPT-2 in 2005? Training GPT-2 required roughly 10²¹ FLOPs, a level of compute that the 2004–2008 IBM BlueGene/L supercomputer at Lawrence Livermore National Laboratory could have delivered in about 41 days, according to an analysis of published estimates. BlueGene/L, built by IBM for $290 million, ran at 7×10¹³ FLOPs/sec when first built and 2.8×10¹⁴ FLOPs/sec by November 2005, while GPT-2's training cost is estimated at 8×10²⁰ to 2.5×10²¹ FLOPs across four cited estimates. The comparison shows the compute barrier to a GPT-2-class model existed two decades earlier, leaving data, algorithms and intent as the missing ingredients. Why didn't we get GPT-2 in 2005? The ancient Romans were never great at building ships and never tried to explore the Atlantic. The basic reason seems to be—why bother? The open ocean has no resources and is a vast plane of death. But imagine that in 146 BC after the Romans killed everyone in Carthage, they found a chest in the basement somewhere. The chest was full of gold and labeled “ Gold from Gold Island which is in the middle of the Atlantic somewhere and has lots of gold ”. It’s plausible the Romans would have figured out how to build great ships and gone to find that island, right? Anyway… GPT-2 probably used around 10²¹ FLOPs. In 2018, GPT-2 was the start of large language models breaking through into public consciousness as being impressive. We don’t know how many calculations / FLOPs it took to train it, but here are four estimates: 1. Some internet people claim with no evidence that it took around 2 days to train GPT-2 on 256 TPUv3 chips which if correct would be around 8 × 10²⁰ FLOPs. 2. The Chinchilla FLOPs formula https://dynomight.net/scaling/ what-about-compute says training a model with GPT-2’s parameters and training data should take 1.9 × 10²⁰ FLOPs but this assumes a single pass over the data, and GPT-2 used multiple passes. 3. The Gopher paper trained a model of the same size as GPT-2 but using a 14× larger dataset, and this used around 2.5 × 10²¹ FLOPs. 4. Sevilla et al. https://arxiv.org/abs/2202.05924 say it took around 2.5 × 10²¹ FLOPs which is cool although I wasn’t able to figure out how they arrived at this number. Let's call it 10²¹. Here are the calculations in more detail. 1. Sevilla et al. https://arxiv.org/abs/2202.05924 say it took around 2.5 × 10²¹ FLOPs to train. However, after reading both their paper and this appendix https://docs.google.com/document/d/1J2BX9jkE5nN5EA1zYRN0lHhdCf1YkiFERc nwiYqCOA/ , I still have basically no idea how they came up with this number. 2. GPT-2 had 1.5 billion parameters. The Gopher paper trained a similar 1.5 billion parameter model using, umm, 2.5 × 10²¹ FLOPs. However, that paper used 300 billion tokens, while GPT-2 used only 21 billion. If we assume the cost is linear in the size of the dataset, this suggests you’d need only 1.75 × 10²⁰ FLOPs. 3. If we plug the number of parameters 9 billion and the number of tokens 21 billion into the Chinchilla FLOPs estimate, that suggests you would need 1.9 × 10²⁰ FLOPs. But this assumes just a single pass through the data, while GPT-2 was trained using multiple epochs. 4. It’s published that GPT-2 was trained on 256 cloud TPUv3 chips, but it’s not said for how long. Some random internet people claim that they needed something like 2 days on the 256 TPUv3 cores. A TPUv3 chip is capable of 123 TFLOPs / sec. But typically training on GPUs / TPUs only has a utilization of around 15% . This suggests a total cost of around 256 cores 2 days 123e12 FLOPs/sec 86400 sec / day 0.15 = 8.2 × 10²⁰ FLOPs. BlueGene/L could have trained GPT-2 in a few weeks. From 2004 to 2008, the most powerful computer in the world was BlueGene/L, located at Lawrence Livermore National Laboratory. It was built by IBM for $290 million https://www.cnet.com/tech/tech-industry/ibm-to-build-fastest-supercomputers/ . When it was first built it was capable of 7×10¹³ FLOPs/sec https://www.top500.org/lists/top500/2004/11/highlights/ . It was expanded to 2.8×10¹⁴ FLOPs/sec https://www.top500.org/lists/top500/2005/11/ by Nov 2005, and later doubled again. Here’s what it looked like: To give a sense of progress, you can buy single GPUs today that can almost do that. Anyway, how long would it have taken for this thing to train GPT-2? That’s just division: That's around 41 days. You might worry about utilization percentages but I don’t think this is a serious problem. When modern LLMs are trained, they aren’t actually able to use GPUs to their theoretical capacity—the chips spend a lot of time waiting for data. This means that they’re only actually busy something like 15% of the time. However, the above FLOPs calculations for BlueGene are based on actual achieved performance running a bunch of giant linear algebra functions. But I suppose you might still want to revise the 41 days figure up by some factor. So why didn’t it? There are two obvious answers: 1. Large language models weren’t invented in 2005. We didn’t know about transformer