NanoGPT Speedrun The NanoGPT speedrun repository reports that a collaborative effort has trained a language model to 3.28 cross-entropy loss on the FineWeb validation set in under 75 seconds on 8 NVIDIA H100 GPUs, a dramatic improvement over Andrej Karpathy's llm.c GPT-2 replication which required 45 minutes and 10B tokens. The speedup, achieved through techniques like rotary embeddings, QK-Norm, ReLU², the Muon optimizer, FP8 precision, and Flash Attention 3, reduces token usage to under 400M. This repository hosts the NanoGPT speedrun , in which we collaboratively|competitively search for the fastest algorithm to use 8 NVIDIA H100 GPUs to train a language model that attains 3.28 cross-entropy loss on the FineWeb https://huggingface.co/datasets/HuggingFaceFW/fineweb validation set. Note: Besides the main track, there is also an optimization track /KellerJordan/modded-nanogpt/blob/master/records/track 3 optimization where we try to minimize steps subject to fixed arch/data/bsz and with unlimited wallclock budget. The target 3.28 validation loss on FineWeb follows Andrej Karpathy's GPT-2 replication in llm.c, which attains that loss after running for 45 minutes https://github.com/karpathy/llm.c/discussions/481 :~:text=By%20the%20end%20of%20the%20optimization%20we%27ll%20get%20to%20about%203.29 . The speedrun code also descends from llm.c's PyTorch trainer https://github.com/karpathy/llm.c/blob/master/train gpt2.py , which itself descends from NanoGPT, hence the name of the repo. Thanks to the efforts of many contributors, this repo now contains a training algorithm which attains the target performance in: - Under 75 seconds on 8xH100 the llm.c GPT-2 replication needed 45 minutes - under 400M tokens the llm.c GPT-2 replication needed 10B This improvement in training speed has been brought about by the following techniques: - Modernized architecture: Rotary embeddings, QK-Norm, and ReLU² - The Muon optimizer writeup https://kellerjordan.github.io/posts/muon/ repo https://github.com/KellerJordan/Muon - Use FP8 for head, and asymmetric rescale and softcap logits - Use FP8 on MLP forward pass - Initialization of projections to zero muP-like - Skip connections from embedding to every block as well as from block 3 to 6 - Extra embeddings which are mixed into the values in attention layers inspired by Zhou et al. 2024 - Flash Attention 3 with long-short sliding window attention pattern inspired by Gemma 2 and window size warmup with YaRN - Align training batch starts with EoS and set a max document length - Accumulate gradients for 2 steps for embedding and lm head before updating parameters - Single activation input for last 3 attention layers - Polar Express implementation in Muon - Smear module to enable 1 token look back - Sparse attention gate - NorMuon - Cautious Weight Decay w/ schedule tied to LR - Exponential decay of residual stream - Batch size schedule - Max seq length schedule - Partial Key Offset - Multi token prediction - Untie embed and lm head at 2/3 of training - Additional gating on value embeddings and skip connection - Paired head attention - Bigram hash embedding on 1/4 of model dim w/ sign trick - MUDD skip connections to residual stream and attention values - Learnable XSA - Lightweight Dynamically Composable MHA - Prefix token prediction auxiliary loss As well as many systems optimizations. Contributors list growing with each new record : @bozavlado https://x.com/bozavlado ; @brendanh0gan https://x.com/brendanh0gan ; @fernbear.bsky.social https://bsky.app/profile/fernbear.bsky.social ; @Grad62304977 https://x.com/Grad62304977 ; @jxbz https://x.com/jxbz ; @kellerjordan0 https://x.com/kellerjordan0 ; @KoszarskyB https://x.com/KoszarskyB ; @leloykun https://x.com/@leloykun ; @YouJiacheng https://x.com/YouJiacheng ; @jadenj3o https://x.com/jadenj3o ; @KonstantinWilleke https://github.com/KonstantinWilleke , @alexrgilbert https://github.com/alexrgilbert , @adricarda https://github.com/adricarda , @tuttyfrutyee https://github.com/tuttyfrutyee , @vdlad https://github.com/vdlad ; @ryanyang0 https://x.com/ryanyang0 , @vagrawal https://github.com/vagrawal , @classiclarryd https://x.com/classiclarryd , @byronxu99 https://github.com/byronxu99 , @varunneal https://x.com/varunneal , @EmelyanenkoK https://github.com/EmelyanenkoK , @bernard24 https://github.com/bernard24 / https://www.hiverge.ai/ https://www.hiverge.ai/ , @Gusarich https://x.com/Gusarich , @li zichong https://x.com/li zichong , @akash5474 https://github.com/akash5474 , @snimu https://x.com/omouamoua , @roeeshenberg https://x.com/roeeshenberg , @ChrisJMcCormick https://x.com/ChrisJMcCormick , @dominikkallusky https://github.com/dominikkallusky , @acutkosky https://github.com/acutkosky , @manikbhandari https://github.com/manikbhandari , @andrewbriand https://x.com/andrewbriand8 , @jrauvola https://x.com/Joshrav21 , @soren dunn https://x.com/soren dunn , @photon mz https://x.com/photon mz , @srashedll https://x.com/srashedll , @dhrvji https://x.com/dhrvji , @EmmettBicker https://github.com/EmmettBicker , @dualverse-ai https://github.com/dualverse-ai , @sisovicm https://x.com/sisovicm , @moof2x https://github.com/moof2x , @samacqua https://github.com/samacqua , @Lisennlp https://github.com/Lisennlp , @ djdumpling https://x.com/ djdumpling , @TrianX https://x.com/TrianX , @aryavohra https://github.com/aryavohra , @cong ml https://x.com/cong ml , @jvarho https://github.com/jvarho , @Mister-dev-oss https://github.com/Mister-dev-oss , @CerovazS https://github.com/CerovazS , @MarioPaerle https://github.com/MarioPaerle , @GabrieleCirillo https://github.com/GabrieleCirillo , @crisostomi https://github.com/crisostomi To run the current record, run the following commands. git clone https://github.com/KellerJordan/modded-nanogpt.git && cd modded-nanogpt pip install -r requirements.txt downloads only the first 900M training tokens to save time python data/cached fineweb10B.py 9 ./run.sh Add torchrun to path if ./run.sh gives error torchrun: command not found . Note: torch.compile will add around 7 minutes of latency the first time you run the code. Official records are timed on 8 NVIDIA H100 GPUs from https://app.primeintellect.ai/ https://app.primeintellect.ai/ . PrimeIntellect has generously sponsored recent validation runs. For cases where CUDA or NCCL versions aren't compatible with your current system setup, Docker can be a helpful alternative. This approach standardizes versions for CUDA, NCCL, CUDNN, and Python, reducing dependency issues and simplifying setup. Note: an NVIDIA driver must already be installed on the system useful if only the NVIDIA driver and Docker are available . git clone https://github.com/KellerJordan/modded-nanogpt.git && cd modded-nanogpt sudo docker build -t modded-nanogpt . sudo docker run -it --rm --gpus all -v $ pwd :/modded-nanogpt modded-nanogpt python data/cached fineweb10B.py 8 sudo docker run -it --rm --gpus all -v $ pwd :/modded-nanogpt modded-nanogpt sh run.sh To get an interactive docker, you can use sudo docker run -it --rm --gpus all -v $ pwd :/modded-nanogpt modded-nanogpt bash The following is the historical progression of world speed records for the following competitive task: Train a neural network to ≤3.28 validation loss on FineWeb using 8x NVIDIA H100s. Note: The 3.28 target was selected to match Andrej Karpathy's GPT-2 small reproduction https://github.com/karpathy/llm.c/discussions/481 . | | Record time | Description | Date | Log | Contributors | |---|---|---|---|---|---| | 1 | 45 minutes | | log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-10-13 llmc/main.log Tuned learning rate & rotary embeddings https://x.com/kellerjordan0/status/1798863559243513937 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-06-06 AdamW/f66d43d7-e449-4029-8adf-e8537bab49ea.log Introduced the Muon optimizer https://x.com/kellerjordan0/status/1842300916864844014 Muon improvements https://x.com/kellerjordan0/status/1844820919061287009 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-10-10 Muon/eb5659d0-fb6a-49e5-a311-f1f89412f726.txt Pad embeddings, ReLU², zero-init projections, QK-norm https://x.com/kellerjordan0/status/1845865698532450646 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-10-14 ModernArch/dabaaddd-237c-4ec9-939d-6608a9ed5e27.txt Distributed the overhead of Muon https://x.com/kellerjordan0/status/1847291684016783746 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-10-17 DistributedMuon/22d24867-eb5a-4fcc-ae2c-263d0277dfd1.txt Upgraded PyTorch 2.5.0 https://x.com/kellerjordan0/status/1847358578686152764 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-10-18 PyTorch25/d4bfb25f-688d-4da5-8743-33926fad4842.txt Untied embedding and head https://x.com/kellerjordan0/status/1853188916704387239 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-11-03 UntieEmbed/d6b50d71-f419-4d26-bb39-a60d55ae7a04.txt Value and embedding skip connections, momentum warmup, logit softcap https://x.com/kellerjordan0/status/1854296101303800108 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-11-06 ShortcutsTweaks/dd7304a6-cc43-4d5e-adb8-c070111464a1.txt Bfloat16 activations https://x.com/kellerjordan0/status/1855267054774865980 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-11-08 CastBf16/a833bed8-2fa8-4cfe-af05-58c1cc48bc30.txt U-net pattern skip connections & double lr https://x.com/kellerjordan0/status/1856053121103093922 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-11-10 UNetDoubleLr/c87bb826-797b-4f37-98c7-d3a5dad2de74.txt 1024-ctx dense causal attention → 64K-ctx FlexAttention https://x.com/kellerjordan0/status/1859331370268623321 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-11-19 FlexAttention/8384493d-dba9-4991-b16b-8696953f5e6d.txt Attention window warmup https://x.com/hi tysam/status/1860851011797053450 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-11-24 WindowWarmup/cf9e4571-c5fc-4323-abf3-a98d862ec6c8.txt Value Embeddings https://x.com/KoszarskyB/status/1864746625572257852 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-12-04 ValueEmbed U-net pattern value embeddings, assorted code optimizations https://x.com/YouJiacheng/status/1865761473886347747 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-12-08 UNetValueEmbedsTweaks Split value embeddings, block sliding window, separate block mask https://x.com/YouJiacheng/status/1866734331559071981 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-12-10 MFUTweaks Sparsify value embeddings, improve rotary embeddings, drop an attn layer https://x.com/YouJiacheng/status/1868938024731787640 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2024-12-17 SparsifyEmbeds Lower logit softcap from 30 to 15 https://x.com/kellerjordan0/status/1876048851158880624 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-01-04 SoftCap/31d6c427-f1f7-4d8a-91be-a67b5dcd13fd.txt FP8 head, offset logits, lr decay to 0.1 instead of 0.0 https://x.com/YouJiacheng/status/1878827972519772241 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-01-13 Fp8LmHead/c51969c2-d04c-40a7-bcea-c092c3c2d11a.txt Merged QKV weights, long-short attention, attention scale, lower Adam epsilon, batched Muon https://x.com/leloykun/status/1880301753213809016 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-01-16 Sub3Min/1d3bd93b-a69e-4118-aeb8-8184239d7566.txt Reduced batch size https://x.com/leloykun/status/1885640350368420160 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-01-26 BatchSize/c44090cc-1b99-4c95-8624-38fb4b5834f9.txt log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-02-01 RuleTweak/eff63a8c-2f7e-4fc5-97ce-7f600dae0bc7.txt updated rules timing-change-after-record-21 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-05-24 StableTorch/89d9f224-3b01-4581-966e-358d692335e0.txt Faster gradient all-reduce https://x.com/KonstantinWille/status/1927137223238909969 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-05-24 FasterReduce/23f40b75-06fb-4c3f-87a8-743524769a35.txt Overlap computation and gradient communication https://x.com/kellerjordan0/status/1927460573098262616 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-05-25 EvenFasterReduce/6ae86d05-5cb2-4e40-a512-63246fd08e45.txt log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-05-30 noallreduce/8054c239-3a18-499e-b0c8-dbd27cb4b3ab.txt log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-07-13 UpgradeTorch190/692f80e0-5e64-4819-97d4-0dc83b7106b9.txt log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-07-12 BosAlign/c1fd8a38-bb9f-45c4-8af0-d37f70c993f3.txt log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-07-18 TritonMuon/record.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/109 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-08-23 SparseAttnGate/020630eb-2191-4ba2-9ee4-4cdc94316943.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/117 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-03 FA3/44fc1276-0510-4961-92c0-730c65e5feba.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/118 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-05 SkipMLPBlocks/07e7ae76-b7d0-4481-b149-01e7d81b5ad4.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/120 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-10 Yarn/0ecdb695-510b-4c3b-b030-09861a162ce8.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/122 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-11 VectSigmoidBFloat16/0d0d9882-c34f-4d82-b961-a17d5659c988.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/125 hiverge.ai https://www.hiverge.ai/ log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-15 AsyncDataLoadAttnFinalWindow/25db37c7-2bab-4ef4-ae63-d593590ef823.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/127 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-18 Smear/18a1e5c7-947e-479d-bc3a-a57a61a98fc9.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/130 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-21 DropAttn/01fc4a96-f2a0-47a1-8a6a-c7d10bac99fe.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/131 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-23 MuonCustomSizing/b067b4ac-72a6-4436-a6f8-ea51c1efeef3.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/132 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-27 BF16CE/08c0770f-17fc-44cd-971d-734a7a28a3e3.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/133 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-29 PolarExpress/0e3f0af5-ad08-47a6-813d-0c709b50d422.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/134 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-09-30 CustomBatching/40b101b1-77ea-45ea-a089-1d3a647daa22.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/136 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-10-04 Backout/514e7581-fbd4-4338-a3e4-e556f9c958ce.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/140 NorMuon https://arxiv.org/pdf/2510.05491 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-10-24 NorMuon/088a77ee-9b67-475a-bbb9-3e92e4698799.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/144 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-10-27 FixMuonLR/14afd380-d3d9-48d7-ad23-4c13cb96754b.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/146 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-11-10 CautiousWD/1aac0132-a891-4ed9-b358-0fd2abd1b019.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/154 Profiling 101 https://blog.underfit.ai/profiling-101-nanogpt log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-10-31 AdamSyncGradientHook/0c17cdfd-772c-4906-8d11-141b370599a0.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/149 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-11-18 RefineSkip/00f4e1e6-0044-4a08-b88a-3b7ec0624081.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/159 Batch size schedule https://x.com/classiclarryd/status/1998212158770065844 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-11-29 BatchSizeSchedule/10e8f7c6-7175-4467-bdb0-a5de25d771a6.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/163 Multiply attn lambda with weight instead of data, fix warmup https://x.com/classiclarryd/status/1999630732814348451 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-10 SALambdaOnWeights/15ef5eaf-56e1-40e1-9ddf-af010027c9dd.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/166 Speed up Muon, additional pre-multiply lambda, reshape matrices, update lr, update NorMuon axis https://x.com/classiclarryd/status/2000272495644152317 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-11 NorMuonOptimsAndFixes/82edf6be-f343-475d-b93a-47c32acf4de2.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/168 Partial Key Offset https://x.com/classiclarryd/status/2000841339299402142 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-14 PartialKeyOffset/150d40bf-c20b-4568-aac9-26eb919e25fd.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/169 Extend Cautious Weight Decay to Adam parameters https://x.com/classiclarryd/status/2002482925741486381 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-18 CautiousWDAdam/1981d492-bc65-4ba9-a0fa-2b30fc5c3eba.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/172 Retie Embed to lm head, retune fp8 scales https://x.com/classiclarryd/status/2003167208483209668 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-19 RetieLMHead/0828d309-ecfe-4442-9ee9-68fed3a4b599.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/175 Smooth scalars via beta increase, decrease smear gate lr, freeze scalars during transitions, adam all reduce https://x.com/classiclarryd/status/2003863282613190656 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-21 SmoothedScalars/12-21-Smoothed-Scalars/0bc6e909-8ee8-4ae3-ac62-0070e151a808.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/177 Multi-token prediction, untie embed/lm head at 2/3 training, lr update, tweak CWD https://x.com/classiclarryd/status/2004248941878296580 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-22 MultiTokenPrediction/17aaf854-f338-4d0d-9767-a5db30fd7980.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/178 Asymmetric Logit Rescale https://x.com/classiclarryd/status/2004791008098480232 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-26 LogitRescale/03e41c2d-2951-4546-a599-24cd723247fc.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/181 Gates on value embeds and skip connection https://x.com/classiclarryd/status/2005659526960492638 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-29 VeSkipGates/2851d7dc-d6a5-4e74-8623-57031425db16.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/186 Optimize and compile Adam, increase Adam buffer precision, move gates from Muon to Adam parameter banks https://x.com/classiclarryd/status/2007882371576873445 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-12-31 GatesToCompiledAdam/12-31-gates-to-adam-20stps/219a5f2f-151e-4c56-ab91-3735ae4610b8.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/187 Bfloat16 attn/mlp weights, mixed precision Muon, interweave Adam/Muon, finer-grain Adam beta https://x.com/classiclarryd/status/2008261904566022590 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-04 MixedPrecisionInterweavedOptimizer/41f606b6-1b9c-46a3-b46e-2beff1521d18.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/190 Paired Head Attention https://x.com/classiclarryd/status/2008963501688324228 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-07 PairedHeadAttention/2a5d5cde-db5f-4aab-a4a8-cc8e183ea671.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/191 Fused triton kernel for linear relu square MLP step https://x.com/classiclarryd/status/2010545452832407943 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-10 FusedLinearReLUSquare/3c47e63b-075e-4b5b-9c76-9dbe7bad9ad4.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/197 Fused triton kernel for softcapped multi-token prediction cross entropy step https://x.com/classiclarryd/status/2012927211448516796 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-16 FusedSoftcappedEntropy/45beba56-93e2-4995-bc5b-caff3cb2c1b5.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/199 Locus https://www.intology.ai/blog/previewing-locus Unified Optimizers and Transposed LM Head https://x.com/classiclarryd/status/2013399457841160702 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-18 UnifiedOptimizers/unified-optimizer/2fc79469-a527-4bde-8540-8426ed3352d1.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/200 Bigram Hash Embedding https://x.com/classiclarryd/status/2013520088297558274 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-19 BigramHashEmbedding/40ec7bb6-14b3-46f8-90b7-bb5ed188faba.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/201 Untie Value Embeds https://x.com/classiclarryd/status/2016968386476200301 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-26-UntieValueEmbeddings/43955d93-6914-40cb-bdf8-786ace93784f.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/209 Tuned nonzero Attn V and O init https://x.com/classiclarryd/status/2017735338601726357 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-30 MimeticValueOutput/runs/0f262f64-20c4-4268-9ae7-d7440c810abd.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/214 Group Value Embeds into single parameter https://x.com/classiclarryd/status/2018057653742920016 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-30 VeFused/0ba09d92-4ef1-440f-85e3-9d2766294db4.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/215 Tune fused softcap kernels and fuse fp8 quantization in LM head https://x.com/classiclarryd/status/2021015642472869978 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-24 ImprovedLMHead/record/73a071ac-522d-4ce0-a4d6-cf3955a376e4.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/207 Move bigram hash to GPU https://x.com/classiclarryd/status/2021450730117460439 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-01-31-BigramHashH2D/112c686e-b0d6-4dc8-814a-1ad1f5d5b274.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/216 Kernel Optimizations https://x.com/classiclarryd/status/2023319358303510719 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-02 KernelTuning/25afb73a-332f-4d69-b9ab-f6261497f2d8.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/217 Aster https://www.asterlab.ai/ Tune value embed layout and ve gates https://x.com/classiclarryd/status/2023319358303510719 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-03 VeTuned/42cbebac-0599-4a89-a00e-26d1c4cad140.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/218 Sparse bigram gradient comms and optimized loading on CPU https://x.com/classiclarryd/status/2023319358303510719 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-06 SparseBigramGradient/02fee7bd-cd22-478b-9e8e-12e857ff3152.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/221 Increase minimum lr and add max seq len schedule https://x.com/classiclarryd/status/2023319358303510719 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-10 ShortWindow/Short-Window 1 1.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/224 Station https://github.com/dualverse-ai/station Partitioned Hyperconnections https://x.com/classiclarryd/status/2026131531207761924 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-12 ParallelResiduals/451050db-d471-49db-b19b-be824bb896d0.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/230 Flattened GPT forward, removed post attention lambdas, added transpose kernels https://x.com/classiclarryd/status/2027228782483182059 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-16 FlattenForward/pr233/2026-02-16 21-30-05 time-362 secs F-inject-post-attn 9f12a3.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/233 Cross Entropy Kernel Optimizations https://x.com/classiclarryd/status/2030087884854939947 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-23 CrossEntropyKernel/1e51be6b-7dd4-41ab-b95d-e57da5814776.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/235 Reuse and tune backward transpose kernel https://x.com/classiclarryd/status/2030403421027852337 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-02-28 TransposeCopyBackward/this pr/14c9cefc-c840-493f-870e-61bb1d2b1d97.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/240 Replace partitioned hyperconnections with single saved activation https://x.com/classiclarryd/status/2030465730718908884 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-03-06 SimplifyHC/0ab4a843-8c3a-4fb4-9fff-8e1d39852646.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/241 Tighten bounds on fa3 max num docs to match fineweb distribution https://x.com/classiclarryd/status/2038077427180851240 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-03-22 VarlenMaxDocs/combined/2026-03-22 20-07-32 time-186 secs 06-mbeta2-max-docs 227ce8.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/246 Fuse Cross Entropy Fwd/Bwk Kernel, to avoid recalc on softcap sigmoid https://x.com/classiclarryd/status/2045270983343485140 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-04-04 FuseCEFwdAndBwd/runs/19ad9161-37c0-4985-8dd4-6db4e27f34b4.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/251 In Muon orthogonize Q and K matrices in pairs of heads, instead of across the full 6 head matrix https://x.com/classiclarryd/status/2046046809609457904 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-04-08 PairedHeadMuon/logs/split qk0-1480.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/253 MUDD Skip Connections https://x.com/classiclarryd/status/2058486428255035457 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-04-22 MuddFormer/this pr v3 , PR https://github.com/KellerJordan/modded-nanogpt/pull/259 Learnable XSA https://x.com/classiclarryd/status/2058975556520329302 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-04-29 XSAGatedLayers/this pr v1-s1410/06563169-6435-48ba-a1ad-f3e61bfcc573.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/264 Sign Trick on Bigram Embed https://x.com/classiclarryd/status/2063061926092099868 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-05-20 BigramsSignTrick/pr299/0cf91274-eda8-49cd-9a97-9369f730f271.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/299 FP8 on MLP up-projection forward pass /KellerJordan/modded-nanogpt/blob/master log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-05-19 FP8MLPUpProj/this record/008bb79d-d5bc-4205-bd4e-5e4ae82e658c.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/306 MUDD gates and Lightweight Dynamically Composable MHA https://x.com/classiclarryd/status/2081451521229881374 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-05-27-MuddGatedAndDC , PR https://github.com/KellerJordan/modded-nanogpt/pull/315 Algebraic rewrite of XSA, same math faster execution https://x.com/classiclarryd/status/2081909171554071027 PR https://github.com/KellerJordan/modded-nanogpt/pull/317 Faster Implementation of Relu^2 Kernel https://x.com/classiclarryd/status/2083739041338630372 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-06-11 RecursiveFromBest/this pr/00088a48-30a3-4ebd-9768-6061011337f4.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/322 Recursive https://www.recursive.com/ Prefix token prediction auxiliary loss https://x.com/classiclarryd/status/2083961001930834419 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-07-13 PrefixTokenPrediction/prefix-1375/1b20ccf2-cb2f-4b6b-bc8a-2d9cd146f549.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/337 MLP down projection in FP8 with efficient delayed scaling metric https://x.com/classiclarryd/status/2086582390135406713 log /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2026-07-17 FP8DownProjection/this pr/11cb620c-daaf-4e85-83fc-258a5eb7ba09.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/342 New records must: - Not modify the train or validation data pipelines. You can change the batch size, sequence length, attention structure etc.; just don't change the underlying streams of tokens. - Attain ≤3.28 mean val loss. Due to inter-run variance, submissions must provide enough run logs to attain a statistical significance level of p<0.01 that their mean val loss is ≤3.28. Example code to compute p-value can be found here /KellerJordan/modded-nanogpt/blob/master/records/track 1 short/2025-01-04 SoftCap softer-softcap . For submissions which improve speed by optimizing the systems performance, without touching the ML, this requirement is waived. - Not use any extra torch. inductor.config or torch.compile flags. These can save a few seconds, but they can also make compilation take 30min. This rule was introduced after the 21st record. - Run faster than the prior record when baselined on the same hardware. Incorporating open PRs into a new record is strongly encouraged. This speeds up merges through peer validation and prevents new PRs from going stale due to conflicts with earlier, still-open PRs. Discretionary reasons why a PR may not be accepted: - Disproportionately degrades the readability of the codebase. A 200 line kernel to drop 300ms is considered worthwhile. 500 lines that convolute the optimizer layout for a 50ms gain will likely be rejected. - The current record is intentionally kept roughly 0.001-0.002 loss below 3.28 to make validation simpler. If a PR substantially consumes this buffer, it should do so in a way that outperforms a simple step count decrease, when measured at equivalent loss. Note: torch. inductor.config.coordinate descent tuning is allowed for GPT-2 Medium track a.k.a. 2.92 track . Other than that, anything and everything is fair game The target metric is cross-entropy loss on the FineWeb val set . To speak mathematically, the goal of the speedrun is to obtain a probability model of language which assigns a probability of at least math.exp -3.28 10485760 to the first 10,485,760 tokens of the FineWeb valset. Hence, e.g., we allow evaluation at any sequence length, so long as we still have a valid probability model of language. After the 21st record, we made two changes to the timing. First, there used to be an initial "grace period" of 10 untimed steps to allow kernel warmup. We replaced this with an explicit kernel-warmup section which is untimed and uses dummy data. This results in an extra runtime of 850ms from the 10 extra timed steps. Second, we banned the use of torch. inductor.config.coordinate descent tuning . This saves ~25min of untimed pre-run compilation, but results in an extra runtime of ~3s. Notable runs: @alexjc's 01/20/2025 2.77-minute TokenMonster-based record https://x.com/alexjc/status/1881410039639863622 . This record is technically outside the rules of the speedrun, since we specified that the train/val tokens must be kept fixed. However, it's very interesting, and worth including. The run is not more data-efficient; rather, the speedup comes from the improved tokenizer allowing the vocabulary size to be reduced nearly halved while preserving the same bytes-per-token, which saves lots of parameters and FLOPs in the head and embeddings. @samacqua's 1/23/2026 test time training run https://github.com/KellerJordan/modded-nanogpt/pull/205 . Sam found that prediction accuracy on the later portions of a given document could be improved by performing a training update on Adam parameters based on the early portion of the document. This 'parameter nudging' is repeated independently for each document. Interestingly, these gradient updates prove effective while only using ~500 tokens, substantially less than the over 200k tokens typically used on a normal training step. While technically a valid probability model, we are not allowing untimed backward passes. Notable forks: The target loss for this track is lowered from 3.28 to 2.92, as per Andrej Karpathy's 350M-parameter llm.c baseline. This baseline generates a model with performance similar to the original GPT-2 Medium, whereas the first track's baseline generates a model on par with GPT-2 Small. All other rules remain the same. Note: torch. inductor.config.coordinate descent tuning is turned on after the record 6 . | | Record time | Description | Date | Log | Contributors | |---|---|---|---|---|---| | 1 | 5.8 hours | | log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-01-18/main.log Initial record based on scaling up the GPT-2 small track speedrun https://x.com/kellerjordan0/status/1881959719012847703 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-01-18/241dd7a7-3d76-4dce-85a4-7df60387f32a.txt Added standard weight decay https://x.com/kellerjordan0/status/1888320690543284449 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-02-08 WeightDecay/b01743db-605c-4326-b5b1-d388ee5bebc5.txt Tuned Muon Newton-Schulz coefficients https://x.com/leloykun/status/1892793848163946799 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-02-14 OptCoeffs/1baa66b2-bff7-4850-aced-d63885ffb4b6.txt Increased learning rate cooldown phase duration /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-03-06 LongerCooldown/779c041a-2a37-45d2-a18b-ec0f223c2bb7.txt log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-03-06 LongerCooldown/779c041a-2a37-45d2-a18b-ec0f223c2bb7.txt 2x MLP wd, qkv norm, all reduce/opt.step overlap, optimized skip pattern https://x.com/YouJiacheng/status/1905861218138804534 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-03-25 ArchOptTweaks/train gpt-20250329.txt Remove FP8 head; ISRU logits softcap; New sharded mixed precision Muon; merge weights https://x.com/YouJiacheng/status/1912570883878842527 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-04-16 Record7/223 3310d0b1-b24d-48ee-899f-d5c2a254a195.txt Cubic sliding window size schedule, 2× max window size 24.84 minutes https://x.com/jadenj3o/status/1914893086276169754 24.5min repro https://x.com/YouJiacheng/status/1915667616913645985 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-04-22 Record8/075 640429f2-e726-4e83-aa27-684626239ffc.txt Add two value embeddings https://snimu.github.io/2025/10/07/modded-nanogpt-value-embeddings.html log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-08-28 NewValemb/036 61ef4351-7b68-4897-b440-a99221a1a629.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/119 Second input embedding https://snimu.github.io/2025/10/10/modded-nanogpt-x0.html log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-09-11 SecondInputEmbed/000 592014ec-6781-4f59-b274-c4af68ccfe75.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/124 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-09-16 Snoo/000 01db7a67-f715-4114-a7b5-6bfe23bac1b1.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/128 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-09-17 UpdateSmoothing/001 8379f695-6bc3-4f76-b58b-8fadd3b6ebb0.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/129 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-09-30 SmoothedSnooMedium/101 5bc91cd0-cb46-428c-a5da-9d8d228f1f97.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/137 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-10-04 GPT2MediumLayerReuse/000 cc3943e4-02b5-4ae3-9441-839d32dfd9b2.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/139 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-11-02-Smear-MTP/000 3b50518d-d542-44bc-8566-3abf633f83ad.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/151 log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-11-12 BlockMaskRedundantOp/000 3b22a9d4-b52e-4916-99bf-3d48b38747a7.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/157/ log /KellerJordan/modded-nanogpt/blob/master/records/track 2 medium/2025-12-31 BulkSmallTrackTransfer/354be270-7d41-44b7-8064-f040923f024f.txt , PR https://github.com/KellerJordan/modded-nanogpt/pull/188 A: The officially stated goal of NanoGPT speedrunning is as follows: gotta go fast . But for something a little more verbose involving an argument for good benchmarking, here's some kind of manifesto, adorned with a blessing from the master. https://x.com/karpathy/status/1846790537262571739 https://x.com/karpathy/status/1846790537262571739 A: Because it is a competitive benchmark. In particular, if you attain a new speed record using whatever method you want , there is an open invitation for you to post that record on arXiv or X and thereby vacuum up all the clout for yourself. I will even help you do it by reposting you as much as I can. Q: NanoGPT speedrunning is cool and all, but meh it probably won't scale and is just overfitting to val loss A: This is hard to refute, since "at scale" is an infinite category what if the methods stop working only for 100T models? , making it impossible to fully prove. Also, I would agree that some of the methods used in the speedrun are unlikely to scale, particularly those which impose additional structure on the network, such as logit softcapping. But if the reader cares about 1.5B models, they might be convinced by this result: Straightforwardly scaling up the speedrun 10/18/24 version to 1.5B parameters yields a model with GPT-2 1.5B -level HellaSwag performance 2.5x more cheaply than @karpathy's baseline $233 instead of $576 : Muon is defined as follows: Where NewtonSchulz5 is the following Newton-Schulz iteration 2, 3 , which approximately replaces G with U @ V.T where U, S, V = G.svd . python @torch.compile def zeroth power via newtonschulz5 G, steps=5, eps=1e-7 : assert len G.shape == 2 a, b, c = 3.4445, -4.7750, 2.0315 X = G.bfloat16 / G.norm + eps if G.size 0 G.size 1 : X = X.T for in range steps : A = X @ X.T B = b A + c A @ A X = a X + B @ X if G.size 0 G.size 1 : X = X.T return X.to G.dtype For this training scenario, Muon has the following favorable properties: - Lower memory usage than Adam - ~1.5x better sample-efficiency - <2% wallclock overhead Many of the choices made to generate this optimizer were obtained experimentally by our pursuit of CIFAR-10 speedrunning https://github.com/KellerJordan/cifar10-airbench . In particular, we experimentally obtained the following practices: - Using Nesterov momentum inside the update, with orthogonalization applied after momentum. - Using a specifically quintic Newton-Schulz iteration as the method of orthogonalization. - Using non-convergent coefficients for the quintic polynomial in order to maximize slope at zero, and thereby minimize the number of necessary Newton-Schulz iterations. It turns out that the variance doesn't actually matter that much, so we end up with a quintic that rapidly converges to the range 0.68, 1.13 upon repeated application, rather than converging more slowly to 1. - Running the Newton-Schulz iteration in bfloat16 whereas Shampoo implementations often depend on inverse-pth-roots run in fp32 or fp64 . Our use of a Newton-Schulz iteration for orthogonalization traces to Bernstein & Newhouse 2024 https://arxiv.org/abs/2409.20325 , who suggested it as a way to compute Shampoo 5, 6 preconditioners, and theoretically explored Shampoo without preconditioner accumulation. In particular, Jeremy Bernstein @jxbz sent us the draft, which caused us to experiment with various Newton-Schulz iterations as the orthogonalization method for this optimizer. If we had used SVD instead of a Newton-Schulz iteration, this optimizer would have been too slow to be useful. Bernstein & Newhouse also pointed out that Shampoo without preconditioner accumulation is equivalent to steepest descent in the spectral norm, and therefore Shampoo can be thought of as a way to smooth out spectral steepest descent. The proposed optimizer can be thought of as a second way of smoothing spectral steepest descent, with a different set of memory and runtime tradeoffs compared to Shampoo. - To run experiments on fewer GPUs, simply modify run.sh to have a different --nproc per node . This should not change the behavior of the training. - If you're running out of memory, you may need to reduce the sequence length for FlexAttention which does change the training. see here https://github.com/KellerJordan/modded-nanogpt/pull/38 for a guide Guilherme Penedo et al. "The fineweb datasets: Decanting the web for the finest text data at scale." arXiv preprint arXiv:2406.17557 2024 . https://arxiv.org/abs/2406.17557 - Nicholas J. Higham. Functions of Matrices. Society for Industrial and Applied Mathematics 2008 . Equation 5.22. - Günther Schulz. Iterative Berechnung der reziproken Matrix. Z. Angew. Math. Mech., 13:57â��59 1933 . Jeremy Bernstein and Laker Newhouse. "Old Optimizer, New Norm: An Anthology." arxiv preprint arXiv:2409.20325 2024 . https://arxiv.org/abs/2409.20325 Vineet Gupta, Tomer Koren, and Yoram Singer. "Shampoo: Preconditioned stochastic tensor optimization." International Conference on Machine Learning. PMLR, 2018. https://arxiv.org/abs/1802.09568 Rohan Anil et al. "Scalable second order optimization for deep learning." arXiv preprint arXiv:2002.09018 2020 . https://arxiv.org/abs/2002.09018 Alexander Hägele et al. "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations." arXiv preprint arXiv:2405.18392 2024 . https://arxiv.org/abs/2405.18392 Zhanchao Zhou et al. "Value Residual Learning For Alleviating Attention Concentration In Transformers." arXiv preprint arXiv:2410.17897 2024 . https://arxiv.org/abs/2410.17897 Team, Gemma, et al. "Gemma 2: Improving open language models at a practical size." arXiv preprint arXiv:2408.00118 2024 . https://arxiv.org/abs/2408.00118 Alec Radford et al. "Language models are unsupervised multitask learners." OpenAI blog 1.8 2019 . https://cdn.openai.com/better-language-models/language models are unsupervised multitask learners.pdf @misc{modded nanogpt 2024, author = {Keller Jordan and Jeremy Bernstein and Brendan Rappazzo and @fernbear.bsky.social and Boza Vlado and You Jiacheng and Franz Cesista and Braden Koszarsky and @Grad62304977}, title = {modded-nanogpt: Speedrunning the NanoGPT baseline}, year = {2024}, url = {https://github.com/KellerJordan/modded-nanogpt} }