{"slug": "why-do-openai-s-gpt-2-weights-beat-mine-part-five-data-quality", "title": "Why do OpenAI's GPT-2 weights beat mine? Part five: data quality", "summary": "Independent LLM-from-scratch experiments found that OpenAI's GPT-2 small (124M parameters) consistently outperformed the author's 163M-parameter models on an instruction fine-tuning task adapted from Sebastian Raschka's book, despite using the same architecture. The author ruled out overtraining and dropout handling as explanations in earlier parts of the series, and is now investigating training data quality, noting OpenAI never released the GPT-2 dataset and only described it in the GPT-2 paper as a web scrape emphasizing document quality, seeded from Reddit outbound links. The author's models differ from GPT-2 small by omitting weight-tying and QKV bias.", "body_md": "## Why do OpenAI's GPT-2 weights beat mine? Part five: data quality\n\nWhen I finished learning how to build an [LLM from scratch](https://www.gilesthomas.com/llm-from-scratch),\nI was left with a mystery: my own models were not as good as OpenAI's original\nGPT-2 models, despite being based on the same architecture.\n\nMy models all had 163M parameters, and followed the design from\n[Sebastian Raschka](https://sebastianraschka.com/)'s book\n\"[Build a Large Language Model (from Scratch)](https://www.manning.com/books/build-a-large-language-model-from-scratch)\".\nThat meant that they were pretty much the same as the setup for the OpenAI GPT-2 \"small\" instance,\nexcept that they did not use weight-tying or bias on the QKV matrices.  Weight-tying means that you re-use\nthe initial embedding matrix as the output head at the end, and using it means that GPT-2 small\nsaved quite a few parameters -- it was 124M rather than 163M -- at, at least in my\nown experiments, a [cost in quality](https://www.gilesthomas.com/2026/03/llm-from-scratch-32g-interventions-weight-tying);\nsimilarly, while I found that QKV bias made a [tiny improvement in loss terms](https://www.gilesthomas.com/2026/02/llm-from-scratch-32d-interventions-adding-attention-bias), I'd\nfelt it was likely within the noise.\n\nBut GPT-2 small consistently beat my models on an instruction\nfine-tuning (IFT) task -- also adapted from Raschka's book.   That test fine-tunes the model\non a subset of the [Alpaca](https://github.com/tatsu-lab/stanford_alpaca) dataset, until\nvalidation loss starts rising, and then runs a test set through the resulting model.\nThe responses to the test set questions are stored, and then I run all of the responses\nfrom all of the models under test past GPT 5.5 in one go to get an aggregate score;\n[more details here](https://www.gilesthomas.com/2026/04/llm-from-scratch-32l-interventions-instruction-fine-tuning-tests).\nGPT-2 small\nalways did better than any of my models on this.\n\nAdditionally, it did surprisingly well on a simpler eval -- one that just measured the cross entropy loss it got on a test set. It scored close to my own best models, and better than many of them. What made this result particularly interesting was that the test set in question was a split of my own training data; my models would not have seen it when training (at least, in theory), but it seems likely that it would be much more similar to their own training data than it was to OpenAI's.\n\nI've checked two things while probing this mystery:\n\n- It seems very likely that the GPT-2 models were overtrained by modern standards; would\n[overtraining my own](https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-3-overtraining) models\nget them closer?  It turned out that no, it probably didn't help with the IFT eval (though there might\nhave been some signal there).  It did help quite a lot with the test loss eval, though.\n- The way I was handling [dropout](https://www.gilesthomas.com/2026/08/why-do-openai-gpt2-weights-beat-mine-4-ift-dropout) in the IFT test might\nhave been unduly benefiting some models while working against others.  I decided\nto standardise on*not* using dropout during this eval, as (counter-intuitively for me) it seemed\nto harm the results of most models, even those that had been pre-trained*with* dropout.\nIn particular, the OpenAI weights were harmed by using dropout, and making a change\nthat benefited them (along with some of my own models) seemed the most conservative\napproach to take in investigating this.\n\nThe next thing I wanted to look into was the training data.\n\nThe exact dataset that the various GPT-2 models were trained on has never been released;\nall we know about it\nis from [the paper](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf),\nwhere they say:\n\n[W]e created a new web scrape which emphasizes document quality. To do this we only scraped web pages which have been curated/filtered by humans. Manually filtering a full web scrape would be exceptionally expensive so as a starting point, we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma. This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny.\n\nThey called it \"WebText\".  There is an [OpenWebText](https://skylion007.github.io/OpenWebTextCorpus/)\nthat tries to replicate it, but although they tried to follow the same procedure as the\noriginal, there's no guarantee that it is all that similar.\n\nBy comparison, I'd normally been training against [FineWeb](https://huggingface.co/datasets/HuggingFaceFW/fineweb).\nWhile this is a general web-scraping dataset, without the \"curation\" provided by\nusing only stuff that was linked from upvoted Reddit posts, it has been refined to\nremove any obvious junk.  I had felt that it was\npretty much equivalent.\n\nBut what if I were wrong about that? I decided to see if I could get better models by using better data.\n\n### The starting point\n\nHere's a table of all of the models I've been comparing to date. The \"Test loss\" column shows how well the model in question did on that held-back cross entropy loss evaluation. The \"IFT epochs\" column shows how many epochs of fine-tuning the model needed before its validation loss started rising, the \"IFT score\" the score that GPT 5.5 gave the model's responses to the test set of my Alpaca data, and the \"IFT rank\" the model's rank in terms of that score. The OpenAI small model is in there in bold, and I've also included the OpenAI medium model for comparison purposes.\n\n|  | Test loss | IFT epochs | IFT score | IFT rank | \n|---|---|---|---|---|\n| **OpenAI weights: medium** | 3.231442 | 2 | 43.75 | 1 | \n| JAX, overtrained one long epoch | 3.324953 | 3 | 19.77 | 4 | \n| JAX, overtrained two normal epochs | 3.326482 | 4 | 19.72 | 5 | \n| JAX, with MHA bias, no dropout | 3.418784 | 4 | 18.69 | 6 | \n| JAX, no MHA bias, no dropout | 3.420089 | 5 | 21.46 | 3 | \n| JAX, no MHA bias, with dropout | 3.476802 | 5 | 13.22 | 15 | \n| **OpenAI weights: small** | 3.499677 | 2 | 26.00 | 2 | \n| `1xrtx3090-stacked-interventions` | 3.538161 | 4 | 13.77 | 14 | \n| `8xa100m40-stacked-interventions-1` | 3.577761 | 4 | 10.76 | 18 | \n| Cloud FineWeb, 8x A100 40 GiB | 3.673623 | 3 | 17.72 | 7 | \n| `1xrtx3090-baseline` | 3.683835 | 4 | 15.74 | 8 | \n| `8xa100m40-baseline` | 3.691526 | 3 | 14.19 | 13 | \n| Cloud FineWeb, 8x H100 80 GiB | 3.724507 | 4 | 14.33 | 12 | \n| Cloud FineWeb, 8x A100 80 GiB | 3.729900 | 3 | 11.34 | 17 | \n| Cloud FineWeb, 8x B200 160 GiB | 3.771478 | 4 | 14.67 | 11 | \n| Local FineWeb train | 3.943522 | 5 | 12.31 | 16 | \n| Local FineWeb-Edu extended train | 4.134991 | 5 | 15.04 | 9 | \n| Local FineWeb-Edu train | 4.166892 | 5 | 14.99 | 10 | \n\nYou can see that the OpenAI small model did pretty well in terms of the test loss, when you consider that it has 39M fewer weights than my models and was being tested against a dataset that differs more from its likely training data than it does from my own models'. Additionally, the specific models that did better than OpenAI's small one were all trained with JAX rather than PyTorch -- my hypothesis for that is that it's a result of the JAX ones getting better initial weights by pure chance.\n\nBut the big difference was in the IFT score. In the specific run that gave the results in this table, the OpenAI small model got 26.00 -- the closest of my own models was more than 4.5 points lower, at 21.46.\n\nThis difference was consistent over all of my other test runs. The GPT-2 small model was always ahead of mine. (GPT-2 medium, of course, beat GPT-2 small and all of my models, but given that it is twice the size of mine, that's not a big surprise.)\n\nNow, quite some time ago, I had tried looking into data quality as a lever to pull for model performance. At the bottom of the table, with the worst test loss of all models, you can see two models:\n\n- \"Local FineWeb-Edu train\"\n- \"Local FineWeb-Edu extended train\"\n\nThese two were (as you might guess from the names) trained on the [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu)\ndataset, which includes just the most \"educational\" data from FineWeb.  They\nscored very badly on the test loss score.  Given that the test dataset is from FineWeb,\nthat's not a big surprise -- as I've written previously:\n\nIf you train a model\n  on Jane Austen and then evaluate against [Chuck Tingle](https://www.chucktingle.com/), then\n  you're not going to get amazing results.\n\nBut again, GPT-2 had the same issue, and did perfectly well on the test loss eval.\n\nOn the other hand, while these FineWeb-Edu models' performance on the IFT eval wasn't stellar -- there are plenty of my other models ahead of them -- they did seem to punch above their weight. Consistently across all of the IFT evals I've done, they have scored higher than many of the others -- despite their poor loss on the test eval.\n\nAdditionally: they were amongst the first models that I trained, before I'd spent\ntime learning about how to [optimise my hyperparameters and training loop](https://www.gilesthomas.com/2026/04/llm-from-scratch-32m-interventions-conclusion).\nThey did not use gradient clipping, they did use dropout, their batch size was just \"whatever\nI could squeeze into the GPU\", and I didn't set the learning\nrate to the right kind of value or schedule it over the course of the training run.\n\nSo maybe a new training run on FineWeb-Edu plus my training improvements would help? And maybe some other tweaks to the training data would be worth looking into?\n\n### The plan\n\nI decided to see what would happen if I trained some models with better-quality data. Specifically, I would train models with my current optimised loop and hyperparameters on four different datasets:\n\n- FineWeb-Edu -- essentially the same as \"Local FineWeb-Edu train\" but with a better training setup. This would test the \"more educational -> better\" hypothesis.\n- A 50:50 split of FineWeb and FineWeb-Edu.  I've read that LLMs can be *helped* by\nhaving a decent amount of lower-quality data in their training loop, as it helps them to generalise.\nPerhaps having some FineWeb in there in addition to the FineWeb-Edu stuff would improve that test loss score while\nalso helping the IFT test?\n- A \"curated\" dataset containing 45% of its contents from FineWeb, 45% from FineWeb-Edu, and 10% from the\n[Simple English Wikipedia](https://simple.wikipedia.org/wiki/Main_Page) .  The full\nWikipedia is huge, and full of obscure facts -- while the Simple English one is\nsmall and hopefully richer in useful information on a per-token basis.  And conveniently,\nAnswer.ai have made[a snapshot of it available on Hugging Face Hub](https://huggingface.co/datasets/answerdotai/simplewiki) .\nMight deliberately putting a bunch of encyclopaedic data into the training set make the model\nbetter at the IFT eval (which has lots of factual questions in it, like\n\"who wrote Pride and Prejudice\")?\n- OpenWebText. Even though I was unsure how well it matched the original WebText, given that it was there, it seemed silly to not try training something on it and see how it matched up.\n\nI would train each model on 3.2B tokens of the chosen dataset; that's the Chinchilla-optimal\namount for my 163M-parameter models.  If there were any interesting results, then I might\nconsider doing [overtrained](https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-3-overtraining) models later on.\n\nI decided to be at least vaguely scientific about this, and to pre-register some predictions:\n\n- The FineWeb-Edu-only model would do pretty badly on the test loss, but better than my older FineWeb-Edu models (90%). It would also punch above its weight on the IFT eval (90%).\n- The 50:50 split: I expected it to do worse on the test eval than my JAX FineWeb-only models (70%), but better than the FineWeb-Edu one (90%). I wasn't sure about how it would do on the IFT eval, but thought it might be somewhere in between the two groups (60%).\n- The curated dataset I had high hopes for in terms of the IFT eval -- let's say 80% chance of it being the best of all of my models. For the test loss eval, I expected it to do about as well as the 50:50 split, maybe a little bit worse (70%).\n- I had no idea how the OpenWebText eval would do! Could be worse, could be better.\n\nHere's how things turned out.\n\n### The FineWeb-Edu model\n\nI already had a [dataset based on FineWeb-Edu](https://huggingface.co/datasets/gpjt/fineweb-edu-gpt2-tokens)\nready to go, from when I trained those two original models.  It is just the 10B-token sample of the original\ndataset at the time I generated it last December, formatted appropriately for my training\nscript (details on the dataset card).\n\nI kicked off a training run with my JAX code (which I've been using for the other posts in this series):\n\n``` bash\ngiles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fineweb-edu datasets/\n2026-09-11 18:11:47.991583 Downloading dataset\nFetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1772.93it/s]\nDownload complete: : 0.00B [00:00, ?B/s]                                                                                                            | 0/4 [00:00<?, ?it/s]\n2026-09-11 18:11:48.226273 Loading dataset into RAM\nDownload complete: : 0.00B [00:00, ?B/s]\n2026-09-11 18:16:29.507646 Creating model\n2026-09-11 18:16:33.042509 Creating optimizer\n2026-09-11 18:16:34.138990 Start train\n  0%|                                                                                                                                           | 0/33165 [00:00<?, ?it/s]\n2026-09-11 18:17:38.486288 Saving checkpoint\n  1%|▌                                                                                                     | 173/33165 [13:22<39:17:03,  4.29s/it, loss=6.897, tps=21,201]\n```\n\n...and just less than 40 hours later, I had a model:\n\n```\nTraining complete in 142,912.226 seconds\n2026-09-13 09:58:26.437276 Tokens seen: 3,260,252,160\n2026-09-13 09:58:26.437284 Throughput: 22,813 tokens/second\n2026-09-13 09:58:26.437302 Final train loss: 3.342\n2026-09-13 09:58:26.437309 Done\n```\n\nI converted the saved JAX safetensors file from the last checkpoint into a format that would be compatible with my PyTorch eval code, and ran my smoke test: how would it complete the sentence \"Every effort moves you\"?\n\n```\nEvery effort moves you closer to God’s Kingdom, and even closer to Him.\nAs we can see in\n```\n\nThat was nice and coherent -- if unusually religious! -- so that was promising. I ran the test eval:\n\n``` bash\ngiles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/checkpoints/latest/pytorch-model.safetensors\nFetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 2758.50it/s]\n100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:52<00:00, 13.74it/s]\nLoss against our test dataset: 3.632900\n```\n\nThat was pretty good, putting it at a better test loss than all of the models I had trained\nwithout optimised hyperparameters, and worse than all of the ones I had trained on\nFineWeb *with* optimised hyperparameters.  So that fit in with my prediction that it\nwould be better than the old FineWeb-Edu models; the fact that it was also better than\nthe non-optimised training runs with FineWeb seemed sensible enough\nthat I felt silly for not having predicted that it would have fallen exactly there :-)\n\nI decided to leave the IFT eval until the end so that I could check all of the models\nfrom these experiments together, so it was time to [upload this one to Hugging Face](https://huggingface.co/gpjt/jax-with-mha-bias-fineweb-edu),\nand move on to the next model.\n\n### 50:50 FineWeb to FineWeb-Edu\n\nI put together a new [repo](https://github.com/gpjt/prepare-llm-training-dataset) with\n[a script to prepare datasets specifically for my training setup](https://github.com/gpjt/prepare-llm-training-dataset/blob/9231eb0ec1d1473663cfa892f7a0c0b241fed355/prepare-dataset.py).\nYou provide it with config that specifies some source datasets along with information about\nhow to process them and how to mix them together, and it uploads a new dataset to\nHugging Face Hub with the required characteristics.\n\nFor example, for the 50:50 FineWeb to FineWeb-Edu split, the config looked like this:\n\n```\n{\n    \"seed\": 42,\n    \"tokens_desired\": 10000000000,\n    \"upload_dataset_name\": \"gpjt/fw-fwedu-5050-gpt2-tokens\",\n    \"sources\": [\n        {\n            \"name\": \"FineWeb\",\n            \"hf_id\": \"HuggingFaceFW/fineweb\",\n            \"hf_name\": \"sample-10BT\",\n            \"hf_split\": \"train\",\n            \"item_field\": \"text\",\n            \"weight\": 50\n        },\n        {\n            \"name\": \"FineWeb-Edu\",\n            \"hf_id\": \"HuggingFaceFW/fineweb-edu\",\n            \"hf_name\": \"sample-10BT\",\n            \"hf_split\": \"train\",\n            \"item_field\": \"text\",\n            \"weight\": 50\n        }\n    ]\n}\n```\n\nThe way the script works is pretty simple: it works out (based on those `weight` s and the\n`tokens_desired`) how many tokens it wants from each source dataset, shuffles the items in the sources, then it loops until\nit has the desired number of tokens or more stored in an output.  In the loop, it works out which source is\ncurrently most under-represented, grabs an item from it, tokenises it, and adds it to the output.\n\nRunning it with that 50:50 config seemed to work fine:\n\n``` bash\ngiles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-5050/\nResolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 89875.56it/s]\nLoading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 133.75it/s]\nResolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 87461.48it/s]\nLoading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 200.09it/s]\n2026-09-13 20:13:22.000187: Generating dataset; per-source counts\n2026-09-13 20:13:22.000217: FineWeb: 5,000,000,000\n2026-09-13 20:13:22.000221: FineWeb-Edu: 5,000,000,000\nFineWeb: 100%|████████████████████████████████████████████████████████████████████████████████████████████████▉| 4999999705/5000000000 [1:01:33<00:00, 1353639.33token/s]\nFineWeb-Edu: 5000000363token [1:01:33, 1353639.47token/s]\n2026-09-13 21:14:55.747239:\n\nDone generating tokens\n2026-09-13 21:14:55.748480: FineWeb: 4,999,999,705 / 5,000,000,000 (1.000, 1 iterators)\n2026-09-13 21:14:55.748487: FineWeb-Edu: 5,000,000,363 / 5,000,000,000 (1.000, 1 iterators)\n2026-09-13 21:14:55.748489: Total: 10,000,000,068\n2026-09-13 21:14:55.748491: Catting...\n2026-09-13 21:16:29.565152: Catted into a tensor of shape torch.Size([10000000068])\n2026-09-13 21:16:29.566663: Saving...\n2026-09-13 21:16:36.006267: Saved\n2026-09-13 21:16:36.009413: Uploading to gpjt/fw-fwedu-5050-gpt2-tokens\nProcessing Files (1 / 1)      : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB,  117MB/s\nNew Data Upload               : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 14.6GB / 14.6GB, 98.1MB/s\n  ...du-5050/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB\n2026-09-13 21:17:59.545875: Done\n```\n\nSo we had almost-perfect 50:50 balance between the datasets, and it saved\n[this dataset on Hugging Face](https://huggingface.co/datasets/gpjt/fw-fwedu-5050-gpt2-tokens-DEPRECATED).\n\nI ran [a script to double-check that it looked sane](https://github.com/gpjt/prepare-llm-training-dataset/blob/15331678ebd5149fc89563e940096193792dfb07/check-dataset.py),\nand it did, so it was time to spin up a training run:\n\n``` bash\ngiles@perry:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.90 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-5050 datasets/\n2026-09-13 21:20:59.880918 Downloading dataset\nFetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [01:13<00:00, 36.70s/it]\nDownload complete: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 1.24GB/s]\n2026-09-13 21:22:13.521745 Loading dataset into RAM\nDownload complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 272MB/s]\n2026-09-13 21:22:33.787720 Creating model\n2026-09-13 21:22:35.501063 Creating optimizer\n2026-09-13 21:22:36.043837 Start train\n  0%|                                                                                                                                           | 0/33165 [00:00<?, ?it/s]\n2026-09-13 21:23:11.437206 Saving checkpoint\n  0%|                                                                                                       | 26/33165 [02:20<38:07:05,  4.14s/it, loss=9.308, tps=18,246]\n```\n\nThat was running on `perry`, my normal workstation, and I kicked it off in parallel\nwith the \"curated\" model training run below on `poppy` my training box, but I'll\nkeep the runs separate for the purposes of this writeup.\n\nWhen this had been running for an hour or so, our power went out.  My guess is that\nhaving the tumble dryer running, the car charging, the kettle boiling, the electric hob\nswitched on, and two machines doing training runs is a bit too much for our electrics...\nwhich might be a problem in the future, especially if (as planned) I make `poppy` a\nmulti-GPU machine.\n\nHowever, as things stand, I was able to kick it off again after switching the circuit breaker back on, and things held up.\n\nAgain, about 40 hours later:\n\n```\nTraining complete in 136,060.457 seconds\n2026-09-15 12:05:26.432638 Tokens seen: 3,227,516,928\n2026-09-15 12:05:26.432642 Throughput: 23,721 tokens/second\n2026-09-15 12:05:26.432650 Final train loss: 3.793\n2026-09-15 12:05:26.432653 Done\n```\n\n(Note that the numbers reported at the end of a restarted run like this only include what happened after the restart.)\n\nI converted it to PyTorch-compatible tensors, and did the smoke test:\n\n```\nEvery effort moves you on to other options—in fact, it’s not even worth that effort. Just make\n```\n\nLooking good! Time for the loss test:\n\n``` bash\ngiles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/checkpoints/latest/pytorch-model.safetensors\nFetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1192.07it/s]\n100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:53<00:00, 13.72it/s]\nLoss against our test dataset: 3.462454\n```\n\nThat was almost in keeping with my prediction that it would do worse than the JAX FineWeb-only models, except that it was better than the worst of those, \"JAX, no MHA bias, with dropout\": it was actually better than I predicted.\n\nSo, a promising model.   Time to [upload it to Hugging Face](https://huggingface.co/gpjt/jax-with-mha-bias-fw-fwedu-5050-DEPRECATED) -- and now let's move on\nto the next one.\n\n### The \"curated\" dataset\n\nWith my dataset-preparation script, this was easy enough to set up:\n\n```\n{\n    \"seed\": 42,\n    \"tokens_desired\": 10000000000,\n    \"upload_dataset_name\": \"gpjt/fw-fwedu-simplewiki-gpt2-tokens\",\n    \"sources\": [\n        {\n            \"name\": \"FineWeb\",\n            \"hf_id\": \"HuggingFaceFW/fineweb\",\n            \"hf_name\": \"sample-10BT\",\n            \"hf_split\": \"train\",\n            \"item_field\": \"text\",\n            \"weight\": 45\n        },\n        {\n            \"name\": \"FineWeb-Edu\",\n            \"hf_id\": \"HuggingFaceFW/fineweb-edu\",\n            \"hf_name\": \"sample-10BT\",\n            \"hf_split\": \"train\",\n            \"item_field\": \"text\",\n            \"weight\": 45\n        },\n        {\n            \"name\": \"Simple English Wikipedia\",\n            \"hf_id\": \"answerdotai/simplewiki\",\n            \"hf_name\": \"articles\",\n            \"hf_split\": \"train\",\n            \"item_field\": \"md\",\n            \"weight\": 10\n        }\n    ]\n}\n```\n\nRunning that worked nicely:\n\n``` bash\ngiles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-simplewiki/\nResolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 90196.13it/s]\nLoading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 358.90it/s]\nResolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 88254.11it/s]\nLoading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 589.23it/s]\n2026-09-13 18:59:04.106327: Generating dataset; per-source counts\n2026-09-13 18:59:04.106387: FineWeb: 4,500,000,000\n2026-09-13 18:59:04.106407: FineWeb-Edu: 4,500,000,000\n2026-09-13 18:59:04.106422: Simple English Wikipedia: 1,000,000,000\nFineWeb: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████▉| 4499997964/4500000000 [59:41<00:00, 1256362.56token/s]\nFineWeb-Edu: 4500000607token [59:41, 1256363.31token/s]\nSimple English Wikipedia: 1000002889token [59:41, 279192.58token/s]\n2026-09-13 19:58:45.874744:\n\nDone generating tokens\n2026-09-13 19:58:45.876043: FineWeb: 4,499,997,964 / 4,500,000,000 (1.000, 1 iterators)\n2026-09-13 19:58:45.876048: FineWeb-Edu: 4,500,000,607 / 4,500,000,000 (1.000, 1 iterators)\n2026-09-13 19:58:45.876052: Simple English Wikipedia: 1,000,002,889 / 1,000,000,000 (1.000, 6 iterators)\n2026-09-13 19:58:45.876054: Total: 10,000,001,460\n2026-09-13 19:58:45.876056: Catting...\n2026-09-13 20:00:18.811748: Catted into a tensor of shape torch.Size([10000001460])\n2026-09-13 20:00:18.813169: Saving...\n2026-09-13 20:00:22.773873: Saved\n2026-09-13 20:00:22.773936: Uploading to gpjt/fw-fwedu-simplewiki-gpt2-tokens\nProcessing Files (1 / 1)      : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB,  143MB/s\nNew Data Upload               : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.8GB / 19.8GB,  142MB/s\n  ...plewiki/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB\n2026-09-13 20:01:59.270021: Done\n```\n\nOne thing that is worth noting in that output is the \"6 iterators\" for the Simple English Wikipedia. If a source dataset runs out of items while we're building up the results in this script, we start iterating over it again (with a different seed for the shuffle so that the ordering is different). The \"6 iterators\" means that it needed to do that 6 times -- the original creation of the iterator at the start of the script, and five more. So that means that the Simple English Wikipedia is repeated (oversampled) somewhere between five and six times in the dataset.\n\nThat's not a bad thing! From what I've read, it's actually quite standard to oversample highly educational content in LLM training datasets. And anyway, the dataset the script generated was 10B tokens, of which we're only using 3.2B for the training run in this post, so it would only appear somewhere between one and two times. The repetition would likely only really cut in if and when we did an overtrained model on the dataset.\n\nAnyway, I ran my check against [the uploaded dataset](https://huggingface.co/datasets/gpjt/fw-fwedu-simplewiki-gpt2-tokens-DEPRECATED) -- the first few items were clearly from\nFineWeb, FineWeb-Edu, and the Simple English Wikipedia.\n\nIt was time to kick off a training run:\n\n``` bash\ngiles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki datasets/\n2026-09-13 20:24:48.037024 Downloading dataset\nDownloading (incomplete total...): 0.00B [00:00, ?B/s]                                                                                                                   Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.            | 0/2 [00:00<?, ?it/s]\nWARNING:huggingface_hub.utils._http:Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.\nFetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [02:51<00:00, 85.85s/it]\nDownload complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 435MB/s]\n2026-09-13 20:27:39.934884 Loading dataset into RAM\nDownload complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 116MB/s]\n2026-09-13 20:31:20.492877 Creating model\n2026-09-13 20:31:24.054143 Creating optimizer\n2026-09-13 20:31:25.100832 Start train\n  0%|                                                                                                                                           | 0/33165 [00:00<?, ?it/s]\n2026-09-13 20:32:29.650379 Saving checkpoint\n  0%|▎                                                                                                     | 107/33165 [08:38<39:05:39,  4.26s/it, loss=7.631, tps=20,293]\n```\n\nAgain, this was interrupted by the power outage that hit the 50:50 training run, but I was able to restart from a checkpoint.\n\nAfter another 22 hours, it crashed with an error that I've seen\n[before](https://www.gilesthomas.com/2026/08/chinchilla-check):\n\n```\njax.errors.JaxRuntimeError: INTERNAL: CUDA error: Failed to end stream capture: CUDA_ERROR_STREAM_CAPTURE_INVALIDATED: operation failed due to a previous error during capture [executable_name='jit_train_step']\n```\n\nI put it aside as a one-off oddity when I hit it last time, but this time I dug in\na bit more.  I noted that it had not ever happened on `perry`, but seemed to be an\nissue on `poppy`, and that `poppy` had an older version of CUDA and the Nvidia drivers\n-- might that be the cause?  I decided to upgrade those before kicking off the next\nrun, but for now just restarted the run from the most recent checkpoint.  (Note for\nanyone who is hitting the same error: it has not occurred since the upgrade, so that's worth\ntrying.)\n\nThis time it completed OK:\n\n```\nTraining complete in 59,564.515 seconds\n2026-09-15 15:56:52.909888 Tokens seen: 1,367,212,032\n2026-09-15 15:56:52.909894 Throughput: 22,953 tokens/second\n2026-09-15 15:56:52.909912 Final train loss: 3.332\n2026-09-15 15:56:52.909959 Done\n```\n\nAgain, these numbers just show what happened after the most recent restart.\n\nI copied it over to `perry`, converted it into a format that was compatible\nwith my PyTorch code, and ran the smoke test:\n\n```\nEvery effort moves you by the air, for it will make you a better athlete, so your body becomes bigger and stronger\n```\n\nCoherent enough -- time for the loss eval:\n\n``` bash\ngiles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/checkpoints/latest/pytorch-model.safetensors\nFetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1007.64it/s]\n100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:57<00:00, 13.48it/s]\nLoss against our test dataset: 3.542460\n```\n\nAgain, in line with my predictions -- worse than the JAX FineWeb-only models, and\nindeed than the very best PyTorch one, `1xrtx3090-stacked-interventions`, and also worse\nthan the 50:50 split, but better than the FineWeb-Edu one.\n\nI [uploaded it to Hugging Face](https://huggingface.co/gpjt/jax-with-mha-bias-fw-fwedu-simplewiki-DEPRECATED),\nand it was time to move on to what was meant to be the final model for this set of experiments.\n\n### The OpenWebText run\n\nAgain, this was a simple enough config to set up:\n\n```\n{\n    \"seed\": 42,\n    \"tokens_desired\": 10000000000,\n    \"upload_dataset_name\": \"gpjt/openwebtext-gpt2-tokens\",\n    \"sources\": [\n        {\n            \"name\": \"OpenWebText\",\n            \"hf_id\": \"Skylion007/openwebtext\",\n            \"hf_name\": \"plain_text\",\n            \"hf_split\": \"train\",\n            \"item_field\": \"text\",\n            \"weight\": 50\n        }\n    ]\n}\n```\n\n...and the build and upload process worked well (and took much less time -- for some reason, sampling randomly from a single dataset is faster than sampling from two or three):\n\n``` bash\ngiles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/openwebtext/\nResolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 32723.26it/s]\nResolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 97940.55it/s]\nLoading dataset shards: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 1200.13it/s]\n2026-09-15 13:16:47.622617: Generating dataset; per-source counts\n2026-09-15 13:16:47.622645: OpenWebText: 10,000,000,000\nResolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 45602.65it/s]\nResolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 67650.06it/s]\nLoading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 307.11it/s]\nOpenWebText: 10000000024token [31:46, 5246208.64token/s]\n2026-09-15 13:48:33.761350:\n\nDone generating tokens\n2026-09-15 13:48:33.762021: OpenWebText: 10,000,000,024 / 10,000,000,000 (1.000, 2 iterators)\n2026-09-15 13:48:33.762026: Total: 10,000,000,024\n2026-09-15 13:48:33.762028: Catting...\n2026-09-15 13:49:33.115508: Catted into a tensor of shape torch.Size([10000000024])\n2026-09-15 13:49:33.115923: Saving...\n2026-09-15 13:49:36.365978: Saved\n2026-09-15 13:49:36.366027: Uploading to gpjt/openwebtext-gpt2-tokens\nProcessing Files (0 / 1)      : 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB,  147MB/s\nNew Data Upload               : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.9GB / 19.9GB,  147MB/s\n  ...webtext/train.safetensors: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB\n2026-09-15 13:51:16.202890: Done\n```\n\nNote that it needed to oversample -- that \"2 iterators\". OpenWebText is about 40 GiB uncompressed, and so that's about 10B GPT-2 tokens -- presumably just a little bit less. Again, given that I was planning to use just the first 3.2B tokens of the dataset, I didn't feel that it would matter.\n\nI ran the check script on the newly-uploaded [Hugging Face dataset](https://huggingface.co/datasets/gpjt/openwebtext-gpt2-tokens)\nand all looked well, so that was all set for the training run.\n\nI upgraded `poppy` first with a `sudo pacman -Syu` to see if that helped with the\nweird error that I got in the previous run (which, as I said, it looks like it did), then kicked it off:\n\n``` bash\ngiles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-openwebtext datasets/\n2026-09-15 16:42:32.606185 Downloading dataset\nFetching 2 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:00<00:00, 941.38it/s]\nDownload complete: : 0.00B [00:00, ?B/s]                                                                                                            | 0/2 [00:00<?, ?it/s]\n2026-09-15 16:42:32.879987 Loading dataset into RAM\nDownload complete: : 0.00B [00:00, ?B/s]\n2026-09-15 16:45:40.438791 Creating model\n2026-09-15 16:45:43.840269 Creating optimizer\n2026-09-15 16:45:44.848351 Start train\n  0%|                                                                                                                                           | 0/33165 [00:00<?, ?it/s]\n2026-09-15 16:46:50.632075 Saving checkpoint\n  1%|█                                                                                                     | 332/33165 [24:33<38:45:54,  4.25s/it, loss=6.623, tps=22,154]\n```\n\nAbout 31 hours in, it crashed again, but this time it was my own dumb fault: `poppy`\nhas a relatively small disk and I ran out of space.  I fixed that and kicked it off\nagain from the most recent checkpoint, and this time it completed:\n\n```\nTraining complete in 33,927.995 seconds\n2026-09-17 11:25:10.835989 Tokens seen: 779,747,328\n2026-09-17 11:25:10.835994 Throughput: 22,982 tokens/second\n2026-09-17 11:25:10.836012 Final train loss: 3.165\n2026-09-17 11:25:10.836018 Done\n```\n\nI converted it to PyTorch for the smoke test:\n\n```\nEvery effort moves you through each phase, so it's not a complete picture.\n\nI'm sure your story was\n```\n\n...which looked solid, so it was time for the test loss eval:\n\n``` bash\ngiles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/checkpoints/latest/pytorch-model.safetensors\nFetching 4 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 674.76it/s]\n100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:59<00:00, 13.37it/s]\nLoss against our test dataset: 4.045255\n```\n\nOur worst score yet in this experiment! Worse than any of my models so far, apart from the two FineWeb-Edu ones I did without optimised hyperparameters.\n\nNow, the first draft of this post went straight to the results from here, but the story wasn't quite over yet...\n\n### Test set contamination\n\nGPT-6 Astra is *relentless*.  Before I publish any of these posts, I run them past\nan [editorial board of LLMs](https://www.gilesthomas.com/2026/07/ai-use) to look for issues.  GPT-6 Astra not\nonly checked the text, it also visited the code I'd linked to to check that out too,\nand spotted something problematic.\n\nIt's obvious in retrospect, but my code to build the new datasets had a high risk\nof including the contents of the -- in theory held-back -- test set.  The way that\nthe test set was generated was that\nI downloaded the 10B sample of FineWeb back in [December](https://www.gilesthomas.com/2025/12/llm-from-scratch-28-training-a-base-model-from-scratch),\nsplitting it into 99% training data and 1% \"validation\".  That validation split\nwas about 100M tokens, and I was only using the first 19M or so for actual validation\nruns during training, so I (somewhat arbitrarily) designated about 19M other tokens\nstarting at position 50M in there as my test set.\n\nNow, my new dataset-generation code was just sampling randomly from the complete 10B sample of FineWeb. So there was nothing stopping it from pulling in data that was in that old validation split! That meant that it was quite likely that my new \"curated\" and \"50:50\" datasets contained at least some of the test set that was meant to have been held back from the models during training.\n\nOn reflection, the problem was potentially even worse. FineWeb-Edu is a subset of FineWeb; my existing FineWeb-Edu dataset came from the 10B sample of the Hugging Face original, and so it also could potentially contain documents that I'd put into the test set.\n\nThe first thing to do was to establish the size of the problem.  I wrote\n[a script](https://github.com/gpjt/prepare-llm-training-dataset/blob/4ddd27b787ce3574e217d476ee6878578b27749e/generate-forbidden-doc-hashes.py) to\ntake in a \"forbidden\" dataset and split; this was assumed to be formatted as one big tensor of GPT-2 tokens,\nwhich is what all of my datasets are.  It would then split it by end-of-text tokens, and\ngenerate a hash and a token count for each resulting \"document\".  Optionally, you could\nrestrict it to only considering a subset -- the  tokens starting at position  -- and\nit would then generate hashes/lengths for the documents inside that slice, or that overlapped\nit at the start or the end.\n\nI ran that to generate a list of hashes for the entire validation set -- the validation\nsplit of [`gpjt/fineweb-gpt2-tokens`](https://huggingface.co/datasets/gpjt/fineweb-gpt2-tokens) -- and then\nused [a second script](https://github.com/gpjt/prepare-llm-training-dataset/blob/4ddd27b787ce3574e217d476ee6878578b27749e/check-training-set-for-contamination.py)\nto check my various training sets (and the validation set itself) to see how much\nof a contamination problem there was.  I got these results:\n\n| **Dataset** | **Split** | **Contamination with validation set** | \n|---|---|---|\n| `gpjt/fineweb-gpt2-tokens` | validation | 102163003 out of 102163003 tokens (100.00%) | \n| `gpjt/fineweb-gpt2-tokens` | train | 636166 out of 102163003 tokens (0.62%) | \n| `gpjt/fineweb-edu-gpt2-tokens` | train | 672189 out of 102163003 tokens (0.66%) | \n| `gpjt/fw-fwedu-5050-gpt2-tokens` | train | 49224580 out of 102163003 tokens (48.18%) | \n| `gpjt/fw-fwedu-simplewiki-gpt2-tokens` | train | 44233824 out of 102163003 tokens (43.30%) | \n| `gpjt/openwebtext-gpt2-tokens` | train | 212 out of 102163003 tokens (0.00%) | \n\nSo:\n\n- The validation set was 100% \"contaminated\" with itself, which was a useful sanity check.\n- The training set of `gpjt/fineweb-gpt2-tokens` had what I felt was a small level\nof contamination.  It was interesting that there was any at all -- I think that must\nmean that there are some repeated documents in the original dataset, and some of\nthem wound up with copies in both my training and validation splits.\n- The `gpjt/fineweb-edu-gpt2-tokens` dataset also had what felt like a reassuringly\nlow level of contamination.\n- Both `gpjt/fw-fwedu-5050-gpt2-tokens` and`gpjt/fw-fwedu-simplewiki-gpt2-tokens` ,\nhowever, looked problematic.  In both cases, the training datasets had more than 40% of\nthe validation/test set in them.\n- `gpjt/openwebtext-gpt2-tokens` was, as you'd expect, almost completely uncontaminated.\nIt looks like maybe one document happened to have been picked up by both the\nOpenWebText and the FineWeb crawls and then included in the bit of FineWeb\nI was using for validation.\n\nHowever, these numbers -- while scary, at least for the 50:50 and the curated datasets -- were not quite the ones to use.  They showed\nhow much of the full validation set showed up in the full training set; what I actually\ncared about was how much of the *test* set -- those 19M tokens starting at position 50M in\nthe validation split -- was in the actual subset of the training datasets that I actually\ntrained on -- the first ~3.2B of them.\n\nI re-ran the script to generate hashes for just the test set, and then re-ran the contamination-checking script, telling it just to look at the appropriate subset of the training tokens, and got this:\n\n| **Dataset** (first 3.2B tokens only) | **Split** | **Contamination with test set** | \n|---|---|---|\n| `gpjt/fineweb-gpt2-tokens` | train | 26557 out of 19632681 tokens (0.14%) | \n| `gpjt/fineweb-edu-gpt2-tokens` | train | 32079 out of 19632681 tokens (0.16%) | \n| `gpjt/fw-fwedu-5050-gpt2-tokens` | train | 2986889 out of 19632681 tokens (15.21%) | \n| `gpjt/fw-fwedu-simplewiki-gpt2-tokens` | train | 2682430 out of 19632681 tokens (13.66%) | \n| `gpjt/openwebtext-gpt2-tokens` | train | None | \n\nIt was clear that there was a problem -- certainly with `gpjt/fw-fwedu-5050-gpt2-tokens`\nand `gpjt/fw-fwedu-simplewiki-gpt2-tokens`.  They'd seen what felt like a significant amount\nof the test set while training, so their results on the test loss eval were dubious\nat best.\n\nI decided to train those two models afresh, and see what the result was in terms of loss.\nIf the difference was huge, I'd look into the risks of the (much smaller) contamination\nof `gpjt/fineweb-gpt2-tokens` and `gpjt/fineweb-edu-gpt2-tokens`.  But if it was pretty\nsmall, I'd not worry about that too much.\n\nI extended the script that [prepared datasets](https://github.com/gpjt/prepare-llm-training-dataset/blob/6dcaa02f5019a78631567748872dd8fdcfc7c35f/prepare-dataset.py)\nso that the config file could specify a `forbidden_dataset`.  Any documents in the\nsource datasets that matched forbidden ones would be excluded from the output.  I then\nupdated the config for `gpjt/fw-fwedu-5050-gpt2-tokens` and `gpjt/fw-fwedu-simplewiki-gpt2-tokens`\nso that the whole validation split of `gpjt/fineweb-gpt2-tokens` was forbidden, and\nre-generated them.  You can see the updated datasets [here](https://huggingface.co/datasets/gpjt/fw-fwedu-5050-gpt2-tokens)\nand [here](https://huggingface.co/datasets/gpjt/fw-fwedu-simplewiki-gpt2-tokens).\nRunning the contamination-checker script against them showed that they were clear.\n\nI then re-did the full training runs for those models; the uncontaminated version\nof the 50:50 split model is [here](https://huggingface.co/gpjt/jax-with-mha-bias-fw-fwedu-5050),\nand the curated one is [here](https://huggingface.co/gpjt/jax-with-mha-bias-fw-fwedu-simplewiki).\n\nAnd the good news: both of them actually did very slightly *better* at the test\nloss eval than their equivalents that had been trained on the contaminated data:\n\n| **Model** | **Contaminated** | **Test loss** | \n|---|---|---|\n| JAX, FineWeb/FineWeb-Edu 50:50 | No | 3.449257 | \n| JAX, FineWeb/FineWeb-Edu 50:50 | Yes | 3.462454 | \n| JAX, curated | No | 3.534068 | \n| JAX, curated | Yes | 3.542460 | \n\nThere are a number of possibilities that come to mind; perhaps learning from the test set just doesn't happen with tiny 163M models like this, or perhaps while the contaminated models were learning, the benefit they got from that was outweighed by the data that they got instead of the test set data being in some way better for training purposes, at least in terms of the loss eval.\n\nBut anyway, I felt that if the effect of seeing more than 10% of the test set data during training was so tiny, then the effect of seeing less than 0.2% -- which is what the FineWeb-Edu model in this set of training runs had, as did all of my other FineWeb-only models from previous experiments -- would be even smaller and I'd disregard it.\n\nThat was excellent news! I didn't need to start all of my experiments from scratch.\n\nFor the rest of this post, I will include the numbers and results for the contaminated models as well as the uncontaminated ones -- they're interesting for several reasons -- but for future posts I'll skip the contaminated ones.\n\nSo -- finally! -- let's start digging into the final results.\n\n### Results\n\nFirstly, I think it's worth taking a look at all of the test loss results in context. Here they are in a table, with the new models in bold:\n\n|  | Test loss | \n|---|---|\n| OpenAI weights: medium | 3.231442 | \n| JAX, overtrained one long epoch | 3.324953 | \n| JAX, overtrained two normal epochs | 3.326482 | \n| JAX, with MHA bias, no dropout | 3.418784 | \n| JAX, no MHA bias, no dropout | 3.420089 | \n| **JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated)** | 3.449257 | \n| **JAX, FineWeb/FineWeb-Edu 50:50 (contaminated)** | 3.462454 | \n| JAX, no MHA bias, with dropout | 3.476802 | \n| OpenAI weights: small | 3.499677 | \n| **JAX, curated (uncontaminated)** | 3.534068 | \n| `1xrtx3090-stacked-interventions` | 3.538161 | \n| **JAX, curated (contaminated)** | 3.542460 | \n| `8xa100m40-stacked-interventions-1` | 3.577761 | \n| **JAX, FineWeb-Edu** | 3.632900 | \n| Cloud FineWeb, 8x A100 40 GiB | 3.673623 | \n| `1xrtx3090-baseline` | 3.683835 | \n| `8xa100m40-baseline` | 3.691526 | \n| Cloud FineWeb, 8x H100 80 GiB | 3.724507 | \n| Cloud FineWeb, 8x A100 80 GiB | 3.729900 | \n| Cloud FineWeb, 8x B200 160 GiB | 3.771478 | \n| Local FineWeb train | 3.943522 | \n| **JAX, openwebtext** | 4.045255 | \n| Local FineWeb-Edu extended train | 4.134991 | \n| Local FineWeb-Edu train | 4.166892 | \n\nI think there's something very clear here: with the new models, the more FineWeb that was in the training mix, the better the model did on this eval. I think I might have been subconsciously expecting that in the predictions I did before running these experiments, but in retrospect it's so incredibly obvious that I feel silly for not mentioning it explicitly!\n\nBut that tells us something interesting.  From the description in the paper, whatever OpenAI\ndid the GPT-2 training run on, it was not like FineWeb.  It was *probably* more\nsimilar to OpenWebText -- and yet, that model was the one that performed the worst on\nthis test eval, so if it is more like OpenWebText, there must be some other factor involved.\n\nBut moving on for now: how about the IFT test -- the one that kicked off all of this work in the first place?\n\nI generated a set of IFT responses for all of the new models, and then ran them (plus responses for all of the other models on that table above) past GPT 5.5, and found that one of my new models was getting quite close to the original GPT-2 small weights! So I did four more runs, so that I could get an average.\n\nHere are the results -- the \"IFT score\" is the average across all five runs of the judge, and the \"IFT rank\" is based on that. The \"IFT epochs\" was from the original result-generation script.\n\n|  | Test loss | IFT epochs | IFT score | IFT rank | \n|---|---|---|---|---|\n| OpenAI weights: medium | 3.231442 | 2 | 42.36 | 1 | \n| JAX, overtrained one long epoch | 3.324953 | 3 | 18.67 | 7 | \n| JAX, overtrained two normal epochs | 3.326482 | 4 | 18.71 | 6 | \n| JAX, with MHA bias, no dropout | 3.418784 | 4 | 17.90 | 8 | \n| JAX, no MHA bias, no dropout | 3.420089 | 5 | 20.50 | 4 | \n| **JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated)** | 3.449257 | 4 | 17.69 | 9 | \n| **JAX, FineWeb/FineWeb-Edu 50:50 (contaminated)** | 3.462454 | 4 | 19.30 | 5 | \n| JAX, no MHA bias, with dropout | 3.476802 | 5 | 13.02 | 21 | \n| OpenAI weights: small | 3.499677 | 2 | 25.19 | 2 | \n| **JAX, curated (uncontaminated)** | 3.534068 | 4 | 16.63 | 10 | \n| `1xrtx3090-stacked-interventions` | 3.538161 | 4 | 13.51 | 19 | \n| **JAX, curated (contaminated)** | 3.542460 | 4 | 13.58 | 18 | \n| `8xa100m40-stacked-interventions-1` | 3.577761 | 4 | 10.19 | 24 | \n| **JAX, FineWeb-Edu** | 3.632900 | 4 | 24.56 | 3 | \n| Cloud FineWeb, 8x A100 40 GiB | 3.673623 | 3 | 16.59 | 11 | \n| `1xrtx3090-baseline` | 3.683835 | 4 | 15.15 | 12 | \n| `8xa100m40-baseline` | 3.691526 | 3 | 13.64 | 16 | \n| Cloud FineWeb, 8x H100 80 GiB | 3.724507 | 4 | 13.59 | 17 | \n| Cloud FineWeb, 8x A100 80 GiB | 3.729900 | 3 | 10.79 | 23 | \n| Cloud FineWeb, 8x B200 160 GiB | 3.771478 | 4 | 13.70 | 15 | \n| Local FineWeb train | 3.943522 | 5 | 11.87 | 22 | \n| **JAX, openwebtext** | 4.045255 | 4 | 13.28 | 20 | \n| Local FineWeb-Edu extended train | 4.134991 | 5 | 14.29 | 14 | \n| Local FineWeb-Edu train | 4.166892 | 5 | 14.69 | 13 | \n\nIf you want to see the full numbers, they're [below](#appendix-all-ift-judge-runs).\n\nThe number that initially surprised me, and made me decide to do multiple LLM-judge runs was the one for the \"JAX, FineWeb-Edu\" model. In my first run it came in at 24.35 vs the OpenAI small weights' 24.93 -- so close that I wondered if it might even beat them on a re-run. However, in the further four runs its score was consistently lower than the OpenAI model's, and the gap extended a bit in some.\n\nSo, was FineWeb-Edu the clear winner here? Perhaps. If you look at the contaminated/uncontaminated pairs, something interesting pops out. For the 50:50 mix, the model trained with the contaminated dataset got 19.30, and the one trained on the uncontaminated one got 17.69 -- a difference of 1.61. For the \"curated\" dataset, the situation was even more interesting: uncontaminated got 16.63, while contaminated got 13.58, a delta of 3.05 points.\n\nRemember, the contamination issue is about whether or not the model saw the held-back test set during training. It was an issue for the test loss that is based on that test set, but is entirely orthogonal to the IFT test.\n\nFrom the IFT perspective, both contaminated and uncontaminated models in each case saw training data that was -- in theory, at least -- essentially the same in terms of quality. Indeed, the uncontaminated run saw almost the same data in the same order as the contaminated one, except that some items were omitted, and then extra ones were added to the end.\n\nThe purpose of this set of experiments was to see how data quality affected the results on the IFT test set. But in the case of the curated model, something that should be unrelated to data quality changed the results by 3.05 points!\n\nIf something as simple as changing which data of the same quality the model is trained with can affect the IFT score so drastically, it makes it a bit harder to be certain as to whether or not data quality really had the effect we were looking for.\n\nOn the other hand, the FineWeb-Edu model came in at 24.56, which is 4.06 points better than the 20.50 that the closest other model got -- more than the 3.05 points we see in difference between the two curated dataset models. And it's worth noting that the model with 20.50 is \"JAX, no MHA bias, no dropout\", which has a subtly different architecture -- no bias on the output projection of the multi-head attention blocks. A better comparison might be \"JAX, with MHA bias, no dropout\", which got a score of 17.90, for a whacking great difference of 6.66 points.\n\nI think that without doing a very large number of training runs on different datasets with different mixes, each one created with a different seed, it would be hard to work out exactly what is in the noise here and what is not.\n\nHowever, that would cost a lot in terms of time. I think that the best thing here is to chalk this up as a fairly decent indication that FineWeb-Edu improves matters for the IFT eval, but far from a certainty. But it's certainly worth noting that whatever the noise is, it has a range of at least 3.05 points -- and the FineWeb-Edu model is just 0.63 points short of GPT-2 small! So there could well be something there. Of course, we don't know whether that model got (by chance) the best possible balance of FineWeb-Edu tokens, and could never win -- or whether it got a bad balance and would actually beat GPT-2 with a better one. So that's certainly worth keeping in mind.\n\nAs an aside, the result for the curated dataset really surprised me. I had expected that it would be the best one, simply because it almost certainly contained more facts. I took a look at its answers to the questions -- one possibility that came to mind might be that it would get better responses to questions like \"What is the chemical symbol for chlorine\" or \"Who wrote Pride and Prejudice\" than the others, but would fail on less knowledge-based tasks. But it was terrible at fact-based questions too:\n\nName the author of 'Pride and Prejudice'.\n\n  The author of 'Pride and Prejudice' is Priscilla Finch.\n\nWhat is the periodic symbol for chlorine?\n\n  The periodic symbol for chlorine is H.\n\nAs I understand it, many real-world training runs do include (often oversampled)\namounts of highly educational training data like this model's dataset did.  But perhaps\nthe models that I'm training are just too small to be able to make use of the data they gained that\nway -- maybe doing things this way and expecting\ngood results is like asking six-year-old children to memorise stuff before they've learned enough\nto be able to make use of it <sup>[1](#fn-1)</sup>.  It's worth noting that the GPT-2 small model also\nfailed on those factual questions.\n\nWell, anyway: I think we have some useful results here, so let's work out what that means for next steps.\n\n### Conclusion\n\nThe results we got in these experiments point in two interesting directions.\n\n1. The perfect connection between the amount of FineWeb in the training set and the result on the (FineWeb-based) test loss eval, while perfectly obvious in retrospect, really does highlight how mysterious it is that the OpenAI small weights do so well on that test.\n2. The fact that FineWeb-Edu did well on the IFT test tells us that there does seem to be value in using richer training data -- though the less-spectacular results of the 50:50 mix and the curated one weaken that a bit, as does the indicator of what the noise due to data selection from equivalently high-quality datasets might be. The OpenWebText result I think I'll ignore, given that -- while in theory it should be similar to what OpenAI trained on -- there are no guarantees, and it might differ in non-obvious ways for non-obvious reasons.\n\nI think that the right direction to take this going forward is to separate these two angles.\n\nI should chase a higher IFT score, and then once I have nailed that down, I should see what (if anything) might allow me to get the resulting model to improve its test score. But I will need to make sure that whatever dataset I use, I use various \"mixes\" of it -- versions created with different random seeds.\n\nIn my earlier experiments\nwith overtraining, I did find that it didn't seem to improve the IFT results -- but it\n*did* improve the test loss.  So perhaps identifying the right combination of other factors\nto boost the IFT score, then overtraining the result, might help?\n\nOf course, my overtraining tests were with FineWeb, so the connection might not hold up as well if the starting model (as seems likely) was trained on a different dataset.\n\nAlso, while working through the results here, I've come to the conclusion that the set of models I'm using is a bit confusing -- there are now different hyperparameter settings, small architectural differences (the MHA bias thing), dropout settings during the pre-training, and now datasets. I think that's OK for now; I should see this part of this series as more ideation than actually running the proper experiments. But at the end, when I have some solid hypotheses with a reasonable amount of backup, I should start from scratch: a baseline model, then staged interventions to build up to what (hopefully) will be a model as good as GPT-2 small.\n\nAnyway, I'll wrap this one up here.  I think that the next lever to pull is (perhaps\nsurprisingly) going to be weight tying.  I had previously kind of disregarded that as\na possibility, but while I was working on this post, something popped into my mind.\nThe OpenAI models were originally trained with weight tying.  My codebase does actually\nsupport doing it -- but because I got the OpenAI weights I'm using from the code in\n\"[Build a Large Language Model (from Scratch)](https://www.manning.com/books/build-a-large-language-model-from-scratch)\",\nwhen I'm running the IFT test, the weights are not actually tied!  We load up a model\nthat has separate but identical embedding and output head matrices, and then we fine-tune\nthat.  So those two matrices can vary independently during fine-tuning -- to put it\nanother way, while GPT-2 small was pre-trained with 124M parameters, the IFT test is\nbeing done on a 163M-parameter version.  Does that\ngive them some non-obvious advantage?  And would adding weight-tying to my own models\nhelp, either with or without the output heads being independent at fine-tuning time?\n\nStay tuned :-)\n\n### Appendix: all IFT judge runs\n\nHere are the numbers for all of the IFT judge runs, included for completeness. You can see that the LLM judge ranks models very consistently between runs, but there is variation -- that is, on some runs it's in what I think of as a \"better mood\" than others, and if that's the case, it will give better scores -- but it will give them almost consistently between models, so all of the models do better. Note that (unlike the table above) this one is sorted by the average IFT score rather than the test loss.\n\n| **Model** | **Run 1** | **Run 2** | **Run 3** | **Run 4** | **Run 5** | **Average** | \n|---|---|---|---|---|---|---|\n| OpenAI weights: medium | 42.24 | 42.16 | 42.95 | 41.83 | 42.61 | 42.36 | \n| OpenAI weights: small | 24.93 | 24.96 | 25.39 | 25.01 | 25.66 | 25.19 | \n| **JAX, FineWeb-Edu** | 24.35 | 24.55 | 24.3 | 24.68 | 24.9 | 24.56 | \n| JAX, no MHA bias, no dropout | 20.5 | 19.9 | 20.76 | 21.25 | 20.07 | 20.50 | \n| **JAX, FineWeb/FineWeb-Edu 50:50 (contaminated)** | 19.16 | 18.86 | 19.61 | 19.17 | 19.7 | 19.30 | \n| JAX, overtrained two normal epochs | 18.47 | 18.29 | 19.17 | 18.69 | 18.91 | 18.71 | \n| JAX, overtrained one long epoch | 18.04 | 18.71 | 19.62 | 18.41 | 18.57 | 18.67 | \n| JAX, with MHA bias, no dropout | 17.49 | 17.35 | 18.33 | 17.73 | 18.62 | 17.90 | \n| **JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated)** | 17.37 | 17.73 | 17.53 | 18.01 | 17.83 | 17.69 | \n| **JAX, curated (uncontaminated)** | 16.77 | 16.03 | 17.3 | 16.08 | 16.96 | 16.63 | \n| Cloud FineWeb, 8x A100 40 GiB | 16.44 | 16.23 | 17.14 | 16.62 | 16.54 | 16.59 | \n| `1xrtx3090-baseline` | 14.85 | 15.07 | 15.19 | 15.14 | 15.51 | 15.15 | \n| Local FineWeb-Edu train | 14.37 | 14.23 | 15.08 | 14.79 | 15 | 14.69 | \n| Local FineWeb-Edu extended train | 14.4 | 14.07 | 13.82 | 14.56 | 14.61 | 14.29 | \n| Cloud FineWeb, 8x B200 160 GiB | 13.37 | 13.05 | 13.85 | 13.67 | 14.57 | 13.70 | \n| `8xa100m40-baseline` | 13.64 | 13.36 | 13.9 | 13.32 | 13.97 | 13.64 | \n| Cloud FineWeb, 8x H100 80 GiB | 13.45 | 13.32 | 13.6 | 13.51 | 14.07 | 13.59 | \n| **JAX, curated (contaminated)** | 13.09 | 13.48 | 13.95 | 13.23 | 14.15 | 13.58 | \n| `1xrtx3090-stacked-interventions` | 13.37 | 13.11 | 14.04 | 13.84 | 13.17 | 13.51 | \n| **JAX, openwebtext** | 12.88 | 12.7 | 13.74 | 13.53 | 13.53 | 13.28 | \n| JAX, no MHA bias, with dropout | 13.19 | 12.86 | 12.98 | 12.85 | 13.24 | 13.02 | \n| Local FineWeb train | 11.75 | 11.75 | 12.21 | 11.46 | 12.19 | 11.87 | \n| Cloud FineWeb, 8x A100 80 GiB | 10.68 | 10.2 | 11.03 | 10.55 | 11.49 | 10.79 | \n| `8xa100m40-stacked-interventions-1` | 9.44 | 9.79 | 10.84 | 10.2 | 10.66 | 10.19 | \n\n1. \nA small boy asleep on his right side, the right arm stuck out, the right hand hanging limp over the edge of the bed. Through a round grating in the side of a box a voice speaks softly. \"The Nile is the longest river in Africa and the second in length of all the rivers of the globe. Although falling short of the length of the Mississippi-Missouri, the Nile is at the head of all rivers as regards the length of its basin, which extends through 35 degrees of latitude …\" At breakfast the next morning, \"Tommy,\" some one says, \"do you know which is the longest river in Africa?\" A shaking of the head. \"But don't you remember something that begins: The Nile is the …\" \"The - Nile - is - the - longest - river - in - Africa - and - the - second - in - length - of - all - the - rivers - of - the - globe …\" The words come rushing out. \"Although - falling - short - of …\" \"Well now, which is the longest river in Africa?\" The eyes are blank. \"I don't know.\" \"But the Nile, Tommy.\" \"The - Nile - is - the - longest - river - in - Africa - and - second …\" \"Then which river is the longest, Tommy?\" Tommy burst into tears. \"I don't know,\" he howls. [*Brave New World*](https://www.huxley.net/bnw/two.html) , Aldous Huxley[↩](#fnref-1)\n\n## Citing this post\n\nThis is a blog, and if you want to link to this post then please do :-) However, if you're writing something more academic and need to do a proper citation, then here's a BibTeX block to make things easier.\n\n```\n@misc{thomas2026oct-why-do-openai-gpt2-weights-beat-mine-5-data-quality,\n  author       = {Thomas, Giles},\n  title        = {{Why do OpenAI's GPT-2 weights beat mine?  Part five: data quality}},\n  year         = {2026},\n  month        = oct,\n  howpublished = {Blog post},\n  url          = {https://www.gilesthomas.com/2026/10/why-do-openai-gpt2-weights-beat-mine-5-data-quality},\n}\n```\n\n", "url": "https://wpnews.pro/news/why-do-openai-s-gpt-2-weights-beat-mine-part-five-data-quality", "canonical_source": "https://www.gilesthomas.com/2026/10/why-do-openai-gpt2-weights-beat-mine-5-data-quality", "published_at": "2026-10-01 16:40:52+00:00", "updated_at": "2026-10-01 16:47:45.472646+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research", "natural-language-processing"], "entities": ["OpenAI", "GPT-2", "Sebastian Raschka", "Alpaca", "GPT 5.5", "Reddit"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/why-do-openai-s-gpt-2-weights-beat-mine-part-five-data-quality", "markdown": "https://wpnews.pro/news/why-do-openai-s-gpt-2-weights-beat-mine-part-five-data-quality.md", "text": "https://wpnews.pro/news/why-do-openai-s-gpt-2-weights-beat-mine-part-five-data-quality.txt", "jsonld": "https://wpnews.pro/news/why-do-openai-s-gpt-2-weights-beat-mine-part-five-data-quality.jsonld"}}