Why do OpenAI's GPT-2 weights beat mine? Part five: data quality Independent LLM-from-scratch experiments found that OpenAI's GPT-2 small (124M parameters) consistently outperformed the author's 163M-parameter models on an instruction fine-tuning task adapted from Sebastian Raschka's book, despite using the same architecture. The author ruled out overtraining and dropout handling as explanations in earlier parts of the series, and is now investigating training data quality, noting OpenAI never released the GPT-2 dataset and only described it in the GPT-2 paper as a web scrape emphasizing document quality, seeded from Reddit outbound links. The author's models differ from GPT-2 small by omitting weight-tying and QKV bias. Why do OpenAI's GPT-2 weights beat mine? Part five: data quality When I finished learning how to build an LLM from scratch https://www.gilesthomas.com/llm-from-scratch , I was left with a mystery: my own models were not as good as OpenAI's original GPT-2 models, despite being based on the same architecture. My models all had 163M parameters, and followed the design from Sebastian Raschka https://sebastianraschka.com/ 's book " Build a Large Language Model from Scratch https://www.manning.com/books/build-a-large-language-model-from-scratch ". That meant that they were pretty much the same as the setup for the OpenAI GPT-2 "small" instance, except that they did not use weight-tying or bias on the QKV matrices. Weight-tying means that you re-use the initial embedding matrix as the output head at the end, and using it means that GPT-2 small saved quite a few parameters -- it was 124M rather than 163M -- at, at least in my own experiments, a cost in quality https://www.gilesthomas.com/2026/03/llm-from-scratch-32g-interventions-weight-tying ; similarly, while I found that QKV bias made a tiny improvement in loss terms https://www.gilesthomas.com/2026/02/llm-from-scratch-32d-interventions-adding-attention-bias , I'd felt it was likely within the noise. But GPT-2 small consistently beat my models on an instruction fine-tuning IFT task -- also adapted from Raschka's book. That test fine-tunes the model on a subset of the Alpaca https://github.com/tatsu-lab/stanford alpaca dataset, until validation loss starts rising, and then runs a test set through the resulting model. The responses to the test set questions are stored, and then I run all of the responses from all of the models under test past GPT 5.5 in one go to get an aggregate score; more details here https://www.gilesthomas.com/2026/04/llm-from-scratch-32l-interventions-instruction-fine-tuning-tests . GPT-2 small always did better than any of my models on this. Additionally, it did surprisingly well on a simpler eval -- one that just measured the cross entropy loss it got on a test set. It scored close to my own best models, and better than many of them. What made this result particularly interesting was that the test set in question was a split of my own training data; my models would not have seen it when training at least, in theory , but it seems likely that it would be much more similar to their own training data than it was to OpenAI's. I've checked two things while probing this mystery: - It seems very likely that the GPT-2 models were overtrained by modern standards; would overtraining my own https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-3-overtraining models get them closer? It turned out that no, it probably didn't help with the IFT eval though there might have been some signal there . It did help quite a lot with the test loss eval, though. - The way I was handling dropout https://www.gilesthomas.com/2026/08/why-do-openai-gpt2-weights-beat-mine-4-ift-dropout in the IFT test might have been unduly benefiting some models while working against others. I decided to standardise on not using dropout during this eval, as counter-intuitively for me it seemed to harm the results of most models, even those that had been pre-trained with dropout. In particular, the OpenAI weights were harmed by using dropout, and making a change that benefited them along with some of my own models seemed the most conservative approach to take in investigating this. The next thing I wanted to look into was the training data. The exact dataset that the various GPT-2 models were trained on has never been released; all we know about it is from the paper https://cdn.openai.com/better-language-models/language models are unsupervised multitask learners.pdf , where they say: W e created a new web scrape which emphasizes document quality. To do this we only scraped web pages which have been curated/filtered by humans. Manually filtering a full web scrape would be exceptionally expensive so as a starting point, we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma. This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny. They called it "WebText". There is an OpenWebText https://skylion007.github.io/OpenWebTextCorpus/ that tries to replicate it, but although they tried to follow the same procedure as the original, there's no guarantee that it is all that similar. By comparison, I'd normally been training against FineWeb https://huggingface.co/datasets/HuggingFaceFW/fineweb . While this is a general web-scraping dataset, without the "curation" provided by using only stuff that was linked from upvoted Reddit posts, it has been refined to remove any obvious junk. I had felt that it was pretty much equivalent. But what if I were wrong about that? I decided to see if I could get better models by using better data. The starting point Here's a table of all of the models I've been comparing to date. The "Test loss" column shows how well the model in question did on that held-back cross entropy loss evaluation. The "IFT epochs" column shows how many epochs of fine-tuning the model needed before its validation loss started rising, the "IFT score" the score that GPT 5.5 gave the model's responses to the test set of my Alpaca data, and the "IFT rank" the model's rank in terms of that score. The OpenAI small model is in there in bold, and I've also included the OpenAI medium model for comparison purposes. | | Test loss | IFT epochs | IFT score | IFT rank | |---|---|---|---|---| | OpenAI weights: medium | 3.231442 | 2 | 43.75 | 1 | | JAX, overtrained one long epoch | 3.324953 | 3 | 19.77 | 4 | | JAX, overtrained two normal epochs | 3.326482 | 4 | 19.72 | 5 | | JAX, with MHA bias, no dropout | 3.418784 | 4 | 18.69 | 6 | | JAX, no MHA bias, no dropout | 3.420089 | 5 | 21.46 | 3 | | JAX, no MHA bias, with dropout | 3.476802 | 5 | 13.22 | 15 | | OpenAI weights: small | 3.499677 | 2 | 26.00 | 2 | | 1xrtx3090-stacked-interventions | 3.538161 | 4 | 13.77 | 14 | | 8xa100m40-stacked-interventions-1 | 3.577761 | 4 | 10.76 | 18 | | Cloud FineWeb, 8x A100 40 GiB | 3.673623 | 3 | 17.72 | 7 | | 1xrtx3090-baseline | 3.683835 | 4 | 15.74 | 8 | | 8xa100m40-baseline | 3.691526 | 3 | 14.19 | 13 | | Cloud FineWeb, 8x H100 80 GiB | 3.724507 | 4 | 14.33 | 12 | | Cloud FineWeb, 8x A100 80 GiB | 3.729900 | 3 | 11.34 | 17 | | Cloud FineWeb, 8x B200 160 GiB | 3.771478 | 4 | 14.67 | 11 | | Local FineWeb train | 3.943522 | 5 | 12.31 | 16 | | Local FineWeb-Edu extended train | 4.134991 | 5 | 15.04 | 9 | | Local FineWeb-Edu train | 4.166892 | 5 | 14.99 | 10 | You can see that the OpenAI small model did pretty well in terms of the test loss, when you consider that it has 39M fewer weights than my models and was being tested against a dataset that differs more from its likely training data than it does from my own models'. Additionally, the specific models that did better than OpenAI's small one were all trained with JAX rather than PyTorch -- my hypothesis for that is that it's a result of the JAX ones getting better initial weights by pure chance. But the big difference was in the IFT score. In the specific run that gave the results in this table, the OpenAI small model got 26.00 -- the closest of my own models was more than 4.5 points lower, at 21.46. This difference was consistent over all of my other test runs. The GPT-2 small model was always ahead of mine. GPT-2 medium, of course, beat GPT-2 small and all of my models, but given that it is twice the size of mine, that's not a big surprise. Now, quite some time ago, I had tried looking into data quality as a lever to pull for model performance. At the bottom of the table, with the worst test loss of all models, you can see two models: - "Local FineWeb-Edu train" - "Local FineWeb-Edu extended train" These two were as you might guess from the names trained on the FineWeb-Edu https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu dataset, which includes just the most "educational" data from FineWeb. They scored very badly on the test loss score. Given that the test dataset is from FineWeb, that's not a big surprise -- as I've written previously: If you train a model on Jane Austen and then evaluate against Chuck Tingle https://www.chucktingle.com/ , then you're not going to get amazing results. But again, GPT-2 had the same issue, and did perfectly well on the test loss eval. On the other hand, while these FineWeb-Edu models' performance on the IFT eval wasn't stellar -- there are plenty of my other models ahead of them -- they did seem to punch above their weight. Consistently across all of the IFT evals I've done, they have scored higher than many of the others -- despite their poor loss on the test eval. Additionally: they were amongst the first models that I trained, before I'd spent time learning about how to optimise my hyperparameters and training loop https://www.gilesthomas.com/2026/04/llm-from-scratch-32m-interventions-conclusion . They did not use gradient clipping, they did use dropout, their batch size was just "whatever I could squeeze into the GPU", and I didn't set the learning rate to the right kind of value or schedule it over the course of the training run. So maybe a new training run on FineWeb-Edu plus my training improvements would help? And maybe some other tweaks to the training data would be worth looking into? The plan I decided to see what would happen if I trained some models with better-quality data. Specifically, I would train models with my current optimised loop and hyperparameters on four different datasets: - FineWeb-Edu -- essentially the same as "Local FineWeb-Edu train" but with a better training setup. This would test the "more educational - better" hypothesis. - A 50:50 split of FineWeb and FineWeb-Edu. I've read that LLMs can be helped by having a decent amount of lower-quality data in their training loop, as it helps them to generalise. Perhaps having some FineWeb in there in addition to the FineWeb-Edu stuff would improve that test loss score while also helping the IFT test? - A "curated" dataset containing 45% of its contents from FineWeb, 45% from FineWeb-Edu, and 10% from the Simple English Wikipedia https://simple.wikipedia.org/wiki/Main Page . The full Wikipedia is huge, and full of obscure facts -- while the Simple English one is small and hopefully richer in useful information on a per-token basis. And conveniently, Answer.ai have made a snapshot of it available on Hugging Face Hub https://huggingface.co/datasets/answerdotai/simplewiki . Might deliberately putting a bunch of encyclopaedic data into the training set make the model better at the IFT eval which has lots of factual questions in it, like "who wrote Pride and Prejudice" ? - OpenWebText. Even though I was unsure how well it matched the original WebText, given that it was there, it seemed silly to not try training something on it and see how it matched up. I would train each model on 3.2B tokens of the chosen dataset; that's the Chinchilla-optimal amount for my 163M-parameter models. If there were any interesting results, then I might consider doing overtrained https://www.gilesthomas.com/2026/07/why-do-openai-gpt2-weights-beat-mine-3-overtraining models later on. I decided to be at least vaguely scientific about this, and to pre-register some predictions: - The FineWeb-Edu-only model would do pretty badly on the test loss, but better than my older FineWeb-Edu models 90% . It would also punch above its weight on the IFT eval 90% . - The 50:50 split: I expected it to do worse on the test eval than my JAX FineWeb-only models 70% , but better than the FineWeb-Edu one 90% . I wasn't sure about how it would do on the IFT eval, but thought it might be somewhere in between the two groups 60% . - The curated dataset I had high hopes for in terms of the IFT eval -- let's say 80% chance of it being the best of all of my models. For the test loss eval, I expected it to do about as well as the 50:50 split, maybe a little bit worse 70% . - I had no idea how the OpenWebText eval would do Could be worse, could be better. Here's how things turned out. The FineWeb-Edu model I already had a dataset based on FineWeb-Edu https://huggingface.co/datasets/gpjt/fineweb-edu-gpt2-tokens ready to go, from when I trained those two original models. It is just the 10B-token sample of the original dataset at the time I generated it last December, formatted appropriately for my training script details on the dataset card . I kicked off a training run with my JAX code which I've been using for the other posts in this series : bash giles@poppy:~/Dev/jax-gpt2-from-scratch main $ XLA PYTHON CLIENT MEM FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fineweb-edu datasets/ 2026-09-11 18:11:47.991583 Downloading dataset Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 00:00<00:00, 1772.93it/s Download complete: : 0.00B 00:00, ?B/s | 0/4 00:00