Fun with low-rank vocab matrices (and a bonus test loss reduction?) Training GPT-2 small-style models with low-rank factorised embeddings on the input side alone reduced test loss, while applying the trick to both embeddings and the output head raised loss by 0.07, 0.10 and 0.14 across three tests, according to an experiment by the model's creator. The author was prompted by Hugging Face user AndrewThompson1233, who reported that at rank 128 on a 50k vocabulary the validation cross-entropy penalty is typically within +0.02 to +0.04 loss, or under 0.5 perplexity. The finding matters because embeddings and the output head account for roughly 39 million of the 163 million parameters in the author's GPT-2 small-style models, about half the model, or 30% under weight tying. Fun with low-rank vocab matrices and a bonus test loss reduction? I was nerdsniped https://xkcd.com/356/ On the HF discussion page for one of the models I created for my previous post https://www.gilesthomas.com/2026/10/why-do-openai-gpt2-weights-beat-mine-5-data-quality , AndrewThompson1233 https://huggingface.co/AndrewThompson1233 asked if I'd considered trying out factorised embeddings https://huggingface.co/gpjt/jax-with-mha-bias-fw-fwedu-5050-DEPRECATED/discussions/1 -- something he's using for his model, Maba, and which was previously used in some other models, like ALBERT https://arxiv.org/abs/1909.11942v6 . It's a really nifty idea; you use a similar trick to LoRA as a way of reducing the number of parameters used for your embeddings and your output head. And for small models, those can be a disproportionate number of the total. For example, with my 163 million parameter GPT-2 small-style models, the embeddings are about 39 million of them https://www.gilesthomas.com/2026/07/llm-parameter-counts ; the output head is the same size, so with those two taken together, that's half the model just getting stuff in and out rather than actually doing the thinking. Even if you use weight tying, like the original GPT-2 models did, you'll wind up spending 30% of your "budget" on the single matrix shared between embeddings and the output head. Similarly, while larger models have a smaller percentage spent that way, even medium-sized MoE models can wind up spending a lot of the active parameters on embeddings; I calculated https://www.gilesthomas.com/2026/07/benchmarking-qwen-3-6-35b-moe-rtx-3090 that for Qwen 3.6 35B MoE, with 3B active parameters, there were about 1B of them used across both the input embeddings and the output head. So anything that can reduce that -- so long as it doesn't come at a high cost in terms of the model's capabilities -- is worth considering. You save on your parameter count, which either means smaller models and potentially quicker training, or allows you to "invest" the savings in more thinking parameters -- that is, a wider network with a higher embedding dimensionality, or a deeper one with more layers. Andrew reported that the impact of using this trick was minimal: At rank 128 on a 50k vocabulary, the penalty on validation cross-entropy is negligible typically within +0.02 to +0.04 loss, or <0.5 perplexity delta . Now, that suggests that he's getting a loss of around 2.5 that being the point where a loss increase of 0.04 implies a perplexity change of 0.5 , which is much lower than I typically get with my GPT-2 architecture 3.5 is more like it for me , but still, perhaps I'd get good numbers too? I decided to train a few models to see what happened. The results were interesting I found that using this trick on both the embeddings and the output head caused a somewhat larger increase in loss than Andrew described -- about 0.07 in my first test, 0.10 in the second, and 0.14 in the third. Those are relatively large, though potentially recoverable from further training or larger models. But more surprisingly, I found that using the trick on the input embeddings only seemed to reduce test loss. Before I'd started, I'd expected it to be less harmful on the input side than it was on the output side, but seeing an improvement was certainly unexpected. Getting loss down by reducing your number of parameters is not what normally happens So, it's definitely worth digging in a bit. Let's start with the theory: what are these factorised embeddings, and how do they work? Like LoRA, they rely on low-rank factors, so firstly I'll define those. Low-rank factors Our models are made up of a large set of matrices -- embeddings, output heads, attention weights, FFN linear layers, and so on. Making them smaller obviously reduces the number of parameters for the model, though you would expect it to come at a cost. The idea behind low-rank factors is that you can replace a larger matrix with two smaller ones that will contain most of the important information. Imagine that you've got a matrix of size . Now, from matrix multiplication, we know that if you multiply an matrix by an one, you'll get a result that is . If we do that, then we've factorised the matrix in the same way as we might factorise 12 into 3 and 4 because , and the value is referred to logically enough as the rank of the factorisation. If is sufficiently smaller than and , then these two matrices -- the low-rank factors -- combined will contain fewer parameters than the full matrix. Let's make this specific, using the output head of a GPT-2 style model without weight tying. It takes in the embeddings that result from the Transformer layers after normalisation , and converts them into logits across the vocabulary. For the GPT-2 small size, our embeddings are 768-dimensional, and the vocab size is 50,257 tokens. So when we use a linear layer to do that mapping, it has parameters. If we were to replace it with two matrices -- say, one of and one of , then combined they would take up parameters. That is almost six times smaller So: the idea is that instead of creating our model with an matrix, we create it with a pair, and , and train those instead. If all goes well, that pair will be almost as capable of learning what we want as the full one would have been, so we'll get results that are close enough. The maths and the implementation work out simply, too. For a linear layer with no bias, we can write the matrix multiplication that takes our inputs and a weight matrix , and produces an output , like this: 1 fn-1 Now, if we're using a pair of low-rank factor matrices and instead of a full one, , we can say the "virtual" weights we want to use for the calculation are such that: So that means that our calculation for this neural network is this: Matrix multiplication is associative, which means that you can rewrite that as: ...which hopefully you can see is the same as feeding through a layer using as its weights, and then feeding the result through a second one using . In code, you've taken something like this: out head = nn.Linear d emb, vocab size, bias=False ... logits = out head embeddings ...and replaced it with this: out head A = nn.Linear d emb, r, bias=False out head B = nn.Linear r, vocab size, bias=False ... intermediate = out head A embeddings logits = out head B intermediate It's probably intuitively obvious, though, that this comes at a cost. If you have fewer parameters then you can store less information, so this part of your model is "dumber". If it's not clear, though, imagine that r in the code above was one. You would be taking the for GPT-2 small 768-dimensional embedding, converting it to a single number, and then expanding that single number out to a 50,257-dimensional set of logits. Stuff is going to get lost -- the idea behind low-rank factors is that, so long as you choose an appropriate value for which is often referred to as the width of the low-rank bottleneck , you won't lose the important stuff. This works surprisingly well in many cases -- in particular, LoRA, which allows you to fine-tune models that are too hard to fully train on your hardware, or to fine-tune them faster, uses it to easily train something like low rank factor "diffs" to your weight matrices