cd /news/large-language-models/vocab-break · home topics large-language-models article
[ARTICLE · art-109578] src=ianbarber.blog ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Vocab Break

Tokenizer enthusiast Sander Land reproduced a tokenizer similar to Claude's current one and found it has only about 15,000 entries, far fewer than Qwen 3.8's 250,000 tokens. Anthropic's tokenizer size has decreased from about 50,000 entries in Claude 3 to about 16,000 today, possibly to avoid a gradient bottleneck in the final softmax layer, as described in the paper 'Lost in Backpropagation: The LM Head is a Gradient Bottleneck.' This smaller vocab reduces memory and avoids glitch tokens, though it increases token count per input by 1.2-2x.

read3 min views3 publishedAug 25, 2026
Vocab Break
Image: Ianbarber (auto-discovered)

Tokenizer enthusiast Sander Land recently reproduced something very like Claude’s current tokenizer, and it appears to only have about 15,000 entries. That is surprising! Qwen 3.8, a very strong release, has about 250k tokens in its vocab. In general the trend had seemed to be more is better in this space.

One theory is that Anthropic have been working around a bottleneck caused by the final softmax layer. There is a recent(ish) paper about this: “Lost in Backpropagation: The LM Head is a Gradient Bottleneck“, but, if this is the reason, then the folks at Throppy have known this for way longer.

The basic idea is that you have to project at the end of the forward pass from the model’s latent space, dimension D, to a much bigger vocabulary space, dimension V, to select a token. Sticking with Qwen, their 2.4T parameter flagship model has a hidden size, D, of 8,192, and a V of that 250k vocab size.

When training, you compare the distribution that model gets to the actual right token. If the model was correct and confident is small, if it was confidently wrong the loss is large. That loss is then propagated through all 250k entries, and from there down to the 8k entries of the hidden dimension. This compression bottlenecks how much information can be fed back into the network. Specifically, the authors show the change in logits has rank at most 2D. So if V is a lot larger than D, we are losing information:

We show both empirically and theoretically that the softmax bottleneck induces lossy compression during backpropagation

The fact this happens is not totally obvious. The correct distribution is just one entry wide (the actual next token), and the hidden dimension can represent that. Over a wide batch, though, you get all kinds of different next tokens. The signals are sparse, but not low rank: if you go over enough examples nearly every token is, at some point, “the next token”.

This means the learning signal coming in is as wide as the vocab, and so the model is sampling a random D-sized subset of it. That isn’t a problem per-se: you can learn to map between them, but there isn’t anything in the process that particularly encourages it to learn that mapping.

Whether this is the reason for the small vocab or not, there is a question of how they can get away with it! Every other model has been increasing, but as far as Land can estimate the folks at Anthropic have been cutting: from ~50k vocab entries in Claude 3 to ~16k today.

So, whatever the gradient bottleneck costs, Anthropic (mostly) aren’t paying it!

This also has a number of other benefits. You don’t need to do funky chunked CE kernels since you don’t have to project to a big, memory eating, space, and you don’t get any solidgoldmagikarp 1 style

glitch tokens, because every token gets trained. They aren’t ignoring the rare tokens and other languages, they’re just using subword tokens and, in the worst cast, fallbacks to UTF-8. That means more tokens per piece of input text, and more attention cost. That said, it only seems to be 1.2-2x more tokens in Land’s testing. It’s not free: decreasing the tokenizer really is costing more execution and more money, but the tradeoff is presumably more than worth it!

  • After which I presume Land username’d himself. Sorry regular magikarp. ↩︎
── more in #large-language-models 4 stories · sorted by recency
── more on @sander land 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/vocab-break] indexed:0 read:3min 2026-08-25 ·