Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter A new arXiv paper (2608.11361v1) shows that the cost-optimal tokenizer vocabulary size for large language models varies by deployment regime, shifting 16x with serving batch size from 32k at batch size 1 to 524k at batch size 64+. The authors, who ran experiments on A10G and A100 GPUs, found that quality is nearly invariant (less than 2% BPB spread) across the optimal range, making vocabulary a pure systems optimization. They recommend on-device deployments use about 32k vocabulary, while datacenter serving should use 131k-262k. arXiv:2608.11361v1 Announce Type: new Abstract: Tokenizer vocabulary size is a foundational design choice in large language model LLM infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis. We show that the cost-optimal vocabulary is not a constant but a function of the serving regime. We formalize total deployment cost as $C {lifecycle} V = C {train} V + \lambda \cdot C {infer} V, B $, where $\lambda$ is inference volume and $B$ is the serving batch size. Through controlled experiments on two GPU families spanning the memory-bound to compute-bound regimes A10G, ridge $\approx$ 117 FLOP/byte; A100, ridge $\approx$ 183 FLOP/byte , we demonstrate: 1 the inference-optimal vocabulary shifts 16x with serving batch, from 32k at $B=1$ to 524k at $B=64+$, driven by amortization of the $V \times d$ unembedding matrix read; 2 at 1.3-2.3B model scale, quality bits per byte, BPB is optimized at $V=65$k, confirming scale-dependent vocabulary preference; 3 the lifecycle-optimal vocabulary diverges from training-optimal by up to 16x for production deployments. Quality is approximately invariant across the optimal range $<$2% BPB spread , making vocabulary a pure systems optimization with no quality penalty in the measured range. Our results provide actionable capacity planning guidance: on-device deployments $B=1$ should use $V \approx 32$k; datacenter serving $B \geq 64$, $\lambda \geq 10$ should use $V \approx 131$-262k.