09:24
2026-09-02
healeycodes.com
large-language-models
What Makes LLM Tokenization Slow?
LLM tokenization, while a small part of overall latency, sits on the hot path and can occur multiple times per request, according to a technical analysis of GPT-2's reference encoder. The analysis sho…