04:00
2026-08-25
arxiv.org
natural-language-processing
Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems
A new arXiv study (2608.21384v1) finds that Ukrainian and other Cyrillic-script languages face 68-121% token overhead on modern tokenizers and 220% on the older cl100k, compared to English, across ninβ¦