cd /news/natural-language-processing/the-price-of-token-boundaries-compre… · home › topics › natural-language-processing › article
[ARTICLE · art-142211] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

The Price of Token Boundaries: Compression Certificates and Prediction

A new arXiv paper (2609.35869v1) measures the compression cost of pre-tokenisation boundary rules, finding that on English Wikipedia such boundaries increase the optimal token count by 28.3–36.8%. The authors report that byte pair encoding lies 2.1% above the constrained lower bound but 10.9% above the unrestricted bound, and that at 85M non-embedding parameters with matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte in all 12 languages studied. The paper also introduces boundary licences, showing that licensing 10% of the vocabulary budget recovers 85.2% of the token-count reduction on English and 100.0% on Chinese.

by read1 min views2 publishedSep 30, 2026

arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8%. Byte pair encoding lies 2.1% above the constrained lower bound, but 10.9% above the unrestricted bound. Compression and prediction favour different dictionaries: at 85M non-embedding parameters and matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte under a common unrestricted decoder in all 12 languages in the paired study and 11 of 12 under independent tuning and evaluation. To study intermediate boundary policies, we introduce boundary licences, which limit the vocabulary entries permitted to cross cuts and admit the same form of certificate. On separate English and Chinese fitting corpora, licensing 10% of the vocabulary budget recovers 85.2% and 100.0% of the achieved token-count reduction from removing all cuts. These results quantify the compression cost of boundaries while separating it from the prediction quality of the resulting token units.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-price-of-token-b…] indexed:0 read:1min 2026-09-30 · —