cd /news/natural-language-processing/tokka-bench-evaluating-tokenizers-ac… · home › topics › natural-language-processing › article
[ARTICLE · art-147315] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages

Researchers introduced Tokka-Bench, an open-source framework that evaluates tokenizers on five metrics — bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition — across 100 natural languages (30+ scripts) and 20 programming languages, per arXiv paper 2610.08794v1. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, the authors found vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available.

by read1 min views2 publishedOct 8, 2026

arXiv:2610.08794v1 Announce Type: new Abstract: Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @tokka-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tokka-bench-evaluati…] indexed:0 read:1min 2026-10-08 · —