cd /news/artificial-intelligence/ctok-reconstructed-claude-tokenizer · home topics artificial-intelligence article
[ARTICLE · art-98252] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Ctok: Reconstructed Claude Tokenizer

Ctok, an unofficial open-source library, reconstructs Anthropic's Claude tokenizer offline, reporting exact token counts for 1,664,940 v3 and 1,722,961 v4.7 texts with zero under-counts. The library supports Claude model generations from v3 (Claude 3 through Opus 4.6) to v5 (Opus 5 and Sonnet 5), with version strings compared component-wise, and is not affiliated with Anthropic.

read3 min views1 publishedAug 15, 2026
Ctok: Reconstructed Claude Tokenizer
Image: Michielbdejong (auto-discovered)

ctok

reconstructs Claude token counts offline, with no API call, network access, or runtime dependencies. It is unofficial and is not affiliated with Anthropic.

The reconstruction targets counts. Claude does not expose token boundaries, so tokenize()

returns one valid minimum-cost tiling, not a claim about Anthropic's exact segmentation. The research behind the model is described in On the biology of Claude's tokenizer.

from ctok import token_count, tokenize

token_count("hello, world")           # 10, using the v3 family
token_count("hello, world", "4.7")    # 15
token_count("hello, world", "5.0")    # 10

tokens = tokenize("NASA likes tokenizers")
assert len(tokens) == token_count("NASA likes tokenizers")

The command-line interface prints the marked stream and its tiling:

ctok "hello, world"

version

is always a string, e.g. "4.7"

, compared component by component — "4.10"

sorts after "4.9"

, not below "4.2"

. A float

can't make that distinction (Python collapses the literal 4.10

to 4.1

before any code here sees it), so a non-str

version raises TypeError

.

requested version family model generation
"3.0" <= version < "4.7"
v3, the default Claude 3 through Opus 4.6
"4.7" <= version < "5.0"
v4.7 Opus 4.7 through 4.9
version >= "5.0"
v5 Opus 5 and Sonnet 5

v5 is v4.7 with a slightly different fixed overhead.

For one user message, ctok

:

  • normalizes the text, including NFC and family-specific quote folding;
  • rewrites it into a stream with word, case, and byte markers;
  • finds a minimum-cost tiling over the measured vocabulary and UTF-8 byte fallback;
  • adds the measured message frame.

token_count(text)

is len(tokenize(text))

. The output notation makes internal structure visible:

notation meaning
⟨bow⟩ , ⟨eow⟩
word boundaries
⟨shift⟩ , ⟨caps⟩
case rewrites
⟨0xNN⟩
a byte-fallback token
⟨pad⟩
part of the single-message frame

These results compare ctok

with recorded count_tokens

responses:

corpus role v3 exact v4.7 exact
Goldfish, 350 languages and 350,000 rows mining 350,000 350,000
MultiPL-E, 22 programming languages held out 22 22
Rosetta Code, 1,741 documents mining 1,741 1,741
Rosetta Code, separate 250 documents mining 250 250
UDHR, 501 languages mining (in-sample since 2026-08-12) 501 501

v5 has the same content result as v4.7 because it uses the same vocabulary.

The stored measurement sets contain no under-counts: 0 of 1,664,940 v3 texts and 0 of 1,722,961 v4.7 texts. This is an empirical result, not a guarantee for arbitrary input. Goldfish, Rosetta and (since its final six pieces were selected by bisecting it) UDHR may select candidates; MultiPL-E never does, and is the one remaining held-out corpus in this table.

Run the public gates with:

uv run pytest
uv run python tests/gates.py --markdown

The two vocabulary files contain 48,645 v3 pieces and 15,240 v4.7 pieces. Every entry has a fixed membership witness or is one of the structural marker atoms checked by the test suite.

from ctok import pieces, witness

len(pieces("4.7"))                       # 15240
witness("⟨bow⟩the⟨eow⟩", "4.7")

A witness says that one marked piece costs one token in a calibrated probe. It does not prove the encoder rewrite or resolve ties between equal-cost tilings. tests/test_witness.py

checks every published witness and requires complete witnessed-or-special coverage.

MIT

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ctok 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ctok-reconstructed-c…] indexed:0 read:3min 2026-08-15 ·