ttok 1.0
Simon Willison released ttok 1.0, changing the token-counting tool's default tokenizer from GPT-4 to the GPT-5/GPT-6 family after running `uv tool upgrade ttok` and finding the old default still in pl…
Simon Willison released ttok 1.0, changing the token-counting tool's default tokenizer from GPT-4 to the GPT-5/GPT-6 family after running `uv tool upgrade ttok` and finding the old default still in pl…
Simon Willison released ttok 0.4, an update to his CLI token-counting tool built on OpenAI's open source tiktoken library. The release fixes a Click warning, updates CI, and adds a --list-models comma…
A developer published a Python tutorial for building a long-document question-answering pipeline that uses a single 1M-token context window instead of a vector database, embeddings, or chunking librar…
Actual Computer released toks 0.3.0, a tokenizer it says runs 13x to 151x faster than Hugging Face tokenizers and beats tiktoken in all 195 cells it can run, while returning identical token ids to Hug…
SerpApi published a comparison of markdown search output across SerpApi, Tavily, Exa, and Firecrawl, running three queries ("coffee", "how to make delicious coffee", and "grok 4.7") through all four p…
Hugging Face's upcoming tokenizers v1 release candidate produces the same token IDs as v0.23 while running often tens of times faster, the company reported, citing benchmarks run from its tokbench rep…
A developer's technical writeup explains that a language model's tokenizer is a frozen part of the trained artifact rather than preprocessing, and walks through tiktoken's educational byte-pair-encodi…
A developer has published a technical explainer on tokenization, detailing how large language models convert text into integer token sequences rather than processing words directly. The writeup covers…
The LectuLibre team built a token-based chunking pipeline for translating full-length books with large language models, using Python, FastAPI, and Anthropic's Claude alongside DeepSeek for simpler pas…
The iLostCount project published a token-conversion cheat sheet stating that 1,000 words of ordinary English is roughly 1,300 tokens, while 1,000 tokens is roughly 750 words or about 4,000 characters.…
A developer measured token usage across 10 web pages and found that raw HTML consumed 3 to 24 times more tokens than the same pages converted to Markdown, with one product page arriving as 267,361 tok…
LectuLibre, an AI-powered book translation platform, developed a chunking and orchestration system to translate 300-page books using Claude API while managing token limits. The system uses paragraph- …
A developer tutorial demonstrates extending local LLM data feeding by storing embeddings in a Neo4j graph database and combining similarity search with full-text search, using a local LLM and Python p…
A developer's guide explains that local token counting for Gemini 3.8 Flash using tiktoken or Hugging Face tokenizers is inaccurate, leading to budget overruns and API errors. The recommended solution…
Crackr released a free, open-source guide on building a byte-level BPE tokenizer from scratch, culminating in a multilingual encoding compatible with OpenAI's tiktoken and a web playground. The guide …
A new open-source tool, Tokwhois, uses 14 probes to identify the tokenizer family behind a large language model (LLM) API, even when the model's weights, logits, and architecture are hidden. The tool,…
A developer measured the token overhead of the Model Context Protocol (MCP) and found that loading 255 tools from 50 MCP servers consumes 71,929 tokens per session, compared to just 123 tokens for the…
LLM tokenizer variance means identical dollar-per-million token prices can produce materially different bills, because each vendor's tokenizer segments the same input into different token counts. Anth…
The skillcheck static analyzer for SKILL.md files received a hardening update, fixing the description scorer, improving error handling for corrupt files, and correcting token estimate documentation. M…
Liquid AI, an AI company, released an open-source byte-pair encoding (BPE) tokenizer trainer called toktoktok on GitHub, built autonomously by two coding agents using Claude Opus 4.5 and Codex with GP…