cd/entity/tiktoken· home› entities› tiktoken
grep -l @tiktoken /news/*.json | wc -l → 46

tiktoken

mentions 46 type Organization page 1/3 feed RSS

// recent coverage 46 mentions

00:34
2026-10-09
simonwillison.net
ai-tools

ttok 1.0

Simon Willison released ttok 1.0, changing the token-counting tool's default tokenizer from GPT-4 to the GPT-5/GPT-6 family after running `uv tool upgrade ttok` and finding the old default still in pl…

23:34
2026-10-08
simonwillison.net
ai-tools

ttok 0.4

Simon Willison released ttok 0.4, an update to his CLI token-counting tool built on OpenAI's open source tiktoken library. The release fixes a Click warning, updates CI, and adds a --list-models comma…

22:22
2026-10-06
actual.inc
ai-infrastructure

toks: the best tokenizer on Earth

Actual Computer released toks 0.3.0, a tokenizer it says runs 13x to 151x faster than Hugging Face tokenizers and beats tiktoken in all 195 cells it can run, while returning identical token ids to Hug…

00:00
2026-09-21
huggingface.co
natural-language-processing

tokenizers v1: encode, decode and scaling, measured

Hugging Face's upcoming tokenizers v1 release candidate produces the same token IDs as v0.23 while running often tens of times faster, the company reported, citing benchmarks run from its tokbench rep…

19:14
2026-09-20
dev.to
natural-language-processing

A model doesn't read text: what a tokenizer decides for you

A developer's technical writeup explains that a language model's tokenizer is a frozen part of the trained artifact rather than preprocessing, and walks through tiktoken's educational byte-pair-encodi…

12:31
2026-09-16
dev.to
large-language-models

Stop Sending Raw HTML to LLMs

A developer measured token usage across 10 web pages and found that raw HTML consumed 3 to 24 times more tokens than the same pages converted to Markdown, with one product page arriving as 267,361 tok…

12:51
2026-09-08
blog.devgenius.io
artificial-intelligence

Feeding Your Local Data to LLMs (II)

A developer tutorial demonstrates extending local LLM data feeding by storing embeddings in a Neo4j graph database and combining similarity search with full-text search, using a local LLM and Python p…

06:01
2026-09-03
dev.to
large-language-models

Your Gemini 3.8 Flash Token Counter Is Wrong

A developer's guide explains that local token counting for Gemini 3.8 Flash using tiktoken or Hugging Face tokenizers is inaccurate, leading to budget overruns and API errors. The recommended solution…

22:32
2026-08-30
github.com
developer-tools

Build a Tokenizer from Scratch

Crackr released a free, open-source guide on building a byte-level BPE tokenizer from scratch, culminating in a multilingual encoding compatible with OpenAI's tiktoken and a web playground. The guide …

01:48
2026-08-27
github.com
ai-tools

Show HN: Tokwhois – 14 probes to name the tokenizer family

A new open-source tool, Tokwhois, uses 14 probes to identify the tokenizer family behind a large language model (LLM) API, even when the model's weights, logits, and architecture are hidden. The tool,…

05:58
2026-08-26
dev.to
developer-tools

How MCP Wastes 4-32x More Tokens Than CLI (and How to Fix It)

A developer measured the token overhead of the Model Context Protocol (MCP) and found that loading 255 tools from 50 MCP servers consumes 71,929 tokens per session, compared to just 123 tokens for the…

00:00
2026-08-24
digitalapplied.com
large-language-models

Tokenizer Variance: Why Identical Prices Cost More

LLM tokenizer variance means identical dollar-per-million token prices can produce materially different bills, because each vendor's tokenizer segments the same input into different token counts. Anth…

18:35
2026-08-18
liquid.ai
artificial-intelligence

Designing Loops for Production-Grade Work

Liquid AI, an AI company, released an open-source byte-pair encoding (BPE) tokenizer trainer called toktoktok on GitHub, built autonomously by two coding agents using Claude Opus 4.5 and Codex with GP…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics