Show HN: Tokwhois – 14 probes to name the tokenizer family A new open-source tool, Tokwhois, uses 14 probes to identify the tokenizer family behind a large language model (LLM) API, even when the model's weights, logits, and architecture are hidden. The tool, released on GitHub by developer fasuizu-br, compares the token counts from these probes against a catalog of 16 public tokenizer families, reporting a family with a confidence heuristic and failing closed if the result is ambiguous. This matters because tokenizers are frozen at training and often reused, making them a fingerprint for stealth LLM deployments. Whois for stealth LLMs. Labs can hide the weights, the logits, and the architecture. They cannot hide the tokenizer they bill you with. git clone https://github.com/fasuizu-br/tokwhois cd tokwhois pip install -e . python3 -m tokwhois demo Zero-install, from the git URL the package is not on PyPI : uvx --from git+https://github.com/fasuizu-br/tokwhois tokwhois demo uvx --from git+https://github.com/fasuizu-br/tokwhois tokwhois https://api.example.com/v1 --model gpt-4o-mini A 14-integer fertility vector . Each probe is a fixed, versioned string. The server returns usage.prompt tokens for a 1-token completion. That integer is compared to a catalog of public tokenizers tiktoken encodings + Hugging Face tokenizer.json , licenses permissive, pinned by commit/version . Catalog v1 is 16 families ; Qwen 2/2.5 is not Qwen3. bash $ python3 -m tokwhois demo family glm4-class confidence heuristic 1.00 L1 distance: 0 runner-up cl100k base margin 35 tokens L1 probe counted glm4 cl100k bas internlm2 o200k base cjk30 22 22 30 23 28 space40 1 1 1 1 1 digit64 43 43 22 32 22 ascii100 26 26 26 27 26 emoji8 21 21 21 29 13 hello leadsp 1 1 1 1 1 nl16 1 1 1 1 1 tab16 1 1 1 1 1 cjk en 8 8 12 8 8 im start 6 6 6 1 6 gmask 1 1 3 3 3 eot 1 1 1 7 1 bot llama 7 7 7 7 7 byte rare 22 22 22 24 23 ────────────────────────────────────────────────────────────────── offset empty 7 subtracted from every prompt count n=1 probe / string K=1 catalog=v1 2026-08-24 discriminating probes vs runner-up: cjk30, digit64, cjk en, gmask It reports a tokenizer family , not a checkpoint, not a lab, not a parameter count. If the top two families land inside the margin, it prints ambiguous and stops. It does not guess. confidence in the output is a heuristic score of L1 distance and margin, not a probability. The package is not on PyPI. Install from the repository: git clone https://github.com/fasuizu-br/tokwhois cd tokwhois pip install -e . or, zero install: uvx --from git+https://github.com/fasuizu-br/tokwhois tokwhois demo Python 3.10+. The demo and selftest run offline stdlib + the embedded catalog . Live mode requires httpx . tiktoken and tokenizers are optional and only used to rebuild the catalog or encode local files. python3 -m tokwhois demo This encodes the v1 probes against the embedded catalog, prints the vectors, and asserts that each family matches itself at export OPENAI API KEY=... python3 -m tokwhois "$OPENAI BASE URL" --model "$MODEL" Fourteen max tokens=1 calls. Fail-closed if usage.prompt tokens is missing. Chat-template framing overhead is subtracted via an empty probe so the live vector can be compared to the local catalog. That comparison is a working hypothesis BPE is not addition; see METHOD.md /fasuizu-br/tokwhois/blob/main/docs/METHOD.md . If the empty probe fails, or a calibrated count is less than 1, the client aborts. There is no silent fallback. python3 -m tokwhois --local path/to/tokenizer.json python3 -m tokwhois demo --json python3 -m tokwhois "$OPENAI BASE URL" --model "$MODEL" --json Tokenizers are frozen at training. Stealth deployments almost always reuse a public tokenizer because training a new one is a research project, not a wrap. Billing requires usage . The combination is a fingerprint the server computes for you. This is not a watermark, not a logit attack, and not stylometry. The empty-prompt offset is a first-order correction, not an identity. BPE is not addition; a chat template is not concatenation. See METHOD.md /fasuizu-br/tokwhois/blob/main/docs/METHOD.md . Not a statement about model quality. Not an identification of a lab, a checkpoint, or a size. Not a benchmark. There is no leaderboard in this repository. Not a request that anyone violate a provider's terms. You run it against endpoints you are authorized to call. Let working hypothesis for encode s no template . The catalog stores ambiguous if the runner-up is within margin default 2 . Every number in the report is an integer count with METHOD.md /fasuizu-br/tokwhois/blob/main/docs/METHOD.md for the full formulation. Copyright 2026 Fabio Suizu / Brainiall. Licensed under the Apache License, Version 2.0. See LICENSE /fasuizu-br/tokwhois/blob/main/LICENSE .