cd /news/developer-tools/show-hn-i-built-a-token-per-second-c… · home topics developer-tools article
[ARTICLE · art-69840] src=github.com ↗ pub= topic=developer-tools verified=true sentiment=· neutral

Show HN: I built a token-per-second counter in Go

A developer built tokps, a Go-based CLI tool that measures token-generation throughput (TPS) for any OpenAI-compatible /chat/completions endpoint. The tool sends streaming test prompts, runs a warmup followed by multiple timed requests, and reports median (p50) TPS along with min–max range, supporting providers such as OpenAI, Z.ai/GLM, and local models. It is available via `go install` or from source on GitHub.

read5 min views1 publishedJul 23, 2026
Show HN: I built a token-per-second counter in Go
Image: source

A tiny CLI that measures the token-generation throughput (TPS) of any OpenAI-compatible /chat/completions

endpoint — OpenAI, Z.ai / GLM, local models, custom gateways, anything that speaks the same wire format.

It sends a test prompt with streaming on, watches the tokens stream back, and prints how fast the model generated them. By default it runs one discarded warmup followed by five timed requests and reports the median (p50) plus the observed min–max range, so a single cold start or network hiccup doesn't skew the result.

tokps — glm-5.2 @ api.z.ai  (5 runs, 1 warmup)

  prompt tokens     39
  output tokens     200   (exact, median)

  TTFT     p50 2.61s   range 2.41s–2.95s
  TPS      p50 73.1   range 69.8–75.4   (generation, N-1)
  e2e      p50 36.8   range 34.1–38.0   (incl. TTFT)

For a single cheap request, pass --runs 1 --warmup 0

— the output falls back to a detailed single-shot block (per-run TTFT, generation, and total wall).

go install github.com/canergulay/tokps@latest

This drops a tokps

binary in $(go env GOPATH)/bin

. Make sure that's on your PATH

.

Or build from source:

git clone https://github.com/canergulay/tokps
cd tokps
go build -o tokps .

Set your key once (API_KEY

is provider-agnostic — it applies to whatever --url

you point at), then run:

export API_KEY=sk-...

tokps --url https://api.openai.com/v1 --model gpt-4o-mini

tokps --url https://api.z.ai/api/paas/v4 --model glm-5.2

tokps --url http://localhost:8080/v1 --model my-local-model

The base URL has /chat/completions

appended automatically, so …/v1

and …/paas/v4

both work. If you pass a full …/chat/completions

URL it's used as-is.

Flag Default Description
--url
(required)
Base URL of the endpoint.
--model
(required)
Model name.
--api-key
API_KEY env, then OPENAI_API_KEY
Bearer token. Flag wins over env.
--prompt
built-in Override the test prompt.
--max-tokens
512
Upper bound on output length.
--runs
5
Number of timed runs; reports p50 + min–max across them.
--warmup
1
Discarded warmup runs before measuring (absorbs cold start).
--detail
false
Also show inter-token latency (ITL) p50/p95.
--json
false
Emit machine-readable JSON instead of the text summary.
--concurrency
1
Parallel streams per run; >1 reports aggregate tok/s under load.
--sweep
Comma-separated concurrency levels to sweep, e.g. 1,2,4,8 .
--timeout
60s
Per-request timeout.

Each invocation sends

--warmup

+--runs

requests (6 by default), so it makes that many billable calls against a metered endpoint. Use--runs 1 --warmup 0

for a single request.

Run tokps

with no flags to see the full list.

Each request is sent with stream: true

and stream_options.include_usage: true

. As chunks arrive, tokps records:

time to first token (TTFT)— request send → first generated token.** generation time**— first → last generated token.** total wall**— request send → last token.

It reports two throughput numbers:

TPS (headline)=(output tokens − 1) ÷ generation time

. The first token is produced during TTFT, so the first-to-last window spansN − 1token intervals — dividing byN − 1

is the standard serving-benchmark definition (vLLM, NVIDIA genai-perf, Anyscale llmperf) and the inverse of mean inter-token latency. This is the pure decode rate, excluding initial latency.end-to-end=output tokens ÷ total wall

, which folds in TTFT.

Runs and percentiles. A model's speed varies run to run (cold replicas, queueing, KV-cache state, network jitter). So tokps runs --warmup

discarded requests, then --runs

timed ones, and reports the median (p50) and the observed min–max range for each metric. Min/max are reported rather than p90/p99 because they're unambiguous for both latency (higher = worse) and throughput (higher = better), and don't oversell percentile resolution at small run counts.

Token counts come from the stream's usage

field when the server sends it (labeled exact

). If a server omits it, tokps estimates from the streamed text length at ~4 chars/token (labeled estimated

) — model-agnostic and independent of how the server chose to chunk the stream.

--detail

adds an inter-token latency (ITL) line — the p50 and p95 of the gaps between successive streamed tokens, pooled across runs. p95 surfaces stalls/jitter that a single averaged rate hides.

--json

emits the full result as machine-readable JSON instead of the text table — the p50/min/max for every metric, ITL, and a runs_detail

array with each run's raw numbers — for CI gates, storing, and diffing over time.

By default tokps measures a single stream — the "how fast is one response" question. --concurrency N

instead fires N streams in parallel per run and adds an aggregate tok/s line (total output tokens across all streams ÷ wall time) alongside the per-stream TTFT/TPS distribution — i.e. throughput under load.

--sweep 1,2,4,8

runs the benchmark at each level in turn and prints the throughput-vs-concurrency curve, so you can see where an endpoint saturates:

tokps — glm-5.2 @ api.z.ai  (sweep, 3 runs, 1 warmup)

  concurrency   aggregate tok/s   TTFT p50   TPS p50/stream
  1             73.1              0.42s      73.1
  2             140.0             0.45s      70.0
  4             250.0             0.61s      62.5
  8             300.0             1.10s      37.5

Reasoning models such as GLM-5.2 and DeepSeek-R1 stream their thinking tokens in delta.reasoning_content

rather than delta.content

, often by default. tokps counts those toward timing and throughput, so TTFT and TPS reflect the full generation including the reasoning phase.

If a server ignores stream: true

and returns a single JSON object, tokps still reports tokens (from usage

) and total time; TTFT and the generation-only rate are shown as n/a

since per-token timing isn't available.

Pure Go standard library, no third-party dependencies.

go test ./...     # unit + integration tests (httptest, no network)
go vet ./...
go build ./...

The code is split into small, independently testable packages:

internal/sse

— turns an SSE byte stream into data payloads.internal/bench

— builds/sends the request, drives the parser, collects metrics.internal/report

— formats the summary block.

MIT © Caner Gulay

── more in #developer-tools 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-i-built-a-to…] indexed:0 read:5min 2026-07-23 ·