cd /news/ai-products/implausible-generation-speed-number-… · home topics ai-products article
[ARTICLE · art-122794] src=kagifeedback.org ↗ pub= topic=ai-products verified=true sentiment=· neutral

Implausible generation speed number in Assistant

A Kagi Assistant user reports that the displayed generation speed of 22 tokens per second contradicts the actual rate of 248 tokens per second calculated from 48,938 tokens generated in 197 seconds, suggesting the UI may be undercounting by an order of magnitude. The user speculates that the metric might exclude reasoning and tool call tokens, but notes this would be misleading given the total request duration.

read1 min views1 publishedSep 8, 2026

Please see this Assistant thread as an example. 48,938 tokens in 197 seconds works out to 248 tokens per second, yet the UI shows 22 tokens per second, so this is confusingly off by an order of magnitude.

(It might be plausible that the 22 tok/s measures only the final output tokens, and disregards reasoning and tool call/response tokens. But it would be a little weird to measure this way over the total request duration, because the LLM is spending most of its time reasoning and doing tool calls.)

I would expect to see a number around 248.

── more in #ai-products 4 stories · sorted by recency
── more on @kagi assistant 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/implausible-generati…] indexed:0 read:1min 2026-09-08 ·