cd /news/ai-agents/ototo-2-8-1 · home › topics › ai-agents › article
[ARTICLE · art-145692] src=ototo.dev ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Otōto 2.8.1

Otōto released version 2.8.1, adding per-question prompt cache isolation via cache_salt after a 27 September incident in which the project's vLLM server shared prompt cache held wrong data for part of Otōto's instructions and every question reusing it failed until a restart. The release also requires a tool call when the small model answers its first turn in prose, runs file-inspection shell commands (grep, cat, sed -n, ls, find) as Otōto's own tools, caches compiled plugins under a user-specific key at ~/.config/ototo/cache.key, and attaches the Claude account e-mail to collected counts unless user.email= is set in otlp_attributes.

read2 min views7 publishedSep 27, 2026
  • Fewer wasted turns. The small model sometimes answered its first turn in prose ("no repository was provided") or called Claude Code's tools by name. Now a prose answer is asked again with a tool call required (on servers that allow it), shell commands that only look at files (grep ,cat ,sed -n ,ls ,find ) run as Otōto's own tools, and a search that finds nothing says why and which kinds of file the repository has.
  • Safe from a model server's stale cache. On 27 September our vLLM server's shared prompt cache held wrong data for part of Otōto's instructions, and every question that reused it went wrong until a restart. Each question now keeps its own cache on the server (cache_salt ), andototo doctor checks whether a server's cache answers as a fresh request would, and says to restart it if not.
  • Running your own vLLM?vllm/README.md in this package has the flags we run, a chat template for coding clients (Qwen's own, plus what those clients send that Qwen's refuses) with a script that checks it, and what to do when the prompt cache goes bad.
  • Plugins, sealed. Compiled plugins are cached with a key of your own (~/.config/ototo/cache.key , made on first use), so code put in the cache by anything else is compiled over, never run. The first start after upgrading compiles them once, in about a second.
  • Reporting: if your team collects Otōto's counts, they now carry your Claude account's e-mail, as Claude Code's own do, so the dashboard shows the two side by side per person without any setting.user.email= inotlp_attributes sends none, andototo doctor says which address goes out.
  • For organisations:ototo managed --install run twice no longer loses the backup of the original settings, andototo doctor no longer gives managed users advice they cannot take.
── more in #ai-agents 4 stories · sorted by recency
── more on @otōto 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ototo-2-8-1] indexed:0 read:2min 2026-09-27 · —