# Otōto 2.8.1

> Source: <https://ototo.dev/changelog#v2.8.1>
> Published: 2026-09-27 14:57:40+00:00

- **Fewer wasted turns.** The small model sometimes answered its first turn in prose ("no repository was provided") or called Claude Code's tools by name. Now a prose answer is asked again with a tool call required (on servers that allow it), shell commands that only look at files (`grep` ,`cat` ,`sed -n` ,`ls` ,`find` ) run as Otōto's own tools, and a search that finds nothing says why and which kinds of file the repository has.
- **Safe from a model server's stale cache.** On 27 September our vLLM server's shared prompt cache held wrong data for part of Otōto's instructions, and every question that reused it went wrong until a restart. Each question now keeps its own cache on the server (`cache_salt` ), and`ototo doctor` checks whether a server's cache answers as a fresh request would, and says to restart it if not.
- **Running your own vLLM?**`vllm/README.md` in this package has the flags we run, a chat template for coding clients (Qwen's own, plus what those clients send that Qwen's refuses) with a script that checks it, and what to do when the prompt cache goes bad.
- **Plugins, sealed.** Compiled plugins are cached with a key of your own (`~/.config/ototo/cache.key` , made on first use), so code put in the cache by anything else is compiled over, never run. The first start after upgrading compiles them once, in about a second.
- **Reporting:** if your team collects Otōto's counts, they now carry your Claude account's e-mail, as Claude Code's own do, so the dashboard shows the two side by side per person without any setting.`user.email=` in`otlp_attributes` sends none, and`ototo doctor` says which address goes out.
- **For organisations:**`ototo managed --install` run twice no longer loses the backup of the original settings, and`ototo doctor` no longer gives managed users advice they cannot take.
