Ollama 0.35 landed September 28 with a local implementation of the Jev API. The new /v1/systemone endpoint mirrors TypeSafe’s hosted only decision API, and two decision models are already in the library with a one-line pull: ollama pull nimble (Bespoke Labs, 9.5GB) and ollama pull tev1 (Together AI, 4.5GB). The release note is unambiguous about the shape: decision models return choices, probabilities, and scores instead of text.
What a decision model is (one paragraph) #
You hand it a short state plus a set of named questions: which label fits this ticket, is this condition true, how urgent on a 1-to-5 rubric. It answers every question in one request, with a probability attached to each allowed answer and nothing to parse out of generated text. The taxonomy comes from TypeSafe’s Jev, whose closed API until tonight defined the category; the house explainers on Jev cover the mechanics (a hundredth of a cent per call, milliseconds, no essay).
The request, verbatim #
The release ships a working example: send state (“Our checkout has returned 500 errors since 9am”) plus a label question of type choice carrying the criteria object, and the response comes back as "choice": "bug" with per-option probabilities (bug 0.9781, billing 0.0125, account 0.0093), a confidence number, and usage of 174 input tokens and 1 output token. One output token is the tell that none of this is generation - the score for each candidate is computed directly. Three question types ship: choice (pick one, probabilities for all), noul (the probability a stated condition is true), and score (a value on an ordered rubric).
Nimble’s pedigree matters more than the runtime #
Nimble is Bespoke Labs’ open recipe for training these models, published two days after Jev’s debut: a LoRA on Qwen3.5-9B trained on 2,676 curated contrastive examples for one epoch, scoring 90.12% agreement with Jev on a 324-example holdout against Jev’s own 93.21%. The data, recipe, and benchmark suite are Apache-2.0 on GitHub, meaning the pull command ships a model any commercial team can run and fine-tune further against their own ticket taxonomy. Tev1 from Together AI fills the small end at 4B (4.5GB, with an 812MB quantized variant on the same page).
Why the runtime matters #
Until now, running a decision model locally meant hosting an inference server yourself and wiring a bespoke scoring path; the API shape was the proprietary part. systemone on your own machine makes the same call the router platforms sell: route a support ticket, gate an agent tool call, pick a model per request - with per-call probabilities that make acceptance thresholds explicit, and the per-request cost being whatever electricity amounts to. For the site’s framing: the local stack now covers judgment calls as a first-class API next to chat completion, which is the last piece of the “small models for small decisions” argument going open runtime.
A demo shows real time: Nimble playing a racing game through systemone calls, deciding steering continuously at 4K resolution - typed decisions with millisecond latency, no essay between states.
Related: the decision-model pieces on this site: Three decision models just ate the router startups covers the hosted side, and Kev: an open decision-model family you can train yourself (coming to this site) covers the train-your-own lane. The models are in the modeldex: Nimble and Tev1 land there this cycle as decision-model type.
Sources: Ollama 0.35.0 release notes - announcement - Nimble repository - Bespoke-Nimble-9B on Hugging Face