# Ollama 0.35 ships the Jev API locally: pull Nimble or Tev1

> Source: <https://tokenstead.ai/guides/ollama-systemone-decision-models>
> Published: 2026-09-30 13:19:58+00:00

Ollama 0.35 landed September 28 with a local implementation of the Jev API. The new `/v1/systemone` endpoint mirrors TypeSafe’s hosted only decision API, and two decision models are already in the library with a one-line pull: `ollama pull nimble` (Bespoke Labs, 9.5GB) and `ollama pull tev1` (Together AI, 4.5GB). The release note is unambiguous about the shape: decision models return choices, probabilities, and scores instead of text.

## What a decision model is (one paragraph)

You hand it a short state plus a set of named questions: which label fits this ticket, is this condition true, how urgent on a 1-to-5 rubric. It answers every question in one request, with a probability attached to each allowed answer and nothing to parse out of generated text. The taxonomy comes from TypeSafe’s Jev, whose closed API until tonight defined the category; the house explainers on Jev cover the mechanics (a hundredth of a cent per call, milliseconds, no essay).

## The request, verbatim

The release ships a working example: send `state` (“Our checkout has returned 500 errors since 9am”) plus a `label` question of type `choice` carrying the criteria object, and the response comes back as `"choice": "bug"` with per-option probabilities (`bug 0.9781`, `billing 0.0125`, `account 0.0093`), a confidence number, and usage of 174 input tokens and 1 output token. One output token is the tell that none of this is generation - the score for each candidate is computed directly. Three question types ship: `choice` (pick one, probabilities for all), `noul` (the probability a stated condition is true), and `score` (a value on an ordered rubric).

## Nimble’s pedigree matters more than the runtime

Nimble is Bespoke Labs’ open recipe for training these models, published two days after Jev’s debut: a LoRA on Qwen3.5-9B trained on 2,676 curated contrastive examples for one epoch, scoring 90.12% agreement with Jev on a 324-example holdout against Jev’s own 93.21%. The data, recipe, and benchmark suite are Apache-2.0 on GitHub, meaning the pull command ships a model any commercial team can run and fine-tune further against their own ticket taxonomy. Tev1 from Together AI fills the small end at 4B (4.5GB, with an 812MB quantized variant on the same page).

## Why the runtime matters

Until now, running a decision model locally meant hosting an inference server yourself and wiring a bespoke scoring path; the API shape was the proprietary part. `systemone` on your own machine makes the same call the router platforms sell: route a support ticket, gate an agent tool call, pick a model per request - with per-call probabilities that make acceptance thresholds explicit, and the per-request cost being whatever electricity amounts to. For the site’s framing: the local stack now covers judgment calls as a first-class API next to chat completion, which is the last piece of the “small models for small decisions” argument going open runtime.

A demo shows real time: Nimble playing a racing game through `systemone` calls, deciding steering continuously at 4K resolution - typed decisions with millisecond latency, no essay between states.

Related: the decision-model pieces on this site: [Three decision models just ate the router startups](https://tokenstead.ai/guides/decision-models-ate-routers) covers the hosted side, and Kev: an open decision-model family you can train yourself (coming to this site) covers the train-your-own lane. The models are in the modeldex: Nimble and Tev1 land there this cycle as decision-model type.

Sources: [Ollama 0.35.0 release notes](https://github.com/ollama/ollama/releases/tag/v0.35.0) - [announcement](https://x.com/ollama/status/2105152056382345544) - [Nimble repository](https://github.com/bespokelabsai/nimble) - [Bespoke-Nimble-9B on Hugging Face](https://huggingface.co/bespokelabs/Bespoke-Nimble-9B)
