Evaluating a decision model: Nimble 9B and its base model Qwen3.5–9B A local evaluation of Bespoke Labs' Nimble 9B decision model against its base model Qwen3.5–9B on a nine-team consumer complaint routing task found the published benchmarks did not come from routing, the task the author would use such a model for. Nimble 9B, released under Apache 2.0 and run through Ollama 0.35's /v1/systemone endpoint on a MacBook Pro M1 Pro, was tested on 900 CFPB Consumer Complaint Database complaints (100 per team, from 2018 onward, 200–2,000 characters) against Qwen3.5–9B with reasoning off and on. Bespoke Labs' Nimble model page reports 75.7% mean accuracy across 13 public decision datasets and a lift from 66% to 90% on its own test set, while Ollama cites 91 ms per decision on a MacBook Pro M5 Max and TypeSafe claims Jev is "193.6x faster, 444.6x cheaper" than hosted frontier models. In September 2026 TypeSafe AI https://typesafe.ai/ released Jev https://typesafe.ai/blog/introducing-system-one-models-and-jev and called it a decision model , or “System One” model. A decision model does not write text. You give it some state, as text or JSON, and a set of typed questions. It returns typed answers: one option out of several, a yes or no, or a score. Every answer comes with probabilities. There is no explanation, and the output cannot be malformed. In this article I compare a decision model called Nimble 9B with Qwen3.5–9B, the model it was trained from. Both ran locally on a MacBook Pro with an M1 Pro chip. The task: read a consumer complaint and pick one of nine teams that should handle it. Two weeks after Jev, Ollama 0.35 added a local endpoint https://ollama.com/blog/ollama-now-supports-jev-style-decision-models for this kind of model, /v1/systemone https://docs.ollama.com/api/systemone , together with open models that use the same API. One of them is Nimble 9B https://ollama.com/library/nimble from Bespoke Labs https://bespokelabs.ai , released under Apache 2.0. That made it possible to evaluate the idea locally, on a task of my own choosing. The published numbers are high. The Nimble model page https://ollama.com/library/nimble gives 75.7% mean accuracy across 13 public decision datasets, and says the training lifts the base model from 66% to 90% correct on Bespoke Labs’ own test set. The Ollama announcement gives 91 ms per decision on a MacBook Pro M5 Max. TypeSafe says Jev is “193.6x faster, 444.6x cheaper” than hosted frontier models in its own workflow evaluations, and calls that “on the higher end of real world gains”. None of these numbers comes from a routing task, and routing is the first thing I would use such a model for. So I tested routing, where the model reads a description of a problem and decides which team should handle it. I had three questions. Nimble’s model page says it is fine-tuned from Qwen3.5–9B https://ollama.com/library/qwen3.5 . That makes a fair comparison possible, because both models have the same architecture and the same starting weights. The main difference is the decision training. So I compare Nimble 9B with Qwen3.5–9B. Qwen runs once with reasoning off, where it answers directly, and once with reasoning on, where it first writes out its thinking and then answers. The complaints come from the Consumer Complaint Database https://www.consumerfinance.gov/data-research/consumer-complaints/ of the US Consumer Financial Protection Bureau CFPB . Consumers file complaints about financial products there. If the consumer agrees, the text is published with personal details replaced by XXXX. The data is in the public domain. I read it through a copy on Hugging Face, BEE-spoke-data/consumer-finance-complaints https://huggingface.co/datasets/BEE-spoke-data/consumer-finance-complaints , which has 1.69 million complaints with text. Every complaint has a product, which the consumer picks from a fixed list when filing. I use that as the label, because the product decides which team at a bank would receive the complaint. So this is a routing decision that a person made on real text. That is why I chose this dataset. It also means the label is the consumer’s own choice, and some complaints fit two products equally well. This becomes visible in the results. CFPB uses twelve product names for nine areas. I mapped them to nine short team names and wrote one line of description for each. The descriptions are my own wording, and both models get exactly the same ones. The raw data is very uneven. Most of the complaints are about credit reporting, and a random sample would mostly measure how well a model recognises those. So I took 100 complaints per team, 900 in total, from 2018 onwards and between 200 and 2,000 characters long. The sample is drawn with fixed random seeds and saved to a file, so every model sees the same complaints. All accuracy figures below are therefore for an equal mix of teams. Everything ran on a MacBook Pro 14-inch 2021 with an Apple M1 Pro 8 CPU cores, 14 GPU cores and 32 GB of memory, with macOS 27.0.1 and Ollama 0.35.1. ollama pull nimbleollama pull qwen3.5:9b-q8 0 The second line needs a comment. The default tag qwen3.5:9b is a 4-bit build see the list of tags https://ollama.com/library/qwen3.5/tags , and nimble is 8-bit. A 4-bit model can be a little less accurate than the 8-bit build of the same weights. If I compared the two default tags, I would not know which difference I was measuring. So I used the 8-bit tag of Qwen. All numbers below are 8-bit against 8-bit. Nimble is not a chat model, so ollama run nimble does not work, and you call it with POST /v1/systemone. That endpoint rejects chat models, so Qwen goes through /api/chat and the test needs two code paths. The first request loads the model and takes several seconds. I send a warm-up request first, and it must not be one of the test items, because that item would later be answered from the cache in 0.1 seconds. This is the smallest request that works, with three of the nine teams: curl -s http://localhost:11434/v1/systemone \ -H 'Content-Type: application/json' -d '{ "model": "nimble", "state": {"complaint": "The escrow amount on my mortgage statement went up without notice."}, "questions": { "team": { "type": "choice", "instructions": "Which product team should handle this complaint?", "criteria": { "Mortgage": "Home loans, mortgage servicing, escrow, foreclosure", "Credit card": "Credit cards and prepaid cards: charges, fees, disputes, rewards", "Bank account": "Checking or savings accounts, deposits, withdrawals, overdrafts" } } }}' { "model": "nimble", "answers": { "team": { "type": "choice", "choice": "Mortgage", "probabilities": { "Mortgage": 0.9985900450531108, "Credit card": 0.00025204474220624415, "Bank account": 0.0011579102046829846 }, "confidence": 0.9896904755809427 } }, "usage": { "input tokens": 211, "output tokens": 1 }} The model counts 211 input tokens for a complaint of one sentence, because it also reads the question and the description of every team. This becomes important in the speed section. The reply also contains a confidence value, which I do not use. In this article, probability always means the number the model gives for the team it chose. In the example above that is 0.9986 for Mortgage. Qwen gets a normal prompt built from the same question, the same nine team names and the same descriptions. A JSON schema limits the answer to the nine team names, so the model cannot return anything else: "format": {"type": "object", "properties": {"team": {"type": "string", "enum": list TEAMS }}, "required": "team" },"think": False,"options": {"temperature": 0}, think switches Qwen’s reasoning on or off. Here is one complaint asked in all three ways: Nimble got 694 of the 900 complaints right, 77.1%. Qwen with reasoning off got 687 right, 76.3%. The margin of error is about 2.8 points for each, so a difference of 0.8 points is well inside it. The vendor reports that training lifts the base model from 66% to 90% on its own test set. On this task the base model already reached 76%, and the training added nothing I could measure. A paired comparison says the same. 53 complaints were right only with Nimble and 46 only with Qwen, and an exact McNemar test https://en.wikipedia.org/wiki/McNemar%27s test on those gives p = 0.55. The green bar, and the green dots in the next chart, are Qwen with reasoning on. I cover that further down. Neither model leads everywhere. The models disagree most on personal loans, money transfers, credit reporting and debt collection. Nimble is clearly better on the first three, and Qwen is clearly better on debt collection 66 of 100 against 53 . Nimble’s 97 of 100 on credit reporting has a cost. It answered “Credit reporting” 196 times, although only 100 complaints belong there. Qwen did so 147 times. On an equal mix this habit costs Nimble accuracy on other teams. On the real mix, where most complaints are about credit reporting, it would help. 160 complaints were wrong with both models, and in 137 of those both chose the same wrong team. In the most common case, 26 complaints filed under debt collection went to Credit reporting. Another 13 filed under money transfer went to Bank account. When I read these complaints, the reason was clear. Here are two that were filed under debt collection: “I was looking at my credit report and noticed my credit score drop a lot and seen that I had two collections on my report that I didn’t know about or authorized. … So I proceeded to contact XXXX and XXXX and the police department and also all three credit bureaus.” “Any debt that I have had with XXXX XXXX XXXX and XXXX XXXX XXXX has been paid. I submitted proof of payment and this is still showing on my report.” Both are about a collection entry on a credit report. A person could file them under either product, and either team could reasonably handle them. About two-thirds of each model’s errors are of this kind, where both models chose the same wrong team. I see these as a problem of the categories more than of the models. With real intake data I would not expect accuracy close to 100%, and I would review the categories before tuning the model. The median time per complaint was 2.25 seconds for Nimble and 2.07 seconds for Qwen with reasoning off. The slowest 5% took more than 3.2 seconds for Nimble and 3.0 for Qwen. The full run took 35 minutes for Nimble and 33 for Qwen. One request at a time, that is about 1,500 complaints per hour for Nimble and about 1,650 for Qwen. So the decision model is not faster than its base model. It is slightly slower. The chart shows why. For both models the time grows in a straight line with the number of tokens the model has to read. Both read at about 240 tokens per second, which makes sense for two models of the same architecture, size and precision. The two models differ in two ways, and the differences cancel each other out. Qwen has to write its answer, about ten tokens of JSON. That gives it a fixed cost of 0.69 seconds per request. Nimble writes nothing, and its fixed cost is 0.12 seconds. But Nimble’s request is longer. For the same complaint and the same nine descriptions, Nimble counts a median of 502 input tokens and Qwen 331. The extra 170 tokens are probably overhead of the typed request format. At 240 tokens per second they cost about 0.7 seconds, which uses up the saving. For a design this means that latency depends on how much the model reads. The number of options and the length of their descriptions are a cost I control, and nine one-line descriptions already make up most of a short request. It also means that once reasoning is off, a general model writes so little that a decision model has almost nothing left to save. The published figure for Nimble is 91 ms per decision, measured in a game demo on a MacBook Pro M5 Max. My median is 2.25 seconds, about 25 times more. A faster chip, a shorter input and caching could each explain part of that. I could test two of the three. For the short input I sent only the first 100 characters of a complaint, with three teams to choose from. That is 228 input tokens, close to the smallest request above. It took 1.06 seconds, which is what the straight line in the speed chart predicts. A short input alone still takes more than ten times the published figure. An identical request is different. Sent a second time, it comes back in 0.10 seconds, which is in the range of the published figure. Any change costs the full time again. I added one sentence at the end of the complaint, then one word at the start, and both took about 2.55 seconds, the same as a new complaint. A different question about the same complaint took 1.82 seconds, and it is faster only because that request is shorter. The cache works on whole requests only. Qwen’s identical request takes 0.64 seconds, because it still has to write the answer token by token. On a cache hit the decision model is therefore six times faster. It is the only case where it made a difference that Nimble writes no text. The sample happens to contain an example. One complaint was filed twice with the CFPB and appears twice in the 900. Nimble answered the second copy in 0.15 seconds. It is the single dot at the bottom of the first chart in this section. So on an M1 Pro I got 100 ms only for inputs that repeat exactly, such as polling an unchanged state, retries, or a game loop where nothing has moved. The published figure is from an M5 Max, which I could not test. For new inputs it helps to ask everything in one request. Three questions about a new complaint took 3.35 seconds in one request, against an estimated 6.2 seconds asked one by one. The request counts 1,870 tokens, but the time is far below what the line predicts for that many, which suggests the complaint is read only once per request. Qwen3.5 is a reasoning model. With think enabled it writes out its reasoning before it answers. I ran that mode on all 900 complaints as well. It got 716 right, 79.6%. Reasoning made the model about 3 points more accurate. It got 29 more complaints right than the same model without reasoning, and 22 more than Nimble. Unlike the gap between Nimble and its base model, this one holds up in a paired test p = 0.001 against reasoning off, p = 0.007 against Nimble . It also took much longer. The run needed 20 hours and 35 minutes. That is 38 times as long as the same model with reasoning off, and 35 times as long as Nimble. I was surprised by the size of that factor, so I looked at where the time goes. A model reads its input in one pass, here at about 240 tokens per second. It writes one token at a time, and every token needs a full pass through the model. That is about 18 tokens per second. Without reasoning, Qwen writes about 10 tokens per complaint. With reasoning the median was 862, with no limit set on the length. Over the whole run the model wrote 1.3 million tokens, against 9,300 without reasoning. Faster hardware raises both speeds, but the reasoning model still has to write that much more. Each dot in this chart is one complaint, and a hollow dot is a wrong answer. Reasoning time is hard to predict. The fastest complaint took 29 seconds and the slowest 9 minutes 51 seconds. 330 of the 900 took longer than a minute, and 108 took longer than three. The two single-pass models stay between 1.4 and about 4 seconds, apart from the one cache hit. The longest runs were more often wrong. The median was 45 seconds when the reasoning model was right and 89 seconds when it was wrong. For the slowest complaint of all the model wrote 10,232 tokens and still ended on the wrong team. I read long reasoning as a sign that the complaint is unclear. Reasoning also does not fix the overlapping categories. Of the 184 complaints the reasoning model got wrong, 144 were also wrong with both other models. The gain comes from three teams: personal loans 63 of 100 against 49 , money transfers 75 against 61 and credit reporting 99 against 88 . On debt collection it got worse 56 against 66 . This is the comparison in which a decision model is much faster. Against a model that reasons first, Nimble needs 35 times less time and gives up about 2.5 points of accuracy. The base model with reasoning off gets the same speed-up, so the speed comes from skipping the reasoning. A decision model is one way to do that, and switching reasoning off is the other. Up to here the decision model has only matched its base model. It does better when I use the probability as well as the answer. If automatically routed complaints must be 95% correct, Nimble can take 43% of the volume without a person. Its base model can take 16%. This matters for a common design, where the model handles the cases it is sure about and sends the rest to a person. For that design I need to know how much of the work the model can take over at the accuracy I need. To find out, I sort all answers by the model’s probability, most confident first, and go down the list. For each required accuracy I report the largest share that still reaches it. Nimble returns a probability for each of the nine teams. Qwen does not, so I built one. I ran Qwen a third time on all 900 complaints, again with reasoning off. In this run every team has a letter from A to I, and Qwen answers with one letter. I asked Ollama for the 20 most likely tokens at that position, took the probabilities of the nine letters among them, and scaled them to add up to 1. Accuracy in this run was 75.6%, within the margin of the first run. The orange line in the chart is this run. Read the chart along the 95% line. The orange dot is at 16% and the blue dot at 43%. The share each model can route at three accuracy targets: Target Nimble Qwen------ ------ ----85% 76% 74%90% 63% 58%95% 43% 16% At 85% the two are nearly equal, and at 90% they are still close. They move apart at the strict end. Both models make about the same number of mistakes overall, but Nimble’s highest probabilities are more reliable. I checked this result in two ways. First, I repeated the comparison 2,000 times on the recorded answers. Each time I drew 900 complaints at random from the original 900, where the same complaint can be drawn more than once. This shows how much the result depends on which complaints happen to be in the sample. Nimble was ahead at the 95% target in 99% of the 2,000 runs. The size of its lead varies, with Nimble’s share between 34% and 56% and Qwen’s between 0% and 45%. Second, Qwen’s probability is something I built myself, so I tried three ways of building it. They gave between 16% and 29% at the 95% target. Nimble’s 43% is above all three. The green line is the reasoning model, with the probability of the answer it wrote after reasoning. The line is flat. The model gives a probability of 0.99 or more for 883 of its 900 answers, and only 80% of those are right. After reasoning, the probability no longer separates right answers from wrong ones. The probabilities rank answers well, but they are not calibrated. Calibrated would mean that answers given with a probability of 0.7 are right about 70% of the time. When Nimble gives a probability between 0.6 and 0.8, it is right in 46% of the cases. Of its 356 answers with a probability of 0.99 or more, 16 were wrong. I would therefore use the probability only to rank answers, and set the cut-off on a labelled sample of my own data, at the accuracy I need. I would check it again when the categories or the data change. Based on these results, I would put a decision model at the front of a flow, where something has to decide what can be handled automatically and what needs more attention. This is an escalation ladder. The cheapest step sees every item. Each step acts on the cases it is confident about and passes the rest on. A person sees the fewest. The test covers the first gate. Nimble’s probability separated safe answers from unsafe ones well enough to act on, at about two seconds per complaint. The larger model in the diagram was not part of this test. A second step with a model of the same size did not pay off. I tried it with the recorded answers. Every complaint where Nimble’s probability is below 0.95 goes to Qwen with reasoning on. That is 299 complaints, a third. On those the reasoning model was right 180 times and Nimble 158 times. On the other 601 each model was right 536 times. Overall accuracy rises from 77.1% to 79.6%, the same as reasoning on everything, and the total time goes from 35 minutes to about 12.5 hours. I would not accept about twelve more hours for about 2.5 points. A second step makes sense only if it is clearly more capable, so a larger model or a person. The second step also cannot use its own probability as a gate, because after reasoning it was 0.99 or more for almost every answer. Checking whether two models agree did not help either. When I route only the complaints where Nimble is at 0.95 or more and Qwen gives the same answer, I get 65% of the complaints at 90.3% accuracy. Nimble alone gives 67% at 89.2%. Two models from the same weights make the same mistakes, so their agreement is not an independent check. I would keep the decisions people make at the end of the ladder as new labels, because they are the cheapest way to keep checking the cut-offs as the data changes. On this task, the decision training did not make the model more accurate, and it did not make it faster than its own base model with reasoning off. It did make the model’s highest probabilities more reliable. That is less than the announcement's promise, but it matters for the gate design above, because more work can be automated at a strict quality level. The clearest case is the first gate in an “automate or escalate” flow, as described above. It is the use where Nimble was ahead of its base model on new input. It also fits when there are many questions about the same item. Putting them in one request is natural in this API, and three questions took 3.35 seconds against an estimated 6.2. And it fits input that repeats exactly, such as polling, retries and state machines. There I measured about 100 ms, six times faster than the base model on a cache hit. It can also replace a same-size reasoning model on a simple, well-defined decision when time matters. That cost about 2.5 points of accuracy and saved 97% of the time, with latency I can predict. I would not replace a non-reasoning call that already returns one word under a schema. Accuracy and speed are the same, and there is a second API to maintain. Valid output is no reason to switch either. Neither model gave a single invalid answer in 900 requests. I would also not use it when someone needs a reason. A decision model returns no explanation, so if an agent or an auditor has to see a justification, a general model is still needed. Code, the 900 complaints and all raw results are on GitHub, in the folder for this article https://github.com/kumarkeviv/articles/tree/main/2026-10-nimble-decision-model of my articles repo. ollama pull nimble && ollama pull qwen3.5:9b-q8 0 about 35 minutespython3 usecase/cfpb.py nimble nimble about 33 minutespython3 usecase/cfpb.py chat qwen3.5:9b-q8 0 reasoning on, about 21 hourspython3 usecase/think.py Qwen with letter answers, about 33 minutespython3 usecase/letters.py about 3 minutespython3 usecase/cache.py tables and chartspython3 usecase/report.py The scripts need only Python 3. The charts also need Google Chrome, and report.py --no-png skips them. The four model runs can be stopped and restarted, and they continue where they stopped. Run one model at a time and nothing else on the GPU, or the times are not comparable. Evaluating a decision model: Nimble 9B and its base model Qwen3.5–9B https://blog.devgenius.io/evaluating-a-decision-model-nimble-9b-and-its-base-model-qwen3-5-9b-a845c2bbd7f0 was originally published in Dev Genius https://blog.devgenius.io on Medium, where people are continuing the conversation by highlighting and responding to this story.