Can Decision Models Abstain? A benchmark of four local System One decision models and the hosted Jev 1.13.0 reference on 18 synthetic customer messages found that only Jev 1.13.0 classified all 18 correctly, while the local models scored 15/18, 13/18, 17/18, and 14/18, with abstention accuracy ranging from 1/6 to 5/6 on the six messages where no current action was clear. The test, run 2026-10-08 through Ollama's System One API and the TryJevAI playground, asked each model to choose cancel_order, check_order_status, or insufficient_information using only the customer's message and no conversation history. All five models handled the 12 clear requests nearly perfectly, but abstention on ambiguous messages such as "I might cancel if it takes much longer" was the main source of error. back to all notes https://elma.dev/notes · applied ai Can Decision Models Abstain? “I might cancel if it takes much longer.” Should a message like this go into a cancellation workflow? I tested four local System One decision models and a hosted Jev reference on that question. Each received a customer message and had to choose: cancellation request, status request, or insufficient information. I wanted to see whether they could leave the decision open when neither action was clearly requested. Four local models. One hosted reference. Eighteen synthetic messages: six cancellations, six status requests, and six cases where no current action is clear. 15/18 clear requests 12/12 · abstentions 3/6 13/18 clear requests 12/12 · abstentions 1/6 17/18 clear requests 12/12 · abstentions 5/6 14/18 clear requests 11/12 · abstentions 3/6 18/18 clear requests 12/12 · abstentions 6/6 I might cancel if it takes much longer. Expected: Abstain Abstain Abstain Read the exact instruction Classify the customer's currently requested action using only their message. Choose cancel order when the customer clearly asks to cancel or stop their order, including clear paraphrases. Choose check order status when the customer asks about the order's current status, shipping progress, or delivery time. Choose insufficient information when neither action is clearly requested. A complaint, a quoted instruction, or a hypothetical future cancellation is not itself a current cancellation request. Pay attention to negation and to the difference between a past request and the current request. Infer the meaning of clear paraphrases, but do not invent an action from dissatisfaction alone. The scope of this benchmark is messages with at most one current action from these two categories; insufficient information is a decision to abstain, not a customer intent. Local results use /v1/systemone; the hosted Jev reference uses TryJevAI. Raw requests and responses are linked below. 2026-10-08. One recorded answer per model/case. Matches are agreement with the stated policy. Scores are not validated probabilities of customer intent. Local API confidence measures score concentration. Hosted confidence and latency are shown as returned by the playground; timings are not directly comparable. A small routing experiment a-small-routing-experiment This is a classification step between an incoming message and a fixed workflow. The model returns a label; the surrounding application decides what happens next. Nothing in this test changed an order. I used Nimble 9B https://ollama.com/library/nimble , Clef Flash 9B https://ollama.com/library/clef-flash , Tev1 4B https://ollama.com/library/tev1 , and Laya 421M https://ollama.com/library/laya , plus Jev 1.13.0 via TryJevAI https://tryjevai.com/ . TryJevAI is an independent playground. Jev's parameter count is not verified here, so this is a hosted-service reference, not an equal-size comparison. There are 18 synthetic English messages: six cancellation requests, six status questions, and six where neither action is clear. The local models use Ollama's System One API https://docs.ollama.com/capabilities/decision . Jev received the same messages, policy, and ordered labels through the playground. Its question field rejects text over 300 characters, so I put the full policy in state beside the message and used a short question. That format difference is part of this comparison. Each message is evaluated alone, with no conversation history . Expand “Read the exact instruction” above to inspect the policy. The rule is to identify the customer's current requested action . Clear paraphrases count. Complaints, quotes, and possible future cancellations don't count as requests by themselves. insufficient information , shown as Abstain , is a third answer choice. I didn't apply a confidence threshold after the model answered. What happened what-happened Nimble, Clef Flash, and Tev1 matched all 12 clear requests. Laya matched 11. On the six messages labelled for abstention, Tev1 matched five, Nimble and Laya three each, and Clef Flash one. The hosted Jev reference matched 18/18 labels: 12/12 clear requests and 6/6 expected abstentions. The explorer opens on the conditional cancellation. Nimble and Tev1 abstained. Clef Flash and Laya chose cancellation, even though the customer only described a possible future action. Jev abstained on this message. “Can you help me with order 314?” is more debatable. A status update could be helpful, but the message doesn't specify what help is wanted. My expected label is abstention. These scores measure agreement with that policy, not an objective reading of someone's hidden intent. “You said ‘cancel the order’ in your last message” refers to missing history. Its expected answer only reflects what can be inferred from the supplied message. What this tells me what-this-tells-me I'd test the review path explicitly before using these labels to trigger a workflow. A single total hides two different mistakes: missing a clear request, and selecting an action when neither is clearly requested. This is an exploratory set, not a held-out benchmark. One changed answer moves a total by 5.6 percentage points. Always abstaining would match 6/18 labels while handling no clear requests. I'd want real messages, independent label review, and prompt and option-order tests before choosing a production model. The bars show the API's label scores. This experiment doesn't establish calibration. The local API confidence value describes score concentration, not correctness. Jev's confidence is shown as returned by the playground. Run details run-details One recorded answer per model and message, grouped by model; fixed option order: cancellation, status, abstention. The Tev1 and Laya runs kept the original inputs and key, using Ollama 0.40.1 defaults on an M3 MacBook Air with 24 GB memory. I restarted Tev1 after an initial connection failure returned no answer. Local timings include loading and HTTP overhead. Jev's displayed latency is the playground's reported value. They aren't directly comparable. The hosted calls were paced to respect the playground's rate limit. Download the requests and responses https://elma.dev/benchmarks/uncertainty/customer-intent-results.jsonl and run details https://elma.dev/benchmarks/uncertainty/run-details.json , including model tags and local digests. I used AI tools to help with the code and edit the text. The results are from local models and the hosted endpoint; the messages are synthetic.