cd /news/artificial-intelligence/can-decision-models-abstain · home › topics › artificial-intelligence › article
[ARTICLE · art-147775] src=elma.dev ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Can Decision Models Abstain?

A benchmark of four local System One decision models and the hosted Jev 1.13.0 reference on 18 synthetic customer messages found that only Jev 1.13.0 classified all 18 correctly, while the local models scored 15/18, 13/18, 17/18, and 14/18, with abstention accuracy ranging from 1/6 to 5/6 on the six messages where no current action was clear. The test, run 2026-10-08 through Ollama's System One API and the TryJevAI playground, asked each model to choose cancel_order, check_order_status, or insufficient_information using only the customer's message and no conversation history. All five models handled the 12 clear requests nearly perfectly, but abstention on ambiguous messages such as "I might cancel if it takes much longer" was the main source of error.

read5 min views3 publishedOct 8, 2026

back to all notes · applied ai

“I might cancel if it takes much longer.” Should a message like this go into a cancellation workflow? I tested four local System One decision models and a hosted Jev reference on that question. Each received a customer message and had to choose: cancellation request, status request, or insufficient information. I wanted to see whether they could leave the decision open when neither action was clearly requested.

Four local models. One hosted reference.

Eighteen synthetic messages: six cancellations, six status requests, and six cases where no current action is clear.

15/18 clear requests 12/12 · abstentions 3/6

13/18 clear requests 12/12 · abstentions 1/6

17/18 clear requests 12/12 · abstentions 5/6

14/18 clear requests 11/12 · abstentions 3/6

18/18 clear requests 12/12 · abstentions 6/6

I might cancel if it takes much longer.

Expected: Abstain

Abstain

Abstain

Read the exact instruction #

Classify the customer's currently requested action using only their message. Choose cancel_order when the customer clearly asks to cancel or stop their order, including clear paraphrases. Choose check_order_status when the customer asks about the order's current status, shipping progress, or delivery time. Choose insufficient_information when neither action is clearly requested. A complaint, a quoted instruction, or a hypothetical future cancellation is not itself a current cancellation request. Pay attention to negation and to the difference between a past request and the current request. Infer the meaning of clear paraphrases, but do not invent an action from dissatisfaction alone. The scope of this benchmark is messages with at most one current action from these two categories; insufficient_information is a decision to abstain, not a customer intent.

Local results use /v1/systemone; the hosted Jev reference uses TryJevAI. Raw requests and responses are linked below. 2026-10-08. One recorded answer per model/case. Matches are agreement with the stated policy. Scores are not validated probabilities of customer intent. Local API confidence measures score concentration. Hosted confidence and latency are shown as returned by the playground; timings are not directly comparable.

A small routing experiment #

This is a classification step between an incoming message and a fixed workflow. The model returns a label; the surrounding application decides what happens next. Nothing in this test changed an order.

I used Nimble 9B, Clef Flash 9B, Tev1 4B, and Laya 421M, plus Jev 1.13.0 via TryJevAI. TryJevAI is an independent playground. Jev's parameter count is not verified here, so this is a hosted-service reference, not an equal-size comparison.

There are 18 synthetic English messages: six cancellation requests, six status questions, and six where neither action is clear. The local models use Ollama's System One API. Jev received the same messages, policy, and ordered labels through the playground. Its question field rejects text over 300 characters, so I put the full policy in state beside the message and used a short question. That format difference is part of this comparison. Each message is evaluated alone, with no conversation history. Expand “Read the exact instruction” above to inspect the policy.

The rule is to identify the customer's current requested action. Clear paraphrases count. Complaints, quotes, and possible future cancellations don't count as requests by themselves. insufficient_information, shown as Abstain, is a third answer choice. I didn't apply a confidence threshold after the model answered.

What happened #

Nimble, Clef Flash, and Tev1 matched all 12 clear requests. Laya matched 11. On the six messages labelled for abstention, Tev1 matched five, Nimble and Laya three each, and Clef Flash one.

The hosted Jev reference matched 18/18 labels: 12/12 clear requests and 6/6 expected abstentions.

The explorer opens on the conditional cancellation. Nimble and Tev1 abstained. Clef Flash and Laya chose cancellation, even though the customer only described a possible future action. Jev abstained on this message.

“Can you help me with order #314?” is more debatable. A status update could be helpful, but the message doesn't specify what help is wanted. My expected label is abstention. These scores measure agreement with that policy, not an objective reading of someone's hidden intent.

“You said ‘cancel the order’ in your last message” refers to missing history. Its expected answer only reflects what can be inferred from the supplied message.

What this tells me #

I'd test the review path explicitly before using these labels to trigger a workflow. A single total hides two different mistakes: missing a clear request, and selecting an action when neither is clearly requested.

This is an exploratory set, not a held-out benchmark. One changed answer moves a total by 5.6 percentage points. Always abstaining would match 6/18 labels while handling no clear requests. I'd want real messages, independent label review, and prompt and option-order tests before choosing a production model.

The bars show the API's label scores. This experiment doesn't establish calibration. The local API confidence value describes score concentration, not correctness. Jev's confidence is shown as returned by the playground.

Run details #

One recorded answer per model and message, grouped by model; fixed option order: cancellation, status, abstention. The Tev1 and Laya runs kept the original inputs and key, using Ollama 0.40.1 defaults on an M3 MacBook Air with 24 GB memory. I restarted Tev1 after an initial connection failure returned no answer. Local timings include and HTTP overhead. Jev's displayed latency is the playground's reported value. They aren't directly comparable. The hosted calls were paced to respect the playground's rate limit.

Download the requests and responses and run details, including model tags and local digests.

I used AI tools to help with the code and edit the text. The results are from local models and the hosted endpoint; the messages are synthetic.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ollama 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-decision-models-…] indexed:0 read:5min 2026-10-08 · —