# Integrating Decision Models into Agent Harnesses

> Source: <http://vivekhaldar.com/articles/llms-do-too-much-offload-to-decision-models/>
> Published: 2026-10-10 00:00:00+00:00

# Integrating Decision Models into Agent Harnesses

By Vivek Haldar

Agentic automation so far has had only one kind of engine: LLMs. We’ve used LLMs as everything-models. In most agentic work they carry out a range of task types: reasoning, planning, natural language understanding and synthesis, classification and prediction. With the recent introduction of Jev, and the creation of the “decision model” category, we’re seeing that some of those tasks (e.g. classification, binary prediction) can be offloaded to a model that is much faster and cheaper.

The design question then becomes: how can we incorporate decision models into our agent harnesses in a way that we’re sending it the right kind of tasks? This should result in deflecting some calls from the mainline LLM, which in turn should help reduce both overall cost and latency, while maintaining or improving task performance.

I spent a weekend trying two ways of embedding Jev in [Proceda](https://github.com/vivekhaldar/proceda/), my state-of-the-art specialized harness for turning standard operating procedures into agents.

## Using Jev inside the harness

Proceda takes a natural-language SOP and executes it one step at a time. That gives each LLM invocation a fairly narrow scope. The harness already knows the current step and manages progression through the procedure. Even within that constrained setting, I suspected the LLM was doing work that could be handed off. (I’ve written more about the architecture in [Anatomy of a SOTA Agentic SOP-Execution Engine](https://enchiridionlabs.online/sop-execution-engine-design.html), and the broader design principles in [Specialized Harness Engineering](https://enchiridionlabs.online/specialized-harness-engineering.html).)

I tested this on three domains from [SOP-Bench](https://github.com/amazon-science/SOP-Bench): referral abuse detection, email intent, and patient intake (a subset of the benchmark, just because I wanted to get a quick read). That’s 360 cases altogether. The main LLM was Qwen 3.8 27B, with hosted Jev as the decision model.

The implementation and experiment reports are on the [`codex/bounded-business-decisions` branch](https://github.com/vivekhaldar/proceda/tree/codex/bounded-business-decisions). The [HTML slide deck I used in the video](https://github.com/vivekhaldar/proceda/blob/codex/bounded-business-decisions/docs/decision-model-experiments.html) is in the same repo; download it and open it in a browser to view the slides.

## Experiment 1: Is this step finished?

After an agent makes a tool call, it will often go back to the LLM to interpret the result and decide whether the current step is done, or needs more work (tool calls or reasoning) to finish. This decision step sounded like a good fit for Jev.

I put Jev after eligible tool calls, giving it the current step, the relevant context, and the tool calls and their results. It answers two questions: have all the necessary tool calls been made, and is the evidence sufficient to complete this step?

If both yes probabilities reach 0.8, Proceda advances. Otherwise, the main LLM continues as usual. The tools still run; the saving comes from avoiding the follow-up LLM turn.

The [completion-check experiment](https://github.com/vivekhaldar/proceda/blob/codex/bounded-business-decisions/experiments/decision-cascades-2026-10-04/qwen-jev-080/REPORT.md) showed:

- **605 fewer LLM calls** , a**23.8% reduction** .
- Task performance was (nearly) the same: **357/360 final tasks correct for Qwen-only, versus 356/360 for the hybrid** .

## Experiment 2: Make the business decision

Experiment 1 used a decision model for the harness’s internal operation. A natural question is: what about steps in the SOP itself are decisions? A lot of business logic questions in SOPs turn out to be the right fit for decision models. This was Typesafe AI’s core hypothesis.

Take email intent. Given a seller’s email, the SOP asks which kind of issue it raises: a generic listing question, incorrect pricing, a product not being listed, an incorrect description, or an inability to decide. This is a great match for the “classify” call in a decision model.

Jev can receive the email and the original policy, choose an outcome, and hand that result back to the harness. If its selected-outcome probability is below 0.8, the decision falls back to Qwen. The rest of the workflow still has to execute.

The [email comparison](https://github.com/vivekhaldar/proceda/blob/codex/bounded-business-decisions/experiments/decision-cascades-2026-10-04/business-decisions-080/REPORT.md#direct-business-decision-result-email) showed:

- **141 of 148 intent decisions handled by Jev** , or**95.3%** .
- **Seven decisions fell back to Qwen** .
- **148/148 final tasks correct with either model** behind the structured interface.
- **25.4% fewer LLM calls across the whole email workflow** , from 556 to 415.

The later LLM call count was identical, so those 141 avoided calls can be attributed directly to offloading intent classification.

## The combined impact

With both changes in place, the [combined experiment](https://github.com/vivekhaldar/proceda/blob/codex/bounded-business-decisions/experiments/decision-cascades-2026-10-04/business-decisions-080/REPORT.md) showed:

- **Nearly a 30% reduction in main-LLM calls** : from 2,537 to 1,785, saving 752 calls (29.6%).
- **Nearly 27% lower estimated execution API cost** : from $6.41 to $4.70 for the 360-case batch, including Jev requests (26.6%).
- **Nearly 20% lower mean case latency** : from 8.64 seconds to 6.94 seconds (19.7%).
- **Higher observed task accuracy in the combined system** : from 357/360 final tasks correct to 360/360.
- **Nearly 25% lower estimated API cost even after preparation** : from $6.41 to $4.83, including the full final preparation cost for the structured SOP plans (24.6%).

Almost a third fewer calls to the main LLM, nearly 27% lower execution API cost, and nearly 20% lower mean latency—with higher observed task accuracy in the combined system—is a useful set of gains. And it came from finding two specific places in the harness where the required answer was much narrower than a general-purpose generation.

## Decision models as a cache

The analogy I keep coming back to is a **cache**. Think of the main LLM as main memory or disk, and the decision model as an L1 or L2 cache that absorbs some of the traffic. Here the smaller model computes an answer rather than retrieving a stored one, but the architectural role is similar: handle the requests you can cheaply, and fall through when you can’t.

Local decision models make this more interesting, as I explored in [Local Decision Models + Jev: Cut Latency With Confidence Cascades](https://www.youtube.com/watch?v=HBVInCcfH1c). In my separate experiments, the response times I observed were:

- **Roughly 30 milliseconds for Laya** running on my M5 MacBook.
- **Roughly 250–300 milliseconds for hosted Jev** . (Though I suspect a lot of that is network latency)

What makes these models useful is the harness around them. It supplies the evidence, defines the choices, validates the result, and handles uncertainty. Once those boundaries are explicit, you can start asking which of the LLM’s jobs really belong there.

In Proceda, two good places to start were “is this step done?” and “which of these outcomes applies?” I suspect a lot of agents are spending LLM calls on those same two questions.
