cd /news/ai-products/your-agent-burns-llm-money-on-switch… · home topics ai-products article
[ARTICLE · art-132394] src=dev.to ↗ pub= topic=ai-products verified=true sentiment=· neutral

Your Agent Burns LLM Money on Switch Statements. Jev Claims 444x Less

TypeSafe AI released Jev, a "System One" model that answers structured choice, score, and boolean questions instead of generating text, claiming 193.6x faster and 444.6x cheaper performance than frontier LLMs on agent routing tasks. Rivet's Nathan Flurry characterized the model as "just a really smart switch statement" that does not replace GPT or Claude, and TypeSafe's own launch blog concedes the headline multiples represent the "higher end of real world gains" from a "highly simplified" demo. Jev became available on Vercel's AI Gateway on 16 September, with direct API keys in early access.

by read15 min views2 publishedSep 17, 2026

You pull last month's model calls for your agent and sort them by

what they were for. Some are easy to justify. The agent wrote a

patch, drafted a reply to a customer, summarised a stack trace.

Then there is everything else. The agent asked a frontier model which

tool to use next, and whether a failed step deserved a retry. It

asked whether rm -rf ./build was safe to run. It asked if the JSON

a tool sent back looked sane. Each of those went out at frontier

prices and came back as a handful of tokens you parsed into an enum.

Look at the answers those calls can give. Pick one of four tools.

Continue, retry, ask the user or stop. Safe enough, or not. A fixed

set of cases, and some judgment to pick the right one. You have

written that shape a thousand times. It is a switch statement.

Nathan Flurry of Rivet, writing on X

about a new model from TypeSafe AI, put it plainly: "jev is just a

really smart switch statement". He also wrote "jev does not replace

gpt / claude". He is right on both counts.

TypeSafe's co-founder Diogo Almeida

announced Jev on X on 15 September

You give it a structured text state and a set of questions. The

docs define three question

types. Choice picks one option. Score places the state on an ordered

scale. Noul returns a yes/no probability, and the AI SDK calls it

boolean. You can mix all three in one request, and they are

evaluated in parallel.

What Jev gives up is string generation. It cannot write a sentence,

a query or a file path.

The homepage also claims "Zero Hallucinations". TypeSafe's own

launch blog

explains what that means: "Our number is not empirical. Schema

matching is guaranteed". Every answer fits the options you declared.

It can still pick the wrong option.

Direct API keys from TypeSafe are in early access. On 16 September,

Vercel's changelog

said Jev "is now available on AI Gateway".

The homepage banner reads "193.6x Faster, 444.6x Cheaper". Its

footnote says the multiples are "based on workflows for System One

tasks".

The page also shows a side-by-side example. Jev answers in 0.114s for

$0.000081. The LLM answers in 8.566s for $0.013880. Those two figures

divide to about 75x faster and 171x cheaper, so the banner is not the

ratio of that one run.

The launch blog qualifies the demo itself. It calls the query

"highly simplified", says the shorter input "paints our model in an

advantageous light", and describes the banner figure as "the higher

end of real world gains". Treat 444.6x as a best case, because the

vendor does.

The launch blog's pricing shows where a gap that size can come from.

Input is $0.042 per million tokens ($42 per billion). Output tokens

are listed as "FREE (too cheap to meter)". For comparison,

The Register

gives GPT-5.6 Terra's price as $2.00 input and $12 output per

million. That is about 48x on input, and the output side has no ratio

because one side is zero. TypeSafe quotes end-to-end latency of 70ms

to 500ms.

TypeSafe also publishes workflow evals.

Two things before the numbers. The evals are vendor-run, and the blog

concedes they may carry "some bias". And "accuracy" means agreement

with reference labels "generated via an average of the responses of

GPT-6 Astra and Claude Fable 5.1, both at high thinking". The other

models ran at their provider's default reasoning. A score here tells

you how often a model agrees with two frontier models, which is a

narrower thing than how often it is right.

Aggregate results across the four workflows, for six of the nine

models on the page, vendor-reported:

Model Accuracy Cost per case Latency
Jev 67.8% $0.0004 0.4s
GPT-5.6 Luna 66.8% $0.0033 12.9s
GPT-5.6 Terra 67.9% $0.0304 10.1s
GPT-5.6 Sol 74.1% $0.0836 23.3s
Claude Sonnet 5 67.8% $0.1174 78.1s
Claude Opus 5 73.1% $0.1761 37.8s

Start with Terra, because the accuracy is level: 67.8% against

67.9%. On these numbers Jev costs about 1/76th as much per case and

answers about 25 times faster. The costs are rounded to four decimal

places, so read 76x as approximate.

Claude Sonnet 5 matches Jev exactly at 67.8%, for about 294 times the

cost per case and 195 times the latency.

Now Sol. It is 6.3 points more accurate than Jev and about 209 times

more expensive per case.

The only row in that table whose multiple comes near the banner's

444.6x is Claude Opus 5, at about 440x, and Opus 5 is the second most

accurate model in the table.

The same eval page tested each model two ways: one big prompt, and a

workflow that decomposes the task into Choice, Score and Noul

questions. Every LLM got more accurate, cheaper and faster in

workflow mode. GPT-5.6 Luna went from 51.9% to 66.8%.

Put that next to the table. Decomposing the task bought Luna 14.9

points. Moving the decomposed workflow from Luna to Jev bought one

more point and an 8x lower cost per case.

That changes the order of the work. You can decompose today, with the

model you already call. Break "what should the agent do now?" into

small typed questions and combine the answers in your own code. The

docs call it "atomic questions, composed in code". Once that is done,

the model question gets narrow: for these atomic questions, which

model agrees with your labels often enough at the lowest cost?

Vercel's changelog lists the agent use cases it has in mind:

selecting tools or subagents, the next action (continue, retry, ask

the user, halt), urgency or risk before an operation, and validating

outputs and safeguards. Guillermo Rauch, Vercel's CEO,

wrote on X about fx, a Vercel tool whose

default auto mode runs a safety reviewer on every command. That

reviewer "runs on GPT Luna today", he wrote, and Jev is "likely new

default". Likely, as of that post.

The design is one decision call per agent step. It asks every

decision question in parallel. Your code branches on the answers.

Anything that needs words stays on the language model: tool

arguments, search queries, replies, patches. So does any decision

Jev is not confident about.

The saving comes from calls that already exist only to decide, like

the retry check or the command reviewer. If one LLM call picks a tool

and writes its arguments today, moving the pick to Jev adds a

request. That pays off only when the pick lets you skip the LLM call

or send the arguments to a smaller, cheaper model.

Install the SDK and the

TypeSafe provider:

pnpm add ai@7 @ai-sdk/typesafe-ai
export TYPESAFE_AI_API_KEY=your-key

The code below type-checks against ai 7.0.105 and

@ai-sdk/typesafe-ai 3.0.2, with allowImportingTsExtensions on for

the ./decide.ts import. It has not been run against the live API,

and direct TypeSafe keys are early access. askJev sends all three

questions in one request:

// decide.ts
import { typeSafeAi } from "@ai-sdk/typesafe-ai";
import { experimental_evaluate } from "ai";

export type AgentState = {
  goal: string;
  lastObservation: string;
  pendingCommand: string;
};

export function askJev(s: AgentState) {
  return experimental_evaluate({
    model: typeSafeAi.evaluationModel("jev-latest"),
    state: s,
    questions: {
      nextTool: {
        type: "choice",
        instructions: "Which tool should the agent use next?",
        criteria: {
          search_docs: { includes: ["Product questions"] },
          run_shell: { includes: ["Builds", "Tests"] },
          edit_file: ["Code changes"],
          none: null,
        },
      },
      nextAction: {
        type: "choice",
        instructions: "What should the loop do now?",
        criteria: {
          continue: ["Last step worked, goal not met"],
          retry: ["Last step failed transiently"],
          ask_user: ["Goal is unclear or needs consent"],
          stop: ["Goal is met"],
        },
      },
      risk: {
        type: "score",
        instructions: "How risky is the pending command?",
        criteria: [
          "Read-only",
          "Reversible change",
          "Deletes data or touches production",
        ],
      },
    },
  });
}

state is the agent's current situation as a plain object, and

pendingCommand is whatever the last LLM turn proposed to run. Score

levels are ordered, least severe first, as in the provider docs'

example. A Choice question takes up to 255 options and

a Score question takes 2 to 10 levels. answers.nextTool.choice comes

back typed as the union of your option names, so a misspelled option

in a comparison is a compile error. The branching needs a few types

of its own:

// decide.ts, continued
type Meta = {
  typesafe?: { confidence?: Record<string, number> };
};

// Missing confidence counts as zero: the LLM decides.
const confidence = (m: unknown, id: string): number =>
  (m as Meta | undefined)?.typesafe?.confidence?.[id] ?? 0;

export type Policy = {
  minConfidence: number;
  riskBlockAt: number;
};

export type Step =
  | { kind: "tool"; tool: string }
  | { kind: "retry" }
  | { kind: "ask_user" }
  | { kind: "stop" }
  | { kind: "approve"; command: string }
  | { kind: "llm"; reason: string };

The provider documents confidence at

providerMetadata.typesafe.confidence, for Choice and Score answers

only. The helper reads it defensively. If the value is absent, the

step goes to the LLM instead of guessing.

Then the switch statement, which is now an actual switch:

// decide.ts, continued
export async function decide(
  s: AgentState,
  p: Policy,
): Promise<Step> {
  const r = await askJev(s);
  const { nextTool, nextAction, risk } = r.answers;
  const sure = (id: string) =>
    confidence(r.providerMetadata, id) >= p.minConfidence;

  if (!sure("nextAction")) {
    return { kind: "llm", reason: "next action unclear" };
  }
  switch (nextAction.choice) {
    case "retry":
      return { kind: "retry" };
    case "ask_user":
      return { kind: "ask_user" };
    case "stop":
      return { kind: "stop" };
  }

  if (!sure("nextTool") || nextTool.choice === "none") {
    return { kind: "llm", reason: "tool choice unclear" };
  }
  if (
    nextTool.choice === "run_shell" &&
    (!sure("risk") || risk.score >= p.riskBlockAt)
  ) {
    return { kind: "approve", command: s.pendingCommand };
  }
  return { kind: "tool", tool: nextTool.choice };
}

The next action is checked first, because a loop that should stop

does not need a tool. The shell gate fails closed on uncertainty. A

risky command goes to a human, and so does a command whose risk Jev

is unsure about. A confident wrong answer still runs the command, so

keep your deterministic guards in front of it, such as a command

denylist and a sandbox with no production credentials.

This gate has the most expensive miss in the loop. Run it in shadow

mode next to your current reviewer, and let it decide alone only

after it agrees with your labelled commands.

Both numbers in Policy are inputs on purpose. The SDK checks that

score is the probability-weighted mean of the level indices, so

with the three levels above it runs from 0 to 2. The right cut-off

inside that range comes from your data. Log raw risk.score values

for a few known-safe and known-dangerous commands first. Then pull a

few hundred steps from your logs, label them, run them through

decide, and set the thresholds where the misses stop being

acceptable.

Label the output of decide as a whole, not each question on its

own. Anthony Maio, in a

skeptical read of the launch,

points out that "individually calibrated judgments do not

automatically compose into a calibrated workflow once you run them

through thresholds, weights, and branches". That function is

thresholds and branches.

The loop calls all of this from one place:

// step.ts
import { decide, type AgentState, type Policy, type Step }
  from "./decide.ts";

// Your existing LLM call. It returns the same Step union.
declare function askLlm(
  s: AgentState,
  reason: string,
): Promise<Step>;

export async function nextStep(
  s: AgentState,
  p: Policy,
): Promise<Step> {
  let step: Step;
  try {
    step = await decide(s, p);
  } catch {
    // Provider or network failure, after the SDK's retries.
    return askLlm(s, "decision model unavailable");
  }
  if (step.kind !== "llm") return step;
  return askLlm(s, step.reason);
}

askLlm is the call your loop makes today, trimmed to return the

same Step union, so the loop never needs to know who decided. The

SDK retries 429 and 529 responses up to maxRetries, which defaults

to 2, before it throws. After that the language model takes over, so

an outage or an access limit slows the agent down without stopping

it. Log usage from every evaluate result, and which branch

answered each step. Without that log, every number you have is

TypeSafe's.

On Vercel AI Gateway, the model is a string: model: "typesafe-ai/jev".

Vercel's example also passes

providerOptions: { gateway: { zeroDataRetention: true } }. Confirm

the confidence metadata comes back on that path before you trust the

threshold. If it does not, the helper above sends every step to the

LLM, which is safe and saves you nothing.

The calculator below runs as-is. It treats one agent step's

batch of decisions as one eval case. TypeSafe's cases are workflows

of several questions too, but they are not your questions, so treat

the mapping as rough.

// cost.ts
// Per-case numbers: TypeSafe's aggregate workflow evals.
// Vendor-reported, vendor-run. Swap in your own logs.
type Row = {
  model: string;
  acc: number; // share of cases matching the labels
  usd: number; // cost per case
  sec: number; // latency per case
};

const ROWS: Row[] = [
  { model: "Jev", acc: 0.678, usd: 0.0004, sec: 0.4 },
  { model: "GPT-5.6 Luna", acc: 0.668, usd: 0.0033, sec: 12.9 },
  { model: "GPT-5.6 Terra", acc: 0.679, usd: 0.0304, sec: 10.1 },
  { model: "GPT-5.6 Sol", acc: 0.741, usd: 0.0836, sec: 23.3 },
  { model: "Claude Sonnet 5", acc: 0.678, usd: 0.1174, sec: 78.1 },
  { model: "Claude Opus 5", acc: 0.731, usd: 0.1761, sec: 37.8 },
];

const perDay = Number(process.argv[2] ?? 10_000);
const days = 30;
const jev = ROWS[0];

const fmt = (n: number) => n.toLocaleString("en-US", {
  maximumFractionDigits: 0,
});

console.log(`${fmt(perDay)} cases/day over ${days} days`);
for (const r of ROWS) {
  const month = r.usd * perDay * days;
  const misses = perDay * (1 - r.acc);
  console.log(
    r.model.padEnd(14),
    `$${fmt(month)}/mo`.padStart(12),
    `${(r.usd / jev.usd).toFixed(0)}x`.padStart(5),
    `${fmt(misses)} misses/day`.padStart(17),
    `${r.sec}s`.padStart(6),
  );
}

Run it with Node's type stripping:

$ node --experimental-strip-types cost.ts
10,000 cases/day over 30 days
Jev                 $120/mo    1x  3,220 misses/day   0.4s
GPT-5.6 Luna        $990/mo    8x  3,320 misses/day  12.9s
GPT-5.6 Terra     $9,120/mo   76x  3,210 misses/day  10.1s
GPT-5.6 Sol      $25,080/mo  209x  2,590 misses/day  23.3s
Claude Sonnet 5   $35,220/mo  294x  3,220 misses/day  78.1s
Claude Opus 5    $52,830/mo  440x  2,690 misses/day  37.8s

A "miss" here is a case that disagrees with TypeSafe's reference

labels, nothing more.

At 10,000 cases a day, Jev runs $120 a month. Claude Sonnet 5 runs

$35,220 at the same 67.8%, and Terra runs $9,120 at 67.9%. Sol cuts

misses from 3,220 a day to 2,590, which is 630 fewer, for $24,960

more a month. Whether that trade is worth it depends on what one

miss costs you. A wrong tool pick that the next step corrects costs a

few seconds. A wrong "this command is safe" can cost you a database.

The latency column matters inside a loop. On these per-case figures,

a 30-step run waits about 12 seconds on Jev decisions and about 5

minutes on Terra's.

The answer is text. Tool arguments, a search query, a commit

message, a reply. Jev can choose search_docs. It cannot write the

query.

There are more than 255 options. Picking one file out of a

2,000-file repo is not a Choice question. Narrow the candidates in

code first, or leave it on the LLM.

The state is an image. TypeSafe says Jev works on structured

text state, not images "(yet…)". A computer-use agent deciding from

screenshots keeps its current model.

You have no key yet. Direct TypeSafe API keys are early access.

Vercel's changelog lists Jev as available on AI Gateway, with no

waitlist mentioned. Build the fallback path first either way, because

it is also what runs when the decision call fails.

The workflow is hard and a miss is expensive. The aggregate hides

a spread. On TypeSafe's invoice processing workflow, Jev scored 61.8%

and GPT-5.6 Sol scored 79.1%, a 17.3-point gap. On security incidents

it was 61.7% against Claude Opus 5's 66.2%. On customer service the

gap narrows to 76.0% against Sol's 78.3%. Anything that moves money

or touches production stays on the language model until your own

labels say otherwise.

Search your agent loop for LLM calls whose output you parse into an

enum, a boolean or a number. Each one is a candidate.

Decompose those first, on the model you already use, and measure

agreement against a labelled sample of your own steps. The evals say

that step alone moves accuracy, cost and latency in the right

direction. Then put Jev behind the same questions, with the

confidence gate and the fallback above, and compare cost per case and

agreement on your labels.

Keep the language model on the words, and on the decisions where one

miss costs more than a month of savings.

I wrote AI That Acts for the loop in this post, where a model picks

a tool and your code runs it. It builds tool calling, function

schemas and a first working agent in TypeScript. My AI That Ships

picks up the confidence gates, fallbacks and cost-per-case logging.

If you want the agent patterns without the build, my AI Agents Pocket Guide is the short read. Start with whichever piece your

AI in TypeScript — five books, one path from your first LLM call to agents in production:

Pocket Guides for Developers — short references you can finish in an evening:

Going deeper on tracing and evals: Observability for LLM Applications.

── more in #ai-products 4 stories · sorted by recency
── more on @typesafe ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/your-agent-burns-llm…] indexed:0 read:15min 2026-09-17 ·