What if one of the most interesting new AI models doesnβt generate text at all?
No chatbot responses.
No code generation.
No essays.
Instead, it is designed to do something much narrower:
Make decisions for software.
Thatβs the idea behind JEV, a new model from TypeSafe AI.
And the numbers TypeSafe AI is reporting immediately caught my attention:
Those are TypeSafe AIβs own workflow comparisons, not universal benchmarks. Different workloads could produce very different results.
But the more interesting question isnβt whether JEV can beat an LLM on every benchmark.
Itβs:
Why are we using generative LLMs for decisions that donβt require generating anything?
That question becomes especially interesting when you start building AI agents.
The easiest way to understand JEV is to compare it with a traditional Large Language Model.
Imagine a customer sends this message:
βMy package was supposed to arrive yesterday, but tracking hasnβt updated.β
You want your system to decide which team should handle it.
Possible options:
billing
shipping
other
Ask an LLM and it might generate something like:
This request should be routed to the shipping department because
the customer is asking about a delayed package.
That works.
But your application doesnβt necessarily need an explanation.
It needs:
shipping
Or, even better, probabilities associated with the available decisions.
Conceptually:
shipping: 0.94
billing: 0.02
other: 0.04
Your software can then decide what happens next.
Thatβs the fundamental difference.
LLMs generate tokens.
JEV evaluates predefined decisions.
The output is primarily intended for software to act on, rather than for a human to read as a conversation.
Modern LLMs are incredibly capable.
They can reason through complex problems, understand large codebases, write software, analyze documents, interact with tools, and orchestrate workflows.
But that flexibility comes with computational cost.
Now consider a typical AI agent.
The agent might need a powerful model to:
But between those large reasoning tasks, agents constantly make smaller decisions.
For example:
Which model should handle this task?
Did the previous step succeed?
Should I retry?
Does this command require approval?
Which tool should run next?
Should this result be escalated?
Today, we often solve these problems withβ¦
another LLM call.
And that works.
But weβre invoking a generative model when the actual output we need might simply be:
retry
or:
needs_approval
use_stronger_model
That is the category of problem TypeSafe AI is targeting with what it calls System One Models.
JEV is its first public model.
Think about the two approaches like this:
| Traditional LLM | JEV |
|---|---|
| Generates text | Evaluates decisions |
| Predicts the next token | Scores predefined options |
| Great for reasoning and generation | Designed for bounded judgments |
| Flexible output | Structured output |
| Human-readable responses | Software-actionable decisions |
| Useful for open-ended problems | Useful for constrained choices |
This distinction matters.
JEV isnβt necessarily trying to become the model inside your AI coding agent that understands an entire repository and writes a feature.
Instead, it could potentially sit beside that model.
A powerful LLM handles:
reasoning
coding
planning
generation
JEV handles:
classification
routing
scoring
triage
bounded decisions
In other words:
The LLM generates. JEV decides.
This is where I think things get particularly interesting.
Consider an agent built using something like Hermes Agent, Claude Code, or another agent framework.
A simplified architecture today might look like:
User
β
Agent
β
LLM
β
Tools
β
LLM
β
Tools
β
LLM
β
Result
The LLM becomes responsible for almost everything.
But imagine separating different responsibilities:
βββ Retrieval / Memory
β
User β Agent βββ JEV β Fast decisions
β
βββ Powerful LLM β Reasoning + Coding
β
βββ Tools β Actions
Now different components handle different types of work.
Thatβs closer to the architecture I would want to experiment with.
Here are three workflows Iβd test first.
This might be the most obvious use case.
Suppose your AI agent has access to multiple models.
Maybe you have:
Fast/Cheap Model
+
Powerful/Expensive Model
A simple task arrives:
Rename this variable across these files.
Do you really need your most expensive reasoning model?
Probably not.
But then another task arrives:
Debug this race condition across multiple distributed services.
Thatβs a very different problem.
Instead of sending everything to the strongest model, JEV could potentially classify the incoming task.
Task
β
JEV
β
βββββββββββββββββββ
β SIMPLE β β Fast/Cheap Model
β COMPLEX β β Powerful Model
β UNCERTAIN β β Powerful Model
βββββββββββββββββββ
This could make model routing much more deliberate.
And importantly, youβd want to measure whether the routing mistakes cost more than the model savings.
Cheap routing isnβt useful if important tasks constantly go to the wrong model.
Thatβs why Iβd want to test:
Cost
Latency
Routing accuracy
Escalation rate
Task success rate
rather than looking only at token pricing.
This one gets even more interesting.
AI agents increasingly interact with real tools.
They can:
Before executing an action, you might want another layer that evaluates what kind of action is being attempted.
Imagine an agent generates:
rm -rf ./build
versus:
rm -rf /
Those commands clearly shouldnβt be treated the same way.
A decision model could evaluate questions such as:
Does this affect production?
Does this modify sensitive files?
Could this expose credentials?
Is this destructive?
Should a human approve this?
The architecture might look like:
AI Agent
β
Proposed Tool Action
β
JEV
β
Risk Classification
β
Policy Engine
β
ALLOW / BLOCK / REQUIRE APPROVAL
Thereβs an important distinction here:
JEV shouldnβt necessarily control permissions.
Your application policy should.
JEV could provide a classification or score, while deterministic code decides what actually happens.
That keeps the security boundary outside the model.
Imagine you have an autonomous agent processing hundreds or thousands of tasks.
Most results might be perfectly normal.
Some wonβt be.
Today, you could send every result through another powerful LLM for verification.
But that potentially means:
1000 tasks
+
1000 verification LLM calls
Instead, a smaller decision model could potentially triage those outputs.
Agent Result
β
JEV
β
βββββββββββββββββ
β Looks Normal β β Finish
β Uncertain β β Stronger LLM
β Suspicious β β Human Review
βββββββββββββββββ
Now your expensive reasoning model only gets involved when necessary.
Again, whether this actually improves the system depends on the accuracy of that triage.
But architecturally, itβs an interesting pattern.
This leads to the bigger idea behind JEV.
Over the last few years, weβve increasingly treated the LLM as the center of everything.
Need classification?
Use an LLM.
Need routing?
Need scoring?
Need validation?
Use another LLM.
Need to judge that LLM?
Use another LLM. π
That architecture works because modern models are incredibly flexible.
But flexibility doesnβt necessarily mean every task should use the same model.
A future AI agent stack could instead look something like:
βββββββββββββββββββββββββββββββ
β Agent Orchestration β
βββββββββββββββββββββββββββββββ€
β β
β Retrieval β Context β
β β
β JEV β Fast Decisions β
β β
β LLM β Reasoning + Coding β
β β
β Tools β Actions β
β β
βββββββββββββββββββββββββββββββ€
β Policy / Control β
βββββββββββββββββββββββββββββββ
Each component has a more specialized responsibility.
The reasoning model doesnβt need to make every tiny routing decision.
The decision model doesnβt need to understand and rewrite an entire codebase.
And tools remain responsible for actually interacting with the outside world.
This is where itβs important to separate an interesting architecture from benchmark hype.
TypeSafe AI reports pricing around:
$0.042 per million input tokens, with no output-token charge.
It also reports latency in the millisecond range.
And on selected company workflows, TypeSafe says JEV achieved results approaching:
200Γ faster
and
400Γ cheaper
than the LLM approaches it compared against.
Those numbers sound impressive.
But they should be interpreted in context.
These arenβt universal statements that JEV is β400Γ better than LLMs.β
Theyβre results from specific workflows and comparisons.
The test Iβd really like to see is much simpler.
Take the same real-world decision workload and run it through:
JEV
vs.
Small LLM
vs.
Frontier LLM
Then compare:
Accuracy
False positives
False negatives
Latency
Cost
Reliability
Because if JEV costs almost nothing but makes significantly more bad routing decisions, the token savings donβt matter.
On the other hand, if it can handle those decisions reliably at a fraction of the latency and costβ¦
then things get interesting very quickly.
And this is the part of JEV that interests me most.
Not whether JEV itself becomes the dominant solution.
Not whether every AI agent suddenly needs it.
And definitely not whether it βreplaces LLMs.β
It probably shouldnβt.
The interesting idea is that not every problem inside an AI agent needs to become another LLM generation.
We might eventually build agent systems where:
Retrieval handles context.
Decision models handle bounded judgments.
Reasoning models handle difficult problems.
Coding models handle software development.
Tools handle actions.
Policy code controls permissions.
Agents orchestrate everything.
Instead of asking one increasingly powerful model to do absolutely everything, we could build systems from specialized intelligence components.
And JEV is an interesting early example of what that architecture could look like.