I tested an open-source alternative to Jev and concluded that I would not use it to route between multiple possible answers. A day later, after fine-tuning it, I had to update that conclusion.
On a banking task, Laya went from 39% accuracy to 78%, compared with 80% for Jev. The more useful change was in its confidence. Before fine-tuning, it was almost always certain, even when it was wrong. Afterwards, its confidence became a much better guide to whether I could trust a decision.
That shift led me to build Laya experts: task-specific versions of Laya for bounded decisions. Here is how I got there, and what I learned from testing both models across banking and cybersecurity.
Results from my banking comparison. These figures describe this task and setup, rather than a general ranking of the models.
Why I started with Jev #
TypeSafe AI launched Jev on 15 September 2026. It is built for structured decisions. In the classification setup I tested, you give it an input and a list of possible answers, and it returns confidence scores rather than generating a prose response.
That is an interesting fit for the small decisions inside a larger workflow: identify a customer’s intent, route a request, or flag a network pattern for inspection.
I skipped the launch demos and ran my own evaluations using public datasets from Kaggle: roughly 31,000 sample records across 12 classification jobs, covering anti-money laundering, customer support intent and malicious network traffic.
Accuracy varied by task. It also varied with how I described the task.
Describe what the model can observe #
In the anti-money laundering evaluation, my first pattern descriptions focused on what a criminal wanted to achieve. I rewrote them to describe what the transactions actually looked like.
Accuracy moved from 65% to 77%.
The distinction matters. Intent is something you infer. Transaction patterns are evidence the model can work with. If the input contains transaction behaviour, the answer descriptions should help the model distinguish that behaviour.
My practical takeaway was to treat the descriptions as part of the system being evaluated. A weak description can make a useful model look worse than it is. Changing the descriptions also changes the experiment, so those changes need to be recorded alongside the results.
Confidence determines how much you can automate #
On customer support intent, Jev returned the right answer 92% of the time. Its confidence also tracked correctness reasonably well: when it said it was about 90% confident, it was right about 90% of the time.
That made a threshold useful. Accepting only predictions above 90% confidence increased accuracy to 97.6% on 82% of cases.
The higher accuracy applies to the accepted subset. The remaining 18% still need another path.
This is the kind of trade-off I care about in a workflow. A model does not have to handle every case by itself to be useful. It needs to handle a meaningful share reliably and give the system a useful signal for when to escalate.
That fallback could be a more capable model, a human reviewer or a request for more information. The right threshold depends on the consequences of an error. The 90% threshold worked as an experiment here; it is not a universal setting.
Inspect disagreements with the dataset #
The network security evaluation produced a different lesson.
Jev’s detection score was 78% using the dataset labels as supplied. One capture contained around 400 windows labelled benign that appeared to be unanswered scans across thousands of hosts. Jev flagged them with confidence scores of roughly 80% to 90%.
Excluding that capture, the score rose to 98%.
I would report both numbers. The 98% result describes a filtered subset, not performance on the original evaluation. It does not establish that every disputed label was wrong.
But the disagreement was worth investigating. Some apparent model errors looked like questionable ground-truth labels. In a security dataset, the most useful next step may be to inspect the traffic behind a disagreement rather than immediately count it as evidence against the model.
Laya disappointed me at first #
Shortly afterwards, I tried Laya, an independent open-source decision model that offers an alternative to Jev. It is not an open-source release of Jev’s weights.
I ran it through the same evaluation harness I had built for Jev. Binary decisions were much closer, but choosing between multiple answers was less convincing.
The bigger problem was overconfidence. In the banking comparison, base Laya reached 39% accuracy while claiming around 97% confidence on almost everything.
That makes confidence-based fallback routing unreliable. If wrong answers are also highly confident, raising the threshold does little to separate cases the model can handle from cases it should pass on.
My initial conclusion was that I would not use that version of Laya as an N-way router: a component choosing between several possible destinations or labels.
Fine-tuning changed the result #
Then I fine-tuned Laya for the banking task. The run took roughly 140 minutes on a laptop.
| Measure | Base Laya | Fine-tuned Laya | Jev |
|---|---|---|---|
| Banking task accuracy | 39% | 78% | 80% |
| Reported latency in this comparison | — | ~90 ms | ~309 ms |
The accuracy gap narrowed to two percentage points. But the confidence behaviour was the part that changed my view most.
After fine-tuning, when Laya reported 97% confidence, it was correct 96% of the time in that group of predictions.
One confidence group is encouraging evidence, but it does not describe calibration across every score or task.
Before fine-tuning, the confidence score gave me little basis for deciding when to trust Laya. Afterwards, it looked much more useful for that decision. That is a substantial change for any system that relies on escalation.
The local model also avoided a per-request API charge. In my comparison, Jev cost approximately $0.07 per 1,000 requests. Local inference still has hardware, electricity and operating costs, and fine-tuning adds its own cost. “No API fee” is the useful distinction here.
The latency figures describe my local and hosted setups, including their different execution paths. They are useful for understanding my workflow, but should not be read as an isolated comparison of model compute speed.
From one fine-tuned model to Laya experts #
These experiments led me to publish Laya experts: versions of Laya fine-tuned for specific tasks, designed to run on a single machine. I observed decisions around 80 ms in the later expert work, separate from the roughly 90 ms banking comparison above.
One expert I trained identifies personally identifiable information (PII). I can see that being useful as a check at different stages of a data workflow, helping route content for appropriate handling or review. Detecting PII is one component of such a workflow; it does not by itself establish compliance.
You can find the release through the Laya experts project link. The important constraint is task specificity. The improvement I saw came from fine-tuning for the task. A banking result does not demonstrate performance on PII detection or IoT cybersecurity. Each expert needs its own evaluation, including a check that its confidence remains useful on the data it will actually encounter.
I plan to cover the fine-tuning process and the IoT cybersecurity expert in a separate technical post.
What I would test before putting one into a workflow #
These experiments made me more interested in a fast “System 1” decision layer: a small component that handles bounded classification and routing decisions inside a larger agent or software system.
For the next implementation, I would focus on five things:
- Observable descriptions. Define the labels using evidence present in the input, and version the descriptions with the evaluation.
- Task-specific performance. Test the actual decision and label set the model will face, including ambiguous examples.
- Confidence and coverage. Measure how often accepted predictions are correct and how much work remains for fallback at each threshold.
- Disputed examples. Review disagreements for model errors, unclear labels and problems in the underlying data.
- An explicit fallback. Decide what happens when confidence is low or the input falls outside the task.
My first Laya result was poor enough that I would not have used it for routing. Fine-tuning changed that assessment. It brought accuracy close to Jev on one banking task and made the confidence scores far more useful.
That is why I am continuing with Laya experts. I want to find out how much useful work a small, task-specific model can take on—and how reliably it can tell the rest of the system when to ask for help.
Benchmark note: These figures are my reported experimental results, not vendor-wide performance claims. The Jev description experiment, support threshold experiment, network evaluation and later banking fine-tuning comparison are separate results. Dataset versions, splits, model versions and full training details should accompany a reproducible technical release.