I tested two small decision models. First I tested TypeSafe AI’s hosted model, Jev. Then I tested Laya, an independent open-source alternative. I started out benchmarking and ended up fine-tuning.
Here is what I found, in the order I found it.
Part 1: Benchmarking Jev #
TypeSafe AI launched Jev on September 15, 2026. It does not write text. You give it a message and a list of possible answers. It tells you how confident it is in each one.
I skipped the launch demos and ran my own evals on public Kaggle datasets:
- 31,369 test samples
- 12 classification tasks
- Three areas: anti-money laundering, customer support intent, and malicious network traffic
- 0.3 s median response time, calling from Singapore to the US
The main lesson: accuracy depends on the task, and even more on how you set the task up.
Describe what the data shows, not what the criminal wants
My first anti-money laundering pattern descriptions said what the criminal was trying to achieve. I rewrote them to describe what the transactions look like.
Accuracy went from 65% to 77%. The model was the same. Only the descriptions changed.
The model sees transactions, not motives. So describe the labels using what’s in the input.
Remove hidden bias from the setup
In account-level laundering detection, I switched to using the same time window for every account. That removed a hidden bias in the setup. Accuracy went from 63% to 75%.
Some tasks have no signal
Not everything worked. Classifying a single transaction as laundering or not gave 54%, which is about chance. One transaction on its own doesn’t carry enough signal. This is a limit of the task, not of the model.
Confidence you can trust
For customer support intent, Jev picked the right answer 92.3% of the time. On banking intents it was 79.8%. The confidence score is honest. When Jev said it was 90% sure, it was right about 90% of the time.
That makes thresholds useful. If you accept only predictions above 90% confidence, accuracy rises to 97.6%, and that still covers 82% of cases. The other 18% go to a fallback: a bigger model, a person, or a request for more information.
Ask the question directly
To catch off-topic messages, I first used a low top score as the signal. That gave 91.7%.
Then I added “is this off-topic?” as a separate yes/no question. That gave 97.7%. It costs no extra time, because everything is answered in one pass.
If you care about a decision, ask for it directly. Don’t infer it from another score.
The dataset can be wrong too
On malicious network traffic, Jev scored 78% against the dataset labels as given.
One capture had about 400 windows labeled “benign”. They were unanswered scans of thousands of hosts. Jev flagged them as malicious with 80% to 90% confidence. Without that capture, the score was 98%.
I report both numbers. The 98% is a filtered result. But the lesson holds: when the model disagrees with the ground truth, look at the data before you count it as a model error.
Part 2: Laya, the open-source alternative #
Less than a week later, an open-source alternative appeared. Laya is a separate project, not a release of Jev’s weights. It claims performance equal to Jev.
I ran it on 11 of the same tasks, using the same harness I built for Jev.
Accuracy
| Task | Jev | Laya (untuned) |
|---|---|---|
| Help-desk requests | 92.3% | 49.7% | | Banking questions | 79.8% | 39.4% | | Money-laundering patterns | 77.3% | 13.5% (no better than guessing) |
Across 11 tasks and 3 setups, Laya never beat Jev. On binary (yes/no) decisions it came much closer. On choosing between many answers it fell well behind.
Confidence was the real problem
The accuracy gap was not the biggest issue. The confidence scores were.
- Help-desk task: 4,173 of 4,500 answers were at 99% confidence, and about half of them were right.
- Banking task: Laya claimed about 97% confidence on almost everything, and got more than half wrong.
Calibration error measures how far confidence is from actual accuracy, where 0 is perfect. Jev scored 0.013. Laya scored 0.486.
That breaks confidence-based routing. With Jev, keeping only answers at 90%+ confidence raised accuracy to 97.6%. With Laya, the same filter dropped 7% of questions and gained only 2 points, to 51.7%.
It also breaks refusals. Of the off-topic questions Laya should have refused, it still answered 88% with 90%+ confidence. Jev did that on 14%.
My conclusion then: I wouldn’t use untuned Laya to route between many options.
Part 3: Fine-tuning changed my mind #
The next day I fine-tuned Laya on the banking task. It took about 140 minutes on my laptop.
Accuracy went from 39.4% to 78.3%, just 1.5 points below Jev’s 79.8%. The task has 77 possible answers, so random guessing would get 1.3%.
The confidence fix mattered more. After fine-tuning, when Laya said it was 97% confident, it was right 96% of the time.
That is one confidence group, not a full calibration curve. But it’s a big change from “99% sure, right half the time.”
| | Jev | Laya (fine-tuned) |
|---|---|---|
| Banking accuracy | 79.8% | 78.3% | | Latency | ~309 ms (hosted) | ~90 ms (local) | | Inference cost | ~$0.07 per 1,000 requests | $0 (runs on my laptop) | | Setup | none | fine-tune per task |
The catch is that you need to fine-tune for each task, and that takes task-specific data. I have not yet tested off-topic detection on the fine-tuned model.
Part 4: Laya experts #
So I built Laya experts: Laya models fine-tuned for specific tasks. Each one is lightweight, runs on a single machine, and makes a decision in about 80 ms.
One expert detects personally identifiable information (PII). On entity F1, the untuned Laya scored 0.156, Jev 0.736, and the Laya PII expert 0.967. A check like this could run at each stage of a data pipeline to support compliance.
Laya experts is published here: Laya experts.
What I’d tell someone trying these models #
These models look like a real “System 1” decision layer. They handle fast classification, routing and other bounded decisions, so an LLM or agent doesn’t have to call a generative model for everything.
To get good results, spend less time on the model and more on the setup around it:
- Descriptions. Describe labels using what’s visible in the input.
- Ask directly. Make important decisions their own question.
- Check confidence. Test whether “90% sure” really means right 90% of the time. If it doesn’t, thresholds won’t work.
- Set thresholds and a fallback. Decide what happens to the cases the model isn’t sure about.
- Inspect disagreements. Sometimes the label is wrong, not the model.
- Fine-tune when needed. An untuned open-source model can be far off. A short fine-tune on task data can close most of the gap.
Benchmark note: all figures are my own results on public benchmarks, with thresholds tuned per task. They are not vendor claims. Jev 1.13.0 (hosted). Laya 421M via laya-mlx, run locally on an Apple M3. The untuned Laya evaluation covered 23 runs and 65,078 requests.