# From benchmarking to fine-tuning: what I learned about small decision models

> Source: <https://gokulakrishna.co/2026/09/30/benchmarking-to-fine-tuning-decision-models/>
> Published: 2026-09-30 04:22:35+00:00

I tested two small decision models. First I tested TypeSafe AI’s hosted model, Jev. Then I tested Laya, an independent open-source alternative. I started out benchmarking and ended up fine-tuning.

Here is what I found, in the order I found it.

## Part 1: Benchmarking Jev

TypeSafe AI [launched Jev on September 15, 2026](https://typesafe.ai/blog/introducing-system-one-models-and-jev). It does not write text. You give it a message and a list of possible answers. It tells you how confident it is in each one.

I skipped the launch demos and ran my own evals on public Kaggle datasets:

- **31,369** test samples
- **12** classification tasks
- Three areas: anti-money laundering, customer support intent, and malicious network traffic
- **0.3 s** median response time, calling from Singapore to the US

The main lesson: accuracy depends on the task, and even more on how you set the task up.

### Describe what the data shows, not what the criminal wants

My first anti-money laundering pattern descriptions said what the criminal was trying to achieve. I rewrote them to describe what the transactions look like.

Accuracy went from **65% to 77%**. The model was the same. Only the descriptions changed.

The model sees transactions, not motives. So describe the labels using what’s in the input.

### Remove hidden bias from the setup

In account-level laundering detection, I switched to using the same time window for every account. That removed a hidden bias in the setup. Accuracy went from **63% to 75%**.

### Some tasks have no signal

Not everything worked. Classifying a **single transaction** as laundering or not gave **54%**, which is about chance. One transaction on its own doesn’t carry enough signal. This is a limit of the task, not of the model.

### Confidence you can trust

For customer support intent, Jev picked the right answer **92.3%** of the time. On banking intents it was **79.8%**.

The confidence score is honest. When Jev said it was 90% sure, it was right about 90% of the time.

That makes thresholds useful. If you accept only predictions above 90% confidence, accuracy rises to **97.6%**, and that still covers **82%** of cases. The other 18% go to a fallback: a bigger model, a person, or a request for more information.

### Ask the question directly

To catch off-topic messages, I first used a low top score as the signal. That gave **91.7%**.

Then I added “is this off-topic?” as a separate yes/no question. That gave **97.7%**. It costs no extra time, because everything is answered in one pass.

If you care about a decision, ask for it directly. Don’t infer it from another score.

### The dataset can be wrong too

On malicious network traffic, Jev scored **78%** against the dataset labels as given.

One capture had about **400 windows labeled “benign”**. They were unanswered scans of thousands of hosts. Jev flagged them as malicious with 80% to 90% confidence. Without that capture, the score was **98%**.

I report both numbers. The 98% is a filtered result. But the lesson holds: when the model disagrees with the ground truth, look at the data before you count it as a model error.

## Part 2: Laya, the open-source alternative

Less than a week later, an open-source alternative appeared. [Laya](https://github.com/NandhaKishorM/laya) is a separate project, not a release of Jev’s weights. It claims performance equal to Jev.

I ran it on 11 of the same tasks, using the same harness I built for Jev.

### Accuracy

| Task | Jev | Laya (untuned) | 
|---|---|---|
| Help-desk requests | 92.3% | 49.7% | 
| Banking questions | 79.8% | 39.4% | 
| Money-laundering patterns | 77.3% | 13.5% (no better than guessing) | 

Across 11 tasks and 3 setups, Laya never beat Jev. On binary (yes/no) decisions it came much closer. On choosing between many answers it fell well behind.

### Confidence was the real problem

The accuracy gap was not the biggest issue. The confidence scores were.

- **Help-desk task:** 4,173 of 4,500 answers were at 99% confidence, and about half of them were right.
- **Banking task:** Laya claimed about 97% confidence on almost everything, and got more than half wrong.

Calibration error measures how far confidence is from actual accuracy, where 0 is perfect. Jev scored **0.013**. Laya scored **0.486**.

That breaks confidence-based routing. With Jev, keeping only answers at 90%+ confidence raised accuracy to 97.6%. With Laya, the same filter dropped 7% of questions and gained only 2 points, to 51.7%.

It also breaks refusals. Of the off-topic questions Laya should have refused, it still answered **88%** with 90%+ confidence. Jev did that on **14%**.

My conclusion then: I wouldn’t use untuned Laya to route between many options.

## Part 3: Fine-tuning changed my mind

The next day I fine-tuned Laya on the banking task. It took about **140 minutes** on my laptop.

Accuracy went from **39.4% to 78.3%**, just 1.5 points below Jev’s 79.8%. The task has 77 possible answers, so random guessing would get 1.3%.

The confidence fix mattered more. After fine-tuning, when Laya said it was 97% confident, it was right **96%** of the time.

That is one confidence group, not a full calibration curve. But it’s a big change from “99% sure, right half the time.”

|  | Jev | Laya (fine-tuned) | 
|---|---|---|
| Banking accuracy | 79.8% | 78.3% | 
| Latency | ~309 ms (hosted) | ~90 ms (local) | 
| Inference cost | ~$0.07 per 1,000 requests | $0 (runs on my laptop) | 
| Setup | none | fine-tune per task | 

The catch is that you need to fine-tune for each task, and that takes task-specific data. I have not yet tested off-topic detection on the fine-tuned model.

## Part 4: Laya experts

So I built **Laya experts**: Laya models fine-tuned for specific tasks. Each one is lightweight, runs on a single machine, and makes a decision in about **80 ms**.

One expert detects personally identifiable information (PII). On entity F1, the untuned Laya scored **0.156**, Jev **0.736**, and the Laya PII expert **0.967**. A check like this could run at each stage of a data pipeline to support compliance.

Laya experts is published here: [Laya experts](https://huggingface.co/goku-san/laya-experts).

## What I’d tell someone trying these models

These models look like a real “System 1” decision layer. They handle fast classification, routing and other bounded decisions, so an LLM or agent doesn’t have to call a generative model for everything.

To get good results, spend less time on the model and more on the setup around it:

1. **Descriptions.** Describe labels using what’s visible in the input.
2. **Ask directly.** Make important decisions their own question.
3. **Check confidence.** Test whether “90% sure” really means right 90% of the time. If it doesn’t, thresholds won’t work.
4. **Set thresholds and a fallback.** Decide what happens to the cases the model isn’t sure about.
5. **Inspect disagreements.** Sometimes the label is wrong, not the model.
6. **Fine-tune when needed.** An untuned open-source model can be far off. A short fine-tune on task data can close most of the gap.

*Benchmark note: all figures are my own results on public benchmarks, with thresholds tuned per task. They are not vendor claims. Jev 1.13.0 (hosted). Laya 421M via laya-mlx, run locally on an Apple M3. The untuned Laya evaluation covered 23 runs and 65,078 requests.*
