{"slug": "from-benchmarking-to-fine-tuning-what-i-learned-about-small-decision-models", "title": "From benchmarking to fine-tuning: what I learned about small decision models", "summary": "A hands-on evaluation of TypeSafe AI's hosted decision model Jev and the open-source alternative Laya found Jev reached 92.3% accuracy on help-desk intent classification and 79.8% on banking intents, while untuned Laya scored 49.7% and 39.4% on the same tasks and never beat Jev across 11 tasks and 3 setups. The tester's own harness, run on 31,369 test samples across 12 classification tasks, showed Jev's confidence scores were calibrated — predictions accepted above 90% confidence hit 97.6% accuracy while still covering 82% of cases — whereas Laya returned 99% confidence on 4,173 of 4,500 help-desk answers. Rewriting anti-money-laundering label descriptions to describe transaction patterns rather than criminal intent raised accuracy from 65% to 77%, and using a uniform time window per account raised it from 63% to 75%.", "body_md": "I tested two small decision models. First I tested TypeSafe AI’s hosted model, Jev. Then I tested Laya, an independent open-source alternative. I started out benchmarking and ended up fine-tuning.\n\nHere is what I found, in the order I found it.\n\n## Part 1: Benchmarking Jev\n\nTypeSafe AI [launched Jev on September 15, 2026](https://typesafe.ai/blog/introducing-system-one-models-and-jev). It does not write text. You give it a message and a list of possible answers. It tells you how confident it is in each one.\n\nI skipped the launch demos and ran my own evals on public Kaggle datasets:\n\n- **31,369** test samples\n- **12** classification tasks\n- Three areas: anti-money laundering, customer support intent, and malicious network traffic\n- **0.3 s** median response time, calling from Singapore to the US\n\nThe main lesson: accuracy depends on the task, and even more on how you set the task up.\n\n### Describe what the data shows, not what the criminal wants\n\nMy first anti-money laundering pattern descriptions said what the criminal was trying to achieve. I rewrote them to describe what the transactions look like.\n\nAccuracy went from **65% to 77%**. The model was the same. Only the descriptions changed.\n\nThe model sees transactions, not motives. So describe the labels using what’s in the input.\n\n### Remove hidden bias from the setup\n\nIn account-level laundering detection, I switched to using the same time window for every account. That removed a hidden bias in the setup. Accuracy went from **63% to 75%**.\n\n### Some tasks have no signal\n\nNot everything worked. Classifying a **single transaction** as laundering or not gave **54%**, which is about chance. One transaction on its own doesn’t carry enough signal. This is a limit of the task, not of the model.\n\n### Confidence you can trust\n\nFor customer support intent, Jev picked the right answer **92.3%** of the time. On banking intents it was **79.8%**.\n\nThe confidence score is honest. When Jev said it was 90% sure, it was right about 90% of the time.\n\nThat makes thresholds useful. If you accept only predictions above 90% confidence, accuracy rises to **97.6%**, and that still covers **82%** of cases. The other 18% go to a fallback: a bigger model, a person, or a request for more information.\n\n### Ask the question directly\n\nTo catch off-topic messages, I first used a low top score as the signal. That gave **91.7%**.\n\nThen I added “is this off-topic?” as a separate yes/no question. That gave **97.7%**. It costs no extra time, because everything is answered in one pass.\n\nIf you care about a decision, ask for it directly. Don’t infer it from another score.\n\n### The dataset can be wrong too\n\nOn malicious network traffic, Jev scored **78%** against the dataset labels as given.\n\nOne capture had about **400 windows labeled “benign”**. They were unanswered scans of thousands of hosts. Jev flagged them as malicious with 80% to 90% confidence. Without that capture, the score was **98%**.\n\nI report both numbers. The 98% is a filtered result. But the lesson holds: when the model disagrees with the ground truth, look at the data before you count it as a model error.\n\n## Part 2: Laya, the open-source alternative\n\nLess than a week later, an open-source alternative appeared. [Laya](https://github.com/NandhaKishorM/laya) is a separate project, not a release of Jev’s weights. It claims performance equal to Jev.\n\nI ran it on 11 of the same tasks, using the same harness I built for Jev.\n\n### Accuracy\n\n| Task | Jev | Laya (untuned) | \n|---|---|---|\n| Help-desk requests | 92.3% | 49.7% | \n| Banking questions | 79.8% | 39.4% | \n| Money-laundering patterns | 77.3% | 13.5% (no better than guessing) | \n\nAcross 11 tasks and 3 setups, Laya never beat Jev. On binary (yes/no) decisions it came much closer. On choosing between many answers it fell well behind.\n\n### Confidence was the real problem\n\nThe accuracy gap was not the biggest issue. The confidence scores were.\n\n- **Help-desk task:** 4,173 of 4,500 answers were at 99% confidence, and about half of them were right.\n- **Banking task:** Laya claimed about 97% confidence on almost everything, and got more than half wrong.\n\nCalibration error measures how far confidence is from actual accuracy, where 0 is perfect. Jev scored **0.013**. Laya scored **0.486**.\n\nThat breaks confidence-based routing. With Jev, keeping only answers at 90%+ confidence raised accuracy to 97.6%. With Laya, the same filter dropped 7% of questions and gained only 2 points, to 51.7%.\n\nIt also breaks refusals. Of the off-topic questions Laya should have refused, it still answered **88%** with 90%+ confidence. Jev did that on **14%**.\n\nMy conclusion then: I wouldn’t use untuned Laya to route between many options.\n\n## Part 3: Fine-tuning changed my mind\n\nThe next day I fine-tuned Laya on the banking task. It took about **140 minutes** on my laptop.\n\nAccuracy went from **39.4% to 78.3%**, just 1.5 points below Jev’s 79.8%. The task has 77 possible answers, so random guessing would get 1.3%.\n\nThe confidence fix mattered more. After fine-tuning, when Laya said it was 97% confident, it was right **96%** of the time.\n\nThat is one confidence group, not a full calibration curve. But it’s a big change from “99% sure, right half the time.”\n\n|  | Jev | Laya (fine-tuned) | \n|---|---|---|\n| Banking accuracy | 79.8% | 78.3% | \n| Latency | ~309 ms (hosted) | ~90 ms (local) | \n| Inference cost | ~$0.07 per 1,000 requests | $0 (runs on my laptop) | \n| Setup | none | fine-tune per task | \n\nThe catch is that you need to fine-tune for each task, and that takes task-specific data. I have not yet tested off-topic detection on the fine-tuned model.\n\n## Part 4: Laya experts\n\nSo I built **Laya experts**: Laya models fine-tuned for specific tasks. Each one is lightweight, runs on a single machine, and makes a decision in about **80 ms**.\n\nOne expert detects personally identifiable information (PII). On entity F1, the untuned Laya scored **0.156**, Jev **0.736**, and the Laya PII expert **0.967**. A check like this could run at each stage of a data pipeline to support compliance.\n\nLaya experts is published here: [Laya experts](https://huggingface.co/goku-san/laya-experts).\n\n## What I’d tell someone trying these models\n\nThese models look like a real “System 1” decision layer. They handle fast classification, routing and other bounded decisions, so an LLM or agent doesn’t have to call a generative model for everything.\n\nTo get good results, spend less time on the model and more on the setup around it:\n\n1. **Descriptions.** Describe labels using what’s visible in the input.\n2. **Ask directly.** Make important decisions their own question.\n3. **Check confidence.** Test whether “90% sure” really means right 90% of the time. If it doesn’t, thresholds won’t work.\n4. **Set thresholds and a fallback.** Decide what happens to the cases the model isn’t sure about.\n5. **Inspect disagreements.** Sometimes the label is wrong, not the model.\n6. **Fine-tune when needed.** An untuned open-source model can be far off. A short fine-tune on task data can close most of the gap.\n\n*Benchmark note: all figures are my own results on public benchmarks, with thresholds tuned per task. They are not vendor claims. Jev 1.13.0 (hosted). Laya 421M via laya-mlx, run locally on an Apple M3. The untuned Laya evaluation covered 23 runs and 65,078 requests.*", "url": "https://wpnews.pro/news/from-benchmarking-to-fine-tuning-what-i-learned-about-small-decision-models", "canonical_source": "https://gokulakrishna.co/2026/09/30/benchmarking-to-fine-tuning-decision-models/", "published_at": "2026-09-30 04:22:35+00:00", "updated_at": "2026-09-30 05:19:19.115834+00:00", "lang": "en", "topics": ["machine-learning", "ai-products", "ai-tools", "ai-research"], "entities": ["TypeSafe AI", "Jev", "Laya", "Kaggle", "NandhaKishorM"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/from-benchmarking-to-fine-tuning-what-i-learned-about-small-decision-models", "markdown": "https://wpnews.pro/news/from-benchmarking-to-fine-tuning-what-i-learned-about-small-decision-models.md", "text": "https://wpnews.pro/news/from-benchmarking-to-fine-tuning-what-i-learned-about-small-decision-models.txt", "jsonld": "https://wpnews.pro/news/from-benchmarking-to-fine-tuning-what-i-learned-about-small-decision-models.jsonld"}}