{"slug": "from-jev-to-laya-experts-what-changed-my-mind-about-small-decision-models", "title": "From Jev to Laya experts: what changed my mind about small decision models", "summary": "Fine-tuning an open-source small decision model called Laya raised its accuracy on a banking classification task from 39% to 78%, approaching the 80% scored by TypeSafe AI's Jev, according to the author's own evaluations across roughly 31,000 Kaggle sample records spanning 12 classification jobs. The author then built \"Laya experts,\" task-specific versions of Laya for bounded decisions, after finding that fine-tuning also made the model's confidence scores a better guide to whether a prediction could be trusted. On customer support intent, Jev answered correctly 92% of the time, and accepting only predictions above 90% confidence raised accuracy to 97.6% on 82% of cases.", "body_md": "I tested an open-source alternative to Jev and concluded that I would not use it to route between multiple possible answers. A day later, after fine-tuning it, I had to update that conclusion.\n\nOn a banking task, Laya went from **39% accuracy to 78%**, compared with **80% for Jev**. The more useful change was in its confidence. Before fine-tuning, it was almost always certain, even when it was wrong. Afterwards, its confidence became a much better guide to whether I could trust a decision.\n\nThat shift led me to build **Laya experts**: task-specific versions of Laya for bounded decisions. Here is how I got there, and what I learned from testing both models across banking and cybersecurity.\n\n*Results from my banking comparison. These figures describe this task and setup, rather than a general ranking of the models.*\n\n## Why I started with Jev\n\nTypeSafe AI [launched Jev on 15 September 2026](https://typesafe.ai/blog/introducing-system-one-models-and-jev). It is built for structured decisions. In the classification setup I tested, you give it an input and a list of possible answers, and it returns confidence scores rather than generating a prose response.\n\nThat is an interesting fit for the small decisions inside a larger workflow: identify a customer’s intent, route a request, or flag a network pattern for inspection.\n\nI skipped the launch demos and ran my own evaluations using public datasets from Kaggle: roughly **31,000 sample records across 12 classification jobs**, covering anti-money laundering, customer support intent and malicious network traffic.\n\nAccuracy varied by task. It also varied with how I described the task.\n\n## Describe what the model can observe\n\nIn the anti-money laundering evaluation, my first pattern descriptions focused on what a criminal wanted to achieve. I rewrote them to describe what the transactions actually looked like.\n\nAccuracy moved from **65% to 77%**.\n\nThe distinction matters. Intent is something you infer. Transaction patterns are evidence the model can work with. If the input contains transaction behaviour, the answer descriptions should help the model distinguish that behaviour.\n\nMy practical takeaway was to treat the descriptions as part of the system being evaluated. A weak description can make a useful model look worse than it is. Changing the descriptions also changes the experiment, so those changes need to be recorded alongside the results.\n\n## Confidence determines how much you can automate\n\nOn customer support intent, Jev returned the right answer **92% of the time**. Its confidence also tracked correctness reasonably well: when it said it was about 90% confident, it was right about 90% of the time.\n\nThat made a threshold useful. Accepting only predictions above 90% confidence increased accuracy to **97.6% on 82% of cases**.\n\n*The higher accuracy applies to the accepted subset. The remaining 18% still need another path.*\n\nThis is the kind of trade-off I care about in a workflow. A model does not have to handle every case by itself to be useful. It needs to handle a meaningful share reliably and give the system a useful signal for when to escalate.\n\nThat fallback could be a more capable model, a human reviewer or a request for more information. The right threshold depends on the consequences of an error. The 90% threshold worked as an experiment here; it is not a universal setting.\n\n## Inspect disagreements with the dataset\n\nThe network security evaluation produced a different lesson.\n\nJev’s detection score was **78% using the dataset labels as supplied**. One capture contained around **400 windows labelled benign** that appeared to be unanswered scans across thousands of hosts. Jev flagged them with confidence scores of roughly 80% to 90%.\n\nExcluding that capture, the score rose to **98%**.\n\nI would report both numbers. The 98% result describes a filtered subset, not performance on the original evaluation. It does not establish that every disputed label was wrong.\n\nBut the disagreement was worth investigating. Some apparent model errors looked like questionable ground-truth labels. In a security dataset, the most useful next step may be to inspect the traffic behind a disagreement rather than immediately count it as evidence against the model.\n\n## Laya disappointed me at first\n\nShortly afterwards, I tried [Laya](https://github.com/NandhaKishorM/laya), an independent open-source decision model that offers an alternative to Jev. It is not an open-source release of Jev’s weights.\n\nI ran it through the same evaluation harness I had built for Jev. Binary decisions were much closer, but choosing between multiple answers was less convincing.\n\nThe bigger problem was overconfidence. In the banking comparison, base Laya reached **39% accuracy** while claiming around **97% confidence on almost everything**.\n\nThat makes confidence-based fallback routing unreliable. If wrong answers are also highly confident, raising the threshold does little to separate cases the model can handle from cases it should pass on.\n\nMy initial conclusion was that I would not use that version of Laya as an N-way router: a component choosing between several possible destinations or labels.\n\n## Fine-tuning changed the result\n\nThen I fine-tuned Laya for the banking task. The run took roughly **140 minutes on a laptop**.\n\n| Measure | Base Laya | Fine-tuned Laya | Jev | \n|---|---|---|---|\n| Banking task accuracy | 39% | 78% | 80% | \n| Reported latency in this comparison | — | ~90 ms | ~309 ms | \n\nThe accuracy gap narrowed to two percentage points. But the confidence behaviour was the part that changed my view most.\n\nAfter fine-tuning, when Laya reported **97% confidence**, it was correct **96% of the time** in that group of predictions.\n\n*One confidence group is encouraging evidence, but it does not describe calibration across every score or task.*\n\nBefore fine-tuning, the confidence score gave me little basis for deciding when to trust Laya. Afterwards, it looked much more useful for that decision. That is a substantial change for any system that relies on escalation.\n\nThe local model also avoided a per-request API charge. In my comparison, Jev cost approximately **$0.07 per 1,000 requests**. Local inference still has hardware, electricity and operating costs, and fine-tuning adds its own cost. “No API fee” is the useful distinction here.\n\nThe latency figures describe my local and hosted setups, including their different execution paths. They are useful for understanding my workflow, but should not be read as an isolated comparison of model compute speed.\n\n## From one fine-tuned model to Laya experts\n\nThese experiments led me to publish **Laya experts**: versions of Laya fine-tuned for specific tasks, designed to run on a single machine. I observed decisions around **80 ms** in the later expert work, separate from the roughly 90 ms banking comparison above.\n\nOne expert I trained identifies **personally identifiable information (PII)**. I can see that being useful as a check at different stages of a data workflow, helping route content for appropriate handling or review. Detecting PII is one component of such a workflow; it does not by itself establish compliance.\n\nYou can find the release through the [Laya experts project link](https://lnkd.in/g6MfzBdE).\n\nThe important constraint is task specificity. The improvement I saw came from fine-tuning for the task. A banking result does not demonstrate performance on PII detection or IoT cybersecurity. Each expert needs its own evaluation, including a check that its confidence remains useful on the data it will actually encounter.\n\nI plan to cover the fine-tuning process and the IoT cybersecurity expert in a separate technical post.\n\n## What I would test before putting one into a workflow\n\nThese experiments made me more interested in a fast “System 1” decision layer: a small component that handles bounded classification and routing decisions inside a larger agent or software system.\n\nFor the next implementation, I would focus on five things:\n\n1. **Observable descriptions.** Define the labels using evidence present in the input, and version the descriptions with the evaluation.\n2. **Task-specific performance.** Test the actual decision and label set the model will face, including ambiguous examples.\n3. **Confidence and coverage.** Measure how often accepted predictions are correct and how much work remains for fallback at each threshold.\n4. **Disputed examples.** Review disagreements for model errors, unclear labels and problems in the underlying data.\n5. **An explicit fallback.** Decide what happens when confidence is low or the input falls outside the task.\n\nMy first Laya result was poor enough that I would not have used it for routing. Fine-tuning changed that assessment. It brought accuracy close to Jev on one banking task and made the confidence scores far more useful.\n\nThat is why I am continuing with Laya experts. I want to find out how much useful work a small, task-specific model can take on—and how reliably it can tell the rest of the system when to ask for help.\n\n*Benchmark note: These figures are my reported experimental results, not vendor-wide performance claims. The Jev description experiment, support threshold experiment, network evaluation and later banking fine-tuning comparison are separate results. Dataset versions, splits, model versions and full training details should accompany a reproducible technical release.*", "url": "https://wpnews.pro/news/from-jev-to-laya-experts-what-changed-my-mind-about-small-decision-models", "canonical_source": "https://gokulakrishna.co/2026/09/30/jev-laya-experts/", "published_at": "2026-09-30 04:22:35+00:00", "updated_at": "2026-09-30 04:49:27.488890+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "ai-tools"], "entities": ["Laya", "Jev", "TypeSafe AI", "Kaggle", "Laya experts"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/from-jev-to-laya-experts-what-changed-my-mind-about-small-decision-models", "markdown": "https://wpnews.pro/news/from-jev-to-laya-experts-what-changed-my-mind-about-small-decision-models.md", "text": "https://wpnews.pro/news/from-jev-to-laya-experts-what-changed-my-mind-about-small-decision-models.txt", "jsonld": "https://wpnews.pro/news/from-jev-to-laya-experts-what-changed-my-mind-about-small-decision-models.jsonld"}}