Jev is the best decision model. Here's what to run when you can't use it. Mark Vange of Autom8ly benchmarked seven open-source alternatives to TypeSafe's Jev decision model on a single 8 GB NVIDIA RTX 4060, adding 4-bit and 8-bit loading via bitsandbytes to three tools that only ship full-precision weights. Using the public jabr/classifier-benchmark v2 suite of 866 synthetic cases across 49 decision tasks, he found at least one local look-alike good enough to build on for high-compliance teams that cannot send data to a hosted API. Jev itself answered in about 135 milliseconds over the internet in his tests. Part 1 of a two-part series. Part 2 continues the story — link to be added once it's published. We put seven open-source Jev look-alikes on a single 8 GB GPU for teams whose data isn't allowed to leave the building. One of them is good enough to build on. Mark Vange · Autom8ly · September 2026 · 9 min read Written mostly by AI. This article and the code behind it were largely generated by an AI assistant Claude , directed by me in response to specific use cases we face in high-compliance environments. Every number was measured on real hardware, and the code is public so you can check it. A few weeks ago TypeSafe released Jev, and it changed how I think about putting language models inside software. Most of what we ask models to do in production isn't writing. It's deciding. Which team should get this ticket? Does this message need a human? Is this claim eligible under our policy? How urgent is this, on a scale of one to three? For years the standard answer was to prompt a general chat model, ask it to reply in JSON, parse whatever came back, and hope it didn't wander off script. Jev skips the text entirely. You send it a piece of state and a set of typed questions pick one, yes/no, or a score , and it sends back probabilities. No parsing, no retries, no invented options. In our tests it answered in about 135 milliseconds, over the internet. The probabilities are the real prize. If a model tells you it is 97% sure a ticket belongs to billing, you can route that ticket automatically and send the 60% cases to a person. That pattern, automate the confident decisions and escalate the rest, is how you put AI into processes where mistakes are expensive. Routing, triage, guardrails, policy checks and agent stop conditions all become cheap, fast function calls. It opens up a lot of new work. I work with teams in high-compliance environments. For many of them, "just call a hosted API" isn't an option. Regulated data may not be allowed to leave a particular network, or a particular country. A new vendor can mean months of security and privacy review. Some systems run on networks with no route to the internet at all. And even where an external call is permitted, auditors want to know exactly which model made each decision, and that it won't change underneath you. When you can't bring your data to the best model, you have to bring a good model to your data. So we went looking for alternatives we could run on our own hardware. Within a week of Jev's release, at least seven open-source look-alikes had appeared. The question was whether any of them were good enough, and whether they would fit on the kind of modest GPU you can actually get approved: a single 8 GB consumer card. All of them copy the same basic idea, and four of them speak Jev's HTTP interface, which made a fair comparison possible. Everything ran on one NVIDIA GeForce RTX 4060 with 8,188 MiB of memory, inside a virtual machine with the GPU passed through. We ran each tool alone, one request at a time, and put every one behind the same /v1/systemone interface so a single scorer could grade them all. For the questions we used an independent public suite: jabr/classifier-benchmark https://github.com/jabr/classifier-benchmark v2, 866 synthetic cases across 49 everyday decision tasks such as support routing, refund and warranty eligibility, phishing and fraud checks, triage, tone and grammar. Its cases were generated and cross-checked by a committee of seven different LLMs, and it is released into the public domain. None of the tools was built around it, and our runs reproduce its author's published scores for Jev, Von and Laya almost exactly. Three of the tools only ship full-precision weights, which don't fit in 8 GB, so we added 4-bit and 8-bit loading to them with bitsandbytes. We also ran Jev itself through TypeSafe's API, and a plain generative model qwen3:8b through Ollama, asked for a JSON answer as the old-fashioned baseline. As a cross-check we ran every setup on a second suite, Kev's own transfer-v4; those results are in the repository and agree on who comes first. Almost everything, once quantized. The only one that didn't fit was Nimble's 9B model, which loads at 7.5 GB before doing any work and runs out of memory on its first request. Peak GPU memory per setup against the 8GB RTX 4060: Peak GPU memory while scoring, sampled every half second. Quantization cost little. Kev-4B scored 0.878 at 8-bit and 0.872 at 4-bit, and Rizzo 0.807 and 0.799; only SemIf lost more, about three points. For a small GPU that trade is easy: Kev-4B needs more than 8 GB in full precision and runs in 3.8 GB at 4-bit. Results — accuracy on 866 questions against median time per question, including a local HTTP hop: | Setup | Accuracy | Wrong at ≥90% conf. | Automatable @5% err | p50 ms | Peak VRAM | |---|---|---|---|---|---| | Jev 1.13 hosted | 0.968 | 0.2% | 100% | 135 | hosted | | Kev-9B · 4-bit | 0.893 | 0.7% | 76% | 223 | 7.4 GB | | Kev-4B · 8-bit | 0.878 | 0.0% | 77% | 385 | 5.3 GB | | Kev-4B · 4-bit | 0.872 | 0.1% | 75% | 188 | 3.8 GB | | SemIf 4B · 8-bit | 0.866 | 3.6% | 74% | 168 | 5.0 GB | | SemIf 4B · 4-bit | 0.837 | 5.0% | 66% | 130 | 3.2 GB | | qwen3:8b via Ollama | 0.827 | n/a | n/a | 207 | 5.9 GB | | Rizzo 4B · q8 | 0.807 | 12.4% | 22% | 71 | 5.7 GB | | Rizzo 4B · q4 | 0.799 | 11.5% | 3% | 74 | 4.1 GB | | Von | 0.724 | 1.7% | 21% | 18 | 3.3 GB | | Kev-0.8B | 0.718 | 0.3% | 14% | 29 | 4.5 GB | | Rizzo 1.7B · q8 | 0.639 | 22.5% | 3% | 35 | 2.5 GB | | Laya | 0.585 | 3.1% | 0% | 21 | 5.5 GB | | NanoJev 0.6B | 0.343 | 0.5% | 0% | 29 | 2.4 GB | | Nimble-9B · 4-bit | did not fit | — | — | — | 7.5 GB idle | Kev-4B at 4-bit is the best all-rounder. It scores 0.872, 9.6 points behind Jev's 0.968, in 3.8 GB at 188 ms per question. Its confidence is trustworthy: only 0.1% of its answers are wrong at 90% confidence or more Jev: 0.2% , so with a 5% error budget you could automate about 75% of its decisions without review. The 8-bit build is a hair more accurate but twice as slow. Kev-9B scores a little higher 0.893 , but it uses 7.4 GB of the 8 GB card, leaving no room for anything else. SemIf is a close second with no training at all. Scoring options straight from an unmodified Qwen3.5-4B reaches 0.866 at 8-bit and 0.837 at 4-bit. That makes it the natural choice if you can't adopt a new model but can run one you already trust. Its confidence needs calibrating first: 3.6–5% of its answers are confidently wrong. Von is the speed pick. It answers in 18 ms in 3.3 GB, ten times faster than Kev-4B, and scores 0.724. It is dependable on tone and routing but drops to coin-flip level on rule-based checks such as warranty eligibility, phishing and suspicious transactions. The rest are harder to recommend today. The generative baseline scores a respectable 0.827 but gives no usable confidence: 17% of its answers are confidently wrong. Rizzo is fast 71 ms and scores about 0.80, but 11–12% of its answers are confidently wrong until you calibrate it, and it struggles with policy rules. Laya 0.585 and Kev-0.8B 0.718 trail Von, and NanoJev's game-trained checkpoint scores near chance. A 9.6-point gap is large, and it widens on the hardest tasks: on grammar checking Jev scores 0.81, while Kev-4B and SemIf score 0.48. But the gap is not just a score. Jev's real advantage shows up when you ask it something new. While writing this article we tried exactly that. Ten different AI models each rewrote the article's opening paragraph to sound more natural, keeping every fact. We added two control paragraphs written as deliberately extreme styles, one stiff corporate prose and one very casual. Then we asked Kev-4B and Jev the same three questions about each paragraph: was it written by a human rather than an AI, how natural does it sound, and which of the ten sounds most human. | Rewrite by | Kev: P human | Jev: P human | Jev: natural | Jev: picked as most human | |---|---|---|---|---| | Fable 5.1 | 0.69 | 0.57 | 0.76 | 12% | | Opus 5.5 | 0.68 | 0.58 | 0.75 | 39% | | Sonnet 5 | 0.70 | 0.60 | 0.76 | 7% | | Opus 4.6 | 0.70 | 0.56 | 0.75 | 5% | | Codex OpenAI | 0.73 | 0.56 | 0.72 | 3% | | Haiku 4.5 | 0.70 | 0.56 | 0.70 | 2% | | hermes3:8b local | 0.71 | 0.45 | 0.67 | 30% | | OpenCode big-pickle | 0.69 | 0.48 | 0.56 | 2% | | qwen3:8b local | 0.69 | 0.47 | 0.43 | 1% | | gemma4:e2b local | 0.70 | 0.49 | 0.40 | 1% | | Control: stiff corporate prose | 0.66 | 0.25 | 0.14 | — | | Control: deliberately casual | 0.68 | 0.59 | 0.84 | — | Sorted by Jev's average rank across its three measures. "Natural" is a 3-level score scaled 0–1. "Picked" is the share of ten rotated pick-one questions; chance is 10%. Kev could not tell. It scored the stiff and casual controls within 0.014 of each other, and its three measures disagreed on the winner. Jev separated the controls sharply 0.25 against 0.59 on the human question, 0.14 against 0.85 on naturalness , and its measures agreed with each other. Its three top picks were all rewrites that kept every fact, and it ranked the two smallest local models' rewrites last on every measure. This is one small experiment with twelve paragraphs, not a benchmark. But it points at a real difference. Kev is very good at the decision shapes it was trained on. Jev appears to generalize further, to questions its builders probably never trained for. If you can use Jev, use Jev. For those of us who can't, the open models are more than a consolation prize, because we control them. Four opportunities stand out. rizzo calibrate command . Calibration doesn't change which answer wins, but it is what makes a confidence threshold something you can defend to an auditor. We aren't the only ones testing this new category, and the independent results agree on the main points: Jev leads, and the small encoders trail the 4B models. What we add is the constraint: a fixed 8 GB memory budget, the larger open models quantized to fit it, and the question a compliance team actually asks, which is how much of this you can safely automate on hardware you control. The scorer, the harness, the quantization patches, the adapters, every result and the humanness experiment are at github.com/Autom8ly/gutcheck-bench https://github.com/Autom8ly/gutcheck-bench . With any /v1/systemone server running, one command scores it: python -m gutcheck.benchmark --endpoint http://127.0.0.1:8008 \ --model kev-latest --suite suites/jabr-v2.jsonl --out runs/my-tool If you run it on different hardware, or on a labelled set from your own domain, I'd like to hear what you find. About this article. This article and the code in the gutcheck-bench repository were mostly generated by AI Claude, by Anthropic , directed by me in response to specific use cases we face in high-compliance environments. The measurements are real: every number comes from runs on the hardware described above, on 25 and 26 September 2026. We are not affiliated with TypeSafe or with any of the projects tested. Our Jev runs cost a few cents in total. How we scored. We wrote a small, tool-neutral scorer gutcheck so that no contestant grades itself. Its metric definitions are based on Kev's open-source scorer Apache-2.0 , and we checked that it reproduces Kev's scores exactly on every run we made with both. Caveats. One GPU, one machine, synthetic decision tasks. Your own decisions may rank these tools differently, so test on a labelled set from your domain before choosing. Jev's latency depends on your distance from TypeSafe's servers. All seven projects were days old when tested and are changing quickly; treat these numbers as a snapshot.