Kev gives the same answer every time. Until you batch it. A developer writing as Mark Vange of Autom8ly re-ran the open-source decision model Kev-4B (4-bit quantized) over 866 questions and found that answers are bit-identical when requests run alone or in batches of two to three, but 1,223 of 2,315 requests changed once batches reached five or more, with two batched runs at concurrency 8 disagreeing with each other on 403 of 866 questions. The effect on final decisions was small — only 2 of 2,598 answers changed their top choice — and batching 16 requests raised GPU throughput 45% while making each request wait nearly ten times as long. Part 2 of the jev-decision-models http://dev.to/guichard/jev-is-the-best-decision-model-heres-what-to-run-when-you-cant-use-it-mme series. Part 1 http://dev.to/guichard/jev-is-the-best-decision-model-heres-what-to-run-when-you-cant-use-it-mme put seven open-source Jev look-alikes on a single 8 GB GPU. This follow-up tests whether the winner holds its answers across reruns, what concurrent load does to a "deterministic" model, and how sure we are about the headline numbers — plus a new entrant, Julia-1. Mark Vange · Autom8ly · September 2026 · 7 min read Written mostly by AI. Like the first article, this one and the analysis behind it were largely generated by our AI assistant Autumn , directed by me. Every number comes from real runs, and the scripts and predictions are public. Late last week I wrote about running Jev look-alikes on a single 8 GB GPU. The best of them, Kev-4B quantized to 4-bit, scored 0.872 against Jev's 0.968, and its confidence looked trustworthy enough to automate about three quarters of its decisions. Taylor Kolasinski , founder of Poisson Labs, read it and asked the right questions. Their team had just published "The same request twice", a replay study of Jev: send identical requests several times and see whether the answers hold. Often they don't. On items close to a decision threshold, 18.5% of Jev's decisions flipped at a 0.5 or 0.6 cut and 5.3% at 0.9, and because the flips went both ways, the aggregate numbers looked perfectly stable while individual decisions changed underneath. Taylor's questions to us, paraphrased: The first question turned out to have two answers, and the second one is the interesting one. On 28 September we sent Kev-4B 4-bit the same 866 questions three more times, one request at a time, and compared every answer with the original run from 26 September. We checked that nothing was faking this. Our gateway has no response cache, and Kev's server caches only the four most recent inputs, so with 866 different inputs between repeats every answer was computed from scratch. Kev's API rounds probabilities to four decimal places, so "identical" means identical at that resolution. Production traffic doesn't arrive one request at a time. Kev's server collects whatever requests are waiting and runs them through the GPU together as one batch. So we ran the full suite again with 2, 4, 8 and 16 requests sent at the same moment, and once more at 8 to see whether batched runs agree with each other. Grouped by the size of the batch each request actually ran in, the line is sharp. In the first run at 8, none of the 116 requests that ran alone or in pairs changed, while 311 of the 735 that ran in batches of seven did. Across every run, the 2,015 requests that ran alone or in batches of two or three were all bit-identical; in batches of five or more, 1,223 of 2,315 changed. No batch of exactly four happened to form. Two batched runs at 8 disagreed with each other on 403 of 866 questions, because which requests land in the same batch is a matter of timing. So under load, Kev stops being replayable: the same request can come back with slightly different numbers depending on what else was running with it. That is the same kind of behaviour Poisson Labs measured in Jev, at a much smaller scale. The effect on decisions was small. Across the three batched runs, 2 of 2,598 answers changed their top choice, and one of those turned a right answer into a wrong one, which is why one run scores 0.8707. Both flips were on questions where two options were within a hair of each other: on one grammar question the top two options moved from 0.3013 and 0.3042 to 0.3043 and 0.3040. No answer anywhere crossed the 0.9 line. Batching barely paid for itself. Sixteen requests at once got 45% more throughput out of the GPU, but each request waited nearly ten times as long. On this small card, the model is already using the GPU well one request at a time. Why it happens is our reading, not something we traced: GPU arithmetic isn't exactly associative, and a bigger batch can be split and summed in a different order. Kev's server also handles larger batches differently, which would fit the sharp line we saw. If you need to replay a decision for an auditor, keep batches small, or log the probabilities the model actually returned. Determinism on a quiet server isn't determinism in production. We should have explained this the first time. "Automatable at 5% error" doesn't use the 90% threshold. It sorts answers from most to least confident and accepts as many as it can while errors among the accepted stay at or under 5%. The threshold is whatever that turns out to be. Taylor's instinct was still right: we picked that threshold on the same 866 questions we scored it on, which flatters it. Done properly, choosing the threshold on a random half and applying it to the other half, 2,000 times over: The coverage holds up. The error rate is the part to watch: aim for 5% and, on a new sample of about 430 decisions, you will land anywhere from 2% to 8.7%. If 5% is a hard limit, set a lower budget, on your own labelled data. The suite is 866 questions across 49 tasks, one question per case. Resampling the cases gives these 95% intervals: The gap to Jev is real: even at the edge of the interval it is more than seven points. Kev's 0.1% confident-error rate is a single mistake, so the true rate could plausibly be a few times higher. It's still low; it just isn't as precise as a decimal suggests. Supersonic Labs released Julia-1 on 26 September: a 144-million-parameter decision model built on a multilingual ModernBERT encoder, Apache-2.0, small enough to run on a laptop CPU. Its model card reports beating Jev on Supersonic's own typed-decision benchmark 73.2% against 72.7% . We ran it through exactly the same tests as everything else, using the official checkpoint and runtime. On our independent suite, Julia-1 scored 0.450, and 30% of its answers were wrong while it was at least 90% sure. It was deterministic across three reruns and very fast, but its yes/no answers didn't separate at all: its average P yes was 0.36 when the right answer was yes and 0.37 when it was no. It is also sensitive to wording in a way the other models aren't. On a clear-cut refund question it gave P yes of 0.005 as a bare question and 0.997 with plain yes/no descriptions attached. Adding "No" and "Yes" descriptions to every yes/no question in the suite left its accuracy unchanged at 0.449 but flipped 101 individual answers. Before publishing we checked the things that could make this our mistake: the weights file matches the checksum on Supersonic's own results, the input-length setting made no difference, and their README example routes correctly. Our best explanation is that Julia-1 is tuned to the distribution of the benchmark it was trained and measured on. Supersonic call it "the first released test of our training system", and a first release deserves to be judged as one. As published today, we wouldn't use it for decisions that matter. Thanks again to Taylor Kolasinski and Poisson Labs. Their question led to the most useful thing we have learned about these models so far. Next on our list is imajev , an open Jev-compatible model that also accepts images. About this article. This article and the code behind it were mostly generated by AI primarily Claude, by Anthropic, with Autom8ly's internal prompts and skills , directed by me. All Kev runs were made on 26 and 28 September 2026 on one NVIDIA RTX 4060 8 GB , Kev-4B loaded in 4-bit with bitsandbytes, scored with our tool-neutral gutcheck scorer on jabr/classifier-benchmark v2. Julia-1 is revision a85b1273 of SupersonicLabs/Julia-1 weights SHA-256 df853bf7… , run with its own Python runtime on CPU, strict encoding. We are not affiliated with TypeSafe, Poisson Labs, Supersonic Labs or any of the projects tested. Method details. Concurrency runs sent bursts of N simultaneous requests through our gateway, paced to about 55 requests a minute. Batch size per request is inferred from requests in the same burst reporting the same server model time. Near-threshold sets are the questions whose top probability in the original run fell in 0.45, 0.55 or 0.85, 0.95 ; a flip is a change of answer or a move across the cut. Intervals are 95% percentile bootstraps over cases; held-out coverage uses 2,000 random 50/50 splits. For Kev's transfer-v4 suite, where some options have no description, our Julia-1 adapter uses the option's name, as Jev does. Code and data. Everything is at github.com/Autom8ly/gutcheck-bench https://github.com/Autom8ly/gutcheck-bench : the rerun and concurrency clients, every run's predictions and the analysis script in stability/ , the Julia-1 adapter in adapters/julia server.py , and its results alongside every other setup in results/ . python stability/analyze.py reproduces the rerun, batching and interval figures.