Kev: An Open-Source Jev Alternative I Ran Locally A developer ran the open-source decision model Kev locally on Apple Silicon, testing the Kev-0.5B, 0.6B, 0.8B, and 4B checkpoints on a shared support-ticket workload and comparing their probability distributions, uncertainty, and latency. Kev, positioned as an open alternative to the closed Jev model, takes a state plus typed questions and returns structured decisions with probabilities rather than generated text. The developer reports that changing checkpoint size altered confidence, uncertainty, and latency rather than simply increasing confidence, with Kev-8B and Kev-9B left for a follow-up. If you spend enough time around the AI and open-source community, you have probably noticed the noise around Jev. Jev appeared with a somewhat different idea: instead of using an AI model mainly to generate text, what if the model was designed to make fast, typed decisions? That immediately caught my attention. But there was another interesting part of the story. Jev itself wasn't released as an open-source model that developers could simply download, inspect, modify, and run however they wanted. That left a natural question for the open-source community: Can we build something similar ourselves? And not long after, projects started appearing around the same general idea. Small decision models. Open implementations. Different architectures. Different training approaches. Different ways of producing probabilities. Projects such as Laya and Kev are part of that growing conversation, along with several other experiments exploring the broader System One approach. I had already spent time looking at Laya and writing about the idea behind Jev. This time, I wanted to take the same approach with Kev. Not just: “Here is another Jev alternative.” I wanted to actually run it. I wanted to understand what was happening under the hood, install the models locally, send them the same questions, look at the probability distributions, see how the model behaves as the size changes, and find out what actually works on consumer hardware. That is what this article is about. We will start with the architecture and the idea behind Kev, then move into a hands-on experiment with its different checkpoints. And rather than treating all of these models as interchangeable copies of Jev, I'll treat them as what they are: independent open-source attempts at building decision-first AI systems. So let's start with the basic question. I’ve been exploring Jev and the growing open-source ecosystem around decision-first AI, and Kev caught my attention because it takes the idea and makes it runnable and customizable. Instead of generating text, Kev takes a state + typed questions and returns structured decisions with probabilities: State ↓ Typed Questions ↓ Kev ↓ Decision + Probability ↓ Application Logic In this walkthrough, I explored the architecture behind Kev and ran Kev-0.5B, 0.6B, 0.8B, and 4B locally on Apple Silicon using the same support-ticket workload. The results were interesting: changing the checkpoint didn't simply make the model “more confident.” The actual probability distributions, uncertainty, and latency changed across models. There are still more checkpoints to explore, especially Kev-8B and Kev-9B, which I'll test in the next article. The bigger idea: AI doesn't always need to generate text. Sometimes, it just needs to make a decision. | Requirement | Details | | ----------------------- | --------------------------------------------------------- | | Python | 3.12 or 3.13 | | Git | Required to clone the Kev repository | | Virtual Environment | Python venv recommended | | Package Manager | pip | | OS | macOS / Linux | | Apple Silicon | MPS + MLX supported for compatible checkpoints | | NVIDIA GPU | Supported for larger checkpoints | | Models Tested | Kev-0.5B, 0.6B, 0.8B, 4B, 8B, 9B | | Storage | Enough space for the selected model and Qwen base weights | | Internet | Required for the first model download from Hugging Face | Model Link Page: https://huggingface.co/jaredpalmer https://huggingface.co/jaredpalmer GitHub: https://github.com/jaredpalmer/kev https://github.com/jaredpalmer/kev Let's start with the simplest possible explanation. A traditional LLM workflow often looks like this: User input ↓ LLM ↓ Generated text ↓ Parse the text ↓ Application logic Maybe we ask the model: Which department should handle this support ticket? And then tell it to return JSON: { "department": "billing" } That looks structured, but the model is still fundamentally generating tokens. Kev approaches the same problem differently. The application provides: The model then produces a decision and its probability distribution. The high-level flow looks like this: The three main decision types are: noul choice score You can think of them as: noul → yes/no choice → select one option score → assign a level So instead of asking Kev to write a paragraph about a customer ticket, we can ask: Which department should handle this? Does this need urgent human attention? How frustrated is the customer? That is a much smaller interface. Imagine a customer sends this: My package arrived two weeks late, the shoes are the wrong size, and I was charged twice. A conventional LLM could summarize the message and explain what should happen. With Kev, we can define three decisions: Department → returns/shipping/billing Urgency → yes/no Frustration → calm / frustrated / very angry Conceptually, the model can return something like: Department returns 0.47 shipping 0.28 billing 0.25 Urgency yes 0.93 Frustration calm 0.00 frustrated 0.56 very angry 0.44 The application can then decide what to do. That last step is important. The model makes the decision. The application owns the action. That distinction becomes very useful when building production systems. This was probably the first question I had when I started looking at these projects. We already have extremely capable LLMs. So why create a separate model for decisions? The answer is not that LLMs suddenly became useless. It is that generation and decision-making are different interfaces. A general LLM is designed to continue a sequence. A decision model can instead expose the thing an application actually wants: decision probability decision probability decision probability The architecture therefore starts looking less like: and more like: That is the core idea behind the entire experiment. One of the first things you notice when looking at the project is that there are several Kev checkpoints. The current family contains: Kev-0.8B Kev-4B Kev-9B Kev-27B The repository also keeps older checkpoints: Kev-0.5B Kev-0.6B Kev-8B The project documentation distinguishes those older releases from the current model family. Kev-0.5B is the older Qwen2.5-based reference model; the 0.6B and 8B checkpoints belong to the previous Qwen3 generation. The current family has moved to Qwen3.5, while Kev-27B uses Qwen3.8. That gives us a nice little history: For this walkthrough, I want to go through the models up to Kev-9B. That means we can look at: | Model | Generation | What we'll do | | -------- | ---------- | --------------- | | Kev-0.5B | Qwen2.5 | Run and inspect | | Kev-0.6B | Qwen3 | Run and inspect | | Kev-0.8B | Qwen3.5 | Run and inspect | | Kev-4B | Qwen3.5 | Run and test | | Kev-8B | Qwen3 | Run and inspect | | Kev-9B | Qwen3.5 | Run and test | I'm leaving Kev-27B out of the local hands-on part because its hardware requirements are in a completely different category. The repository lists 80 GB-class GPU hardware for that checkpoint. This is probably the most interesting part of the project. According to the repository, Kev checkpoints use a rank-16 LoRA adapter and a pointer head on top of a Qwen base model. The pointer head scores option representations against a decision representation, and a softmax converts those scores into probabilities. A simplified version looks like this: Suppose the pointer head produces these scores: returns 1.72 shipping 1.21 billing 1.08 Softmax turns those scores into something like: returns 0.47 shipping 0.28 billing 0.25 Now our application has something it can reason about directly. It can say: if returns probability 0.80: route to returns else: send to human The model is no longer responsible for deciding what the entire application should do. It provides the signal. This is another detail that makes Kev different from simply asking an LLM three questions in a prompt. Imagine we have: Question 1 → Which department? Question 2 → Is it urgent? Question 3 → How frustrated is the customer? The same state can be reused for those decisions while keeping the questions isolated. Conceptually: The Kev repository describes the questions as sharing the text but not reading one another. For the newer Qwen3.5/Qwen3.8 models, the implementation runs each question as its own row because the underlying Gated DeltaNet layers are recurrent and do not follow attention masks in the same way as attention-only models. The state can still be computed once and reused through caching. That sounds complicated, but the practical idea is simple: One piece of state can feed many independent decisions. The previous generation is worth understanding because it explains how the design evolved. For attention-only Qwen3 bases, the repository describes a sequence structure roughly like: