{"slug": "kev-an-open-source-jev-alternative-i-ran-locally", "title": "Kev: An Open-Source Jev Alternative I Ran Locally", "summary": "A developer ran the open-source decision model Kev locally on Apple Silicon, testing the Kev-0.5B, 0.6B, 0.8B, and 4B checkpoints on a shared support-ticket workload and comparing their probability distributions, uncertainty, and latency. Kev, positioned as an open alternative to the closed Jev model, takes a state plus typed questions and returns structured decisions with probabilities rather than generated text. The developer reports that changing checkpoint size altered confidence, uncertainty, and latency rather than simply increasing confidence, with Kev-8B and Kev-9B left for a follow-up.", "body_md": "If you spend enough time around the AI and open-source community, you have probably noticed the noise around Jev.\n\nJev appeared with a somewhat different idea: instead of using an AI model mainly to generate text, what if the model was designed to make fast, typed decisions?\n\nThat immediately caught my attention.\n\nBut there was another interesting part of the story.\n\nJev itself wasn't released as an open-source model that developers could simply download, inspect, modify, and run however they wanted. That left a natural question for the open-source community:\n\nCan we build something similar ourselves?\n\nAnd not long after, projects started appearing around the same general idea.\n\nSmall decision models.\n\nOpen implementations.\n\nDifferent architectures.\n\nDifferent training approaches.\n\nDifferent ways of producing probabilities.\n\nProjects such as Laya and Kev are part of that growing conversation, along with several other experiments exploring the broader System One approach.\n\nI had already spent time looking at Laya and writing about the idea behind Jev. This time, I wanted to take the same approach with Kev.\n\nNot just:\n\n“Here is another Jev alternative.”\n\nI wanted to actually run it.\n\nI wanted to understand what was happening under the hood, install the models locally, send them the same questions, look at the probability distributions, see how the model behaves as the size changes, and find out what actually works on consumer hardware.\n\nThat is what this article is about.\n\nWe will start with the architecture and the idea behind Kev, then move into a hands-on experiment with its different checkpoints.\n\nAnd rather than treating all of these models as interchangeable copies of Jev, I'll treat them as what they are: independent open-source attempts at building decision-first AI systems.\n\nSo let's start with the basic question.\n\nI’ve been exploring Jev and the growing open-source ecosystem around decision-first AI, and Kev caught my attention because it takes the idea and makes it runnable and customizable.\n\nInstead of generating text, Kev takes a state + typed questions and returns structured decisions with probabilities:\n\n```\nState\n  ↓\nTyped Questions\n  ↓\nKev\n  ↓\nDecision + Probability\n  ↓\nApplication Logic\n```\n\nIn this walkthrough, I explored the architecture behind Kev and ran Kev-0.5B, 0.6B, 0.8B, and 4B locally on Apple Silicon using the same support-ticket workload.\n\nThe results were interesting: changing the checkpoint didn't simply make the model “more confident.” The actual probability distributions, uncertainty, and latency changed across models.\n\nThere are still more checkpoints to explore, especially Kev-8B and Kev-9B, which I'll test in the next article.\n\nThe bigger idea: AI doesn't always need to generate text. Sometimes, it just needs to make a decision.\n\n```\n| Requirement             | Details                                                   |\n| ----------------------- | --------------------------------------------------------- |\n| **Python**              | 3.12 or 3.13                                              |\n| **Git**                 | Required to clone the Kev repository                      |\n| **Virtual Environment** | Python `venv` recommended                                 |\n| **Package Manager**     | `pip`                                                     |\n| **OS**                  | macOS / Linux                                             |\n| **Apple Silicon**       | MPS + MLX supported for compatible checkpoints            |\n| **NVIDIA GPU**          | Supported for larger checkpoints                          |\n| **Models Tested**       | Kev-0.5B, 0.6B, 0.8B, 4B, 8B, 9B                          |\n| **Storage**             | Enough space for the selected model and Qwen base weights |\n| **Internet**            | Required for the first model download from Hugging Face   |\n```\n\nModel Link Page: [https://huggingface.co/jaredpalmer](https://huggingface.co/jaredpalmer)\n\nGitHub: [https://github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev)\n\nLet's start with the simplest possible explanation.\n\nA traditional LLM workflow often looks like this:\n\n```\nUser input\n    ↓\nLLM\n    ↓\nGenerated text\n    ↓\nParse the text\n    ↓\nApplication logic\n```\n\nMaybe we ask the model:\n\nWhich department should handle this support ticket?\n\nAnd then tell it to return JSON:\n\n```\n{\n  \"department\": \"billing\"\n}\n```\n\nThat looks structured, but the model is still fundamentally generating tokens.\n\nKev approaches the same problem differently.\n\nThe application provides:\n\nThe model then produces a decision and its probability distribution.\n\nThe high-level flow looks like this:\n\nThe three main decision types are:\n\n```\nnoul\nchoice\nscore\n```\n\nYou can think of them as:\n\n```\nnoul   → yes/no\nchoice → select one option\nscore  → assign a level\n```\n\nSo instead of asking Kev to write a paragraph about a customer ticket, we can ask:\n\nWhich department should handle this?\n\nDoes this need urgent human attention?\n\nHow frustrated is the customer?\n\nThat is a much smaller interface.\n\nImagine a customer sends this:\n\n`My package arrived two weeks late, the shoes are the wrong size, and I was charged twice.`\n\nA conventional LLM could summarize the message and explain what should happen.\n\nWith Kev, we can define three decisions:\n\n```\nDepartment\n→ returns/shipping/billing\n\nUrgency\n→ yes/no\n\nFrustration\n→ calm / frustrated / very angry\n```\n\nConceptually, the model can return something like:\n\nDepartment\n\n```\nreturns     0.47\nshipping    0.28\nbilling     0.25\n\nUrgency\n\nyes         0.93\n\nFrustration\n\ncalm        0.00\nfrustrated  0.56\nvery angry  0.44\n```\n\nThe application can then decide what to do.\n\nThat last step is important.\n\nThe model makes the decision. The application owns the action.\n\nThat distinction becomes very useful when building production systems.\n\nThis was probably the first question I had when I started looking at these projects.\n\nWe already have extremely capable LLMs.\n\nSo why create a separate model for decisions?\n\nThe answer is not that LLMs suddenly became useless.\n\nIt is that generation and decision-making are different interfaces.\n\nA general LLM is designed to continue a sequence.\n\nA decision model can instead expose the thing an application actually wants:\n\n```\ndecision\nprobability\ndecision\nprobability\ndecision\nprobability\n```\n\nThe architecture therefore starts looking less like:\n\nand more like:\n\nThat is the core idea behind the entire experiment.\n\nOne of the first things you notice when looking at the project is that there are several Kev checkpoints.\n\nThe current family contains:\n\n```\nKev-0.8B\nKev-4B\nKev-9B\nKev-27B\n```\n\nThe repository also keeps older checkpoints:\n\n```\nKev-0.5B\nKev-0.6B\nKev-8B\n```\n\nThe project documentation distinguishes those older releases from the current model family. Kev-0.5B is the older Qwen2.5-based reference model; the 0.6B and 8B checkpoints belong to the previous Qwen3 generation. The current family has moved to Qwen3.5, while Kev-27B uses Qwen3.8.\n\nThat gives us a nice little history:\n\nFor this walkthrough, I want to go through the models up to Kev-9B.\n\nThat means we can look at:\n\n```\n| Model    | Generation | What we'll do   |\n| -------- | ---------- | --------------- |\n| Kev-0.5B | Qwen2.5    | Run and inspect |\n| Kev-0.6B | Qwen3      | Run and inspect |\n| Kev-0.8B | Qwen3.5    | Run and inspect |\n| Kev-4B   | Qwen3.5    | Run and test    |\n| Kev-8B   | Qwen3      | Run and inspect |\n| Kev-9B   | Qwen3.5    | Run and test    |\n```\n\nI'm leaving Kev-27B out of the local hands-on part because its hardware requirements are in a completely different category. The repository lists 80 GB-class GPU hardware for that checkpoint.\n\nThis is probably the most interesting part of the project.\n\nAccording to the repository, Kev checkpoints use a rank-16 LoRA adapter and a pointer head on top of a Qwen base model. The pointer head scores option representations against a decision representation, and a softmax converts those scores into probabilities.\n\nA simplified version looks like this:\n\nSuppose the pointer head produces these scores:\n\n```\nreturns     1.72\nshipping    1.21\nbilling     1.08\n```\n\nSoftmax turns those scores into something like:\n\n```\nreturns     0.47\nshipping    0.28\nbilling     0.25\n```\n\nNow our application has something it can reason about directly.\n\nIt can say:\n\n```\nif returns_probability > 0.80:\n    route_to_returns()\nelse:\n    send_to_human()\n```\n\nThe model is no longer responsible for deciding what the entire application should do.\n\nIt provides the signal.\n\nThis is another detail that makes Kev different from simply asking an LLM three questions in a prompt.\n\nImagine we have:\n\n```\nQuestion 1 → Which department?\nQuestion 2 → Is it urgent?\nQuestion 3 → How frustrated is the customer?\n```\n\nThe same state can be reused for those decisions while keeping the questions isolated.\n\nConceptually:\n\nThe Kev repository describes the questions as sharing the text but not reading one another. For the newer Qwen3.5/Qwen3.8 models, the implementation runs each question as its own row because the underlying Gated DeltaNet layers are recurrent and do not follow attention masks in the same way as attention-only models. The state can still be computed once and reused through caching.\n\nThat sounds complicated, but the practical idea is simple:\n\nOne piece of state can feed many independent decisions.\n\nThe previous generation is worth understanding because it explains how the design evolved.\n\nFor attention-only Qwen3 bases, the repository describes a sequence structure roughly like:\n\n```\n<state> ...state...\n\n<q> instructions\n    <opt> option 1 </opt>\n    <opt> option 2 </opt>\n    <opt> option 3 </opt>\n    <decide>\n\n<q> instructions\n    <opt> option 1 </opt>\n    <opt> option 2 </opt>\n    <opt> option 3 </opt>\n    <decide>\n```\n\nThe attention mask prevents one question from reading another.\n\nWe can visualize that as:\n\nFor the current Qwen3.5/Qwen3.8-based models, the implementation takes a different route because of the recurrent components of those base models.\n\nSo there is an important distinction:\n\n```\nOlder Qwen3 Kev\n→ attention masking\n\nCurrent Qwen3.5 / Qwen3.8 Kev\n→ independent rows + shared cached state\n```\n\nThat distinction comes directly from the current project implementation and is worth preserving rather than collapsing all Kev versions into one generic architecture description.\n\nNow we can zoom in one level further.\n\nThe pointer head is the piece that turns the model's representations into option scores.\n\nVery roughly:\n\nThis is why the output is naturally suited to a choice question.\n\nInstead of asking the model to generate:\n\n`The most likely department is returns because...`\n\nwe can compare the available options and derive a probability distribution over them.\n\nThe same overall idea is used for noul and score, with their own output interpretation.\n\nLet's make this concrete.\n\nA binary decision.\n\n```\nIs this customer asking for a refund?\n\nyes → 0.91\n```\n\nThe value represents the probability of the positive outcome.\n\nChoose one item from several options.\n\n```\nWhich department should handle this?\n\nreturns     0.47\nshipping    0.28\nbilling     0.25\n```\n\nKev returns the selected choice along with the distribution.\n\nInstead of picking one label, the model can place probability across ordered levels.\n\nFor example:\n\n```\nHow frustrated is the customer?\n\n0 → Calm\n1 → Frustrated\n2 → Very angry\n```\n\nThe result can contain:\n\n```\n0 → 0.00\n1 → 0.56\n2 → 0.44\n```\n\nand the resulting score can be represented as a weighted level.\n\nThis is useful when the distinction is not simply yes/no or category A/category B.\n\nThis is one of the details I care about most.\n\nA normal classifier might simply return:\n\n```\nbilling\n```\n\nThat gives the application very little information about uncertainty.\n\nKev can return something closer to:\n\n```\nbilling    0.51\nreturns    0.29\nshipping   0.20\n```\n\nNow the application can make its own decision.\n\nBut there is an important caveat here.\n\nProbability is not the same thing as correctness.\n\nA model can be highly confident and still be wrong.\n\nThat is why the Kev repository includes calibration, and why I would not build a production workflow that blindly says:\n\n```\nif confidence > 0.9:\n    trust_the_model()\n```\n\nA threshold needs to be validated against the actual workload.\n\nNow we can finally get away from the architecture diagrams and run Kev.\n\nThe current project recommends Python 3.12 or 3.13 and uv. The repository's .python-version currently points the normal uv workflow toward Python 3.13.\n\nI prefer starting with a clean checkout rather than mixing it into another Python environment:\n\n```\ngit clone https://github.com/jaredpalmer/kev.git\ncd kev\n```\n\nI'd use Python 3.13:\n\n```\npython3.13 --version\n\nYou should see something like:\n\nPython 3.13.x\npython3.13 -m venv .venv\n\nThis creates:\n\nkev/\n├── .venv/\n├── kev/\n├── README.md\n├── pyproject.toml\n└── ...\nsource .venv/bin/activate\n```\n\nYour terminal should now look roughly like:\n\n```\n(.venv) (base) ayushkumar@Ayushs-Mac-mini-2 kev %\npython --version\nwhich python\n```\n\nThe second command should point inside your project:\n\n```\n.../kev/.venv/bin/python\n```\n\nSince Kev defines its serving dependencies as a serve extra, we can install the project directly into this virtual environment:\n\n```\npython -m pip install --upgrade pip\npip install -e \".[serve]\"\n```\n\nThe serve extra includes FastAPI, the TypeSafe SDK, Uvicorn, and the Apple Silicon MLX backend when running on an ARM Mac.\n\n``` python\npython -c \"import kev; print('Kev installed successfully')\"\n```\n\nThen check the package:\n\n```\npip show kev\npython -m kev.serve \\\n  --run jaredpalmer/kev-0.8b \\\n  --port 8009\n```\n\nThen, in another terminal:\n\n```\ncurl http://localhost:8009/v1/models\n```\n\nAt this point, Kev was installed, and the local server was up.\n\nThe next thing I wanted to verify was not whether the process was alive, but which model and backend were actually being used.\n\nI ran:\n\n```\ncurl http://localhost:8009/v1/models\n```\n\nThe response showed that my Mac was running:\n\n```\nModel:       jaredpalmer/kev-0.8b\nBase:        Qwen/Qwen3.5-0.8B-Base\nLoRA rank:   16\nDevice:      MPS\nBackend:     MLX\nDtype:       bfloat16\nTemperature: 2.351...\n```\n\nSo the local setup looked like this:\n\nThe part I found interesting here was the backend.\n\nKev was not running through CUDA on my machine. Since this is an Apple Silicon Mac, it was using the MLX backend with MPS and bfloat16.\n\nThere was also a jev-latest entry in the response.\n\nThat does not mean TypeSafe's hosted Jev was running locally. In this setup, that name points to the same local Kev checkpoint. So for the rest of the article, I'll use kev-latest to avoid confusing the two.\n\nThe server was ready.\n\nNow it was time to actually ask Kev a question.\n\nUntil now, we had only verified that the model loaded.\n\nA model running successfully is one thing.\n\nGetting it to make a useful decision is another.\n\nSo I created a small support-ticket example.\n\nHere is the situation:\n\n`“My laptop was delivered three days late, the screen is damaged, and I was charged twice.”`\n\nThere are several different pieces of information in that single sentence.\n\nInstead of asking an LLM to explain the problem, I want Kev to answer three specific questions:\n\n```\n1. Which team should handle this issue?\n2. Does this require urgent human attention?\n3. How serious is the issue?\n```\n\nThat maps directly to Kev's three decision primitives:\n\n```\nchoice → department\nnoul   → urgency\nscore  → severity\n```\n\nSo the request becomes:\n\nThis is where the API becomes interesting.\n\nThe server was running, so the next step was simple: give Kev an actual problem to solve.\n\nI used a support-ticket scenario with three different types of questions:\n\n```\nState:\n\"My laptop was delivered three days late,\nthe screen is damaged, and I was charged twice.\"\n```\n\nThen I asked Kev:\n\n```\n1. Which team should handle this issue?\n2. Does this require urgent human attention?\n3. How serious is this customer issue?\n```\n\nThat gives us one `choice`, one `noul`, and one `score` question in the same request.\n\nHere is the request I sent:\n\n```\ncurl -s http://localhost:8009/v1/systemone \\\n  -H 'content-type: application/json' \\\n  -d '{\n    \"state\": \"My laptop was delivered three days late, the screen is damaged, and I was charged twice.\",\n    \"model\": \"kev-latest\",\n    \"questions\": {\n      \"department\": {\n        \"type\": \"choice\",\n        \"instructions\": \"Which team should handle this issue?\",\n        \"criteria\": {\n          \"support\": \"Hardware problems, damaged devices, technical issues\",\n          \"shipping\": \"Delivery delays, tracking, lost packages\",\n          \"billing\": \"Charges, invoices, payment problems\"\n        }\n      },\n      \"urgent\": {\n        \"type\": \"noul\",\n        \"instructions\": \"Does this require urgent human attention?\"\n      },\n      \"severity\": {\n        \"type\": \"score\",\n        \"instructions\": \"How serious is this customer issue?\",\n        \"criteria\": [\n          \"Low\",\n          \"Medium\",\n          \"High\"\n        ]\n      }\n    }\n  }'\n```\n\nAnd this time, instead of using a made-up response, let's look at what my local Kev-0.8B instance actually returned.\n\n```\n{\n  \"model\": \"kev-latest\",\n  \"answers\": {\n    \"department\": {\n      \"type\": \"choice\",\n      \"choice\": \"support\",\n      \"confidence\": 0.5144,\n      \"probabilities\": {\n        \"support\": 0.6763,\n        \"shipping\": 0.1844,\n        \"billing\": 0.1393\n      }\n    },\n    \"urgent\": {\n      \"type\": \"noul\",\n      \"noul\": 0.6559\n    },\n    \"severity\": {\n      \"type\": \"score\",\n      \"score\": 1.3554,\n      \"legend\": {\n        \"0\": \"Low\",\n        \"1\": \"Medium\",\n        \"2\": \"High\"\n      },\n      \"probabilities\": {\n        \"0\": 0.1693,\n        \"1\": 0.306,\n        \"2\": 0.5247\n      },\n      \"confidence\": 0.0332\n    }\n  },\n  \"usage\": {\n    \"input_tokens\": 95,\n    \"output_tokens\": 173\n  },\n  \"latency_ms\": 1520.1\n}\n```\n\nThis one request produced three different kinds of outputs:\n\nThis is a nice demonstration of why the System One-style interface is interesting.\n\nOne piece of state can produce several independent decision signals in one request.\n\nThe response also included:\n\n```\nlatency_ms: 1520.1\n```\n\nSo this particular request took approximately:\n\n1.52 seconds\n\naccording to the API's reported latency.\n\nThe request contained:\n\n```\n95 input tokens\n173 output tokens\n```\n\nThere is an important detail here, though.\n\nThis is one local run, not a benchmark.\n\nA single request doesn't tell us what the normal p50 or p95 latency looks like, and it certainly isn't enough to compare Kev with another model.\n\nWe'll collect more measurements later.\n\nFor now, this number simply tells us what happened during this particular experiment.\n\nThe most interesting thing about this first test wasn't actually the support result.\n\nIt was the fact that one model call gave the application three different signals:\n\n```\nDepartment → support\nUrgency    → 0.6559\nSeverity   → 1.3554\n```\n\nA conventional LLM could certainly produce the same information, but we'd normally be asking it to generate some structured response and then parse that response.\n\nHere, the interface is built around the decisions themselves.\n\nThat changes the programming model.\n\nInstead of:\n\n`\"Tell me what you think.\"`\n\nwe are effectively asking:\n\n`\"Answer these specific decisions.\"`\n\nAnd that is much closer to how normal application logic works.\n\nNow let's make the experiment slightly more practical.\n\nSuppose our application has these rules:\n\n```\nHigh urgency\n→ human escalation\n\nStrong support classification\n→ technical support\n\nEverything uncertain\n→ manual review\n```\n\nWe can represent that with a simple flow:\n\nThis is where the separation between the model and the application becomes important.\n\nKev doesn't need to know what your business process is.\n\nIt provides the decision signals.\n\nYour software decides whether those signals are enough to trigger an action.\n\nOur first experiment also reinforces a rule that is easy to forget.\n\nA probability is not a guarantee.\n\nIn this run:\n\n`support = 0.6763`\n\ndoesn't mean:\n\n`“The model is 67.63% correct.”`\n\nAnd:\n\n`urgent = 0.6559`\n\ndoesn't mean:\n\n`“There is definitely a 65.59% chance that a human should intervene.”`\n\nThese values describe the model's output distribution. Whether a particular probability threshold is actually useful needs to be evaluated against the real workload.\n\nThis becomes especially important once we start automating actions.\n\nKev-0.8B has now passed the first basic test.\n\nBut one model isn't enough for what I want to explore.\n\nThe repository contains several generations and sizes, and I want to see how the same decision workload behaves across them.\n\nSo next I'm going to run the same experiment with:\n\n```\nKev-0.5B\nKev-0.6B\nKev-0.8B\nKev-4B\nKev-8B\nKev-9B\n```\n\nWe'll keep the state, questions, and criteria the same.\n\nThat gives us a much cleaner experiment:\n\nKev-0.5B is the earliest checkpoint in this family and is based on Qwen2.5-0.5B.\n\nFor this test, I used the same support-ticket request from the previous section.\n\nStart it with:\n\n```\npython -m kev.serve \\\n  --run jaredpalmer/kev-0.5b \\\n  --port 8009\n```\n\nThen verify:\n\n```\ncurl http://localhost:8009/v1/models\n```\n\nOnce the server is up, send the same /v1/systemone request.\n\nThe local server loaded it successfully on my Mac:\n\n```\nModel:       jaredpalmer/kev-0.5b\nBase:        Qwen/Qwen2.5-0.5B\nDevice:      MPS\nBackend:     PyTorch\nDtype:       bfloat16\nLoRA rank:   16\nTemperature: 1.0\n```\n\nThe /v1/models endpoint also reports a prefix cache of up to 4 states, with caching enabled for states of at least 384 tokens.\n\nOne small detail worth noting: kev-latest and jev-latest both point to the same local Kev-0.5B checkpoint. The latter is simply the server's compatibility alias; it is not TypeSafe's hosted Jev.\n\nAt this stage, the important thing is that the 0.5B model runs locally on Apple Silicon, giving us a lightweight baseline before moving to the newer checkpoints.\n\nNext, I moved to Kev-0.6B, an older-generation Kev checkpoint built on Qwen3-0.6B-Base.\n\nRun:\n\n```\npython -m kev.serve \\\n  --run jaredpalmer/kev-0.6b \\\n  --port 8009\n```\n\nThen:\n\n```\ncurl http://localhost:8009/v1/models\n```\n\nThe model loaded successfully on my Mac:\n\n```\nModel:       jaredpalmer/kev-0.6b\nBase:        Qwen/Qwen3-0.6B-Base\nDevice:      MPS\nBackend:     PyTorch\nDtype:       bfloat16\nLoRA rank:   16\nTemperature: 1.0\n```\n\nSo the main change from Kev-0.5B is the underlying Qwen generation:\n\n```\nKev-0.5B\nQwen2.5-0.5B\n\n        ↓\n\nKev-0.6B\nQwen3-0.6B\n```\n\nThe server again exposes the same kev-latest and jev-latest compatibility names, both pointing to the local Kev-0.6B checkpoint.\n\nAt this stage, the setup was working exactly as expected. The next step was to send the same support-ticket request we used for Kev-0.8B and see whether moving from Qwen2.5 to Qwen3 changes the actual decisions or probability distribution.\n\nQuick comparison so far\n\n```\nModel   Base    Backend Device  Dtype\nKev-0.5B    Qwen2.5-0.5B    PyTorch MPS BF16\nKev-0.6B    Qwen3-0.6B  PyTorch MPS BF16\nKev-0.8B    Qwen3.5-0.8B    MLX MPS BF16\n```\n\nAfter testing the smaller checkpoints, I moved to Kev-4B, one of the current-generation models in the Kev family.\n\nKev-4B is built on Qwen3.5-4B-Base and, on my Apple Silicon machine, it loaded through the MLX backend.\n\nThe /v1/models response showed:\n\n```\nModel:       jaredpalmer/kev-4b\nBase:        Qwen/Qwen3.5-4B-Base\nLoRA rank:   16\nDevice:      MPS\nBackend:     MLX\nDtype:       bfloat16\nTemperature: 2.406...\n```\n\nThe response came back with:\n\n```\n{\n  \"department\": {\n    \"choice\": \"support\",\n    \"confidence\": 0.1233,\n    \"probabilities\": {\n      \"support\": 0.4155,\n      \"shipping\": 0.1831,\n      \"billing\": 0.4014\n    }\n  },\n  \"urgent\": {\n    \"noul\": 0.5069\n  },\n  \"severity\": {\n    \"score\": 1.7751,\n    \"probabilities\": {\n      \"0\": 0.0402,\n      \"1\": 0.1445,\n      \"2\": 0.8153\n    },\n    \"confidence\": 0.6627\n  }\n}\n```\n\nThe model reported:\n\n```\nInput tokens:  95\nOutput tokens: 174\nLatency:       21549.1 ms\n```\n\nThe interesting part was not simply that Kev-4B selected support.\n\nThe probability distribution was much closer between two options:\n\n```\nsupport     0.4155\nbilling     0.4014\nshipping    0.1831\n```\n\nSo although support was the selected option, the model was not particularly decisive about the department.\n\nThat is actually useful information.\n\nThe ticket contains three different issues:\n\n```\nDamaged laptop → support\nLate delivery  → shipping\nDouble charge  → billing\n```\n\nKev-4B reflected that ambiguity in its probability distribution instead of completely ignoring the other possibilities.\n\nThe urgency decision was almost evenly split:\n\n```\nurgent → 0.5069\n```\n\nAgain, this is far from a strong signal.\n\nThe severity result was different:\n\n```\nLow      0.0402\nMedium   0.1445\nHigh     0.8153\n```\n\nHere the model was much more decisive.\n\nThe resulting score was:\n\n```\n1.7751\n```\n\nwith High receiving the largest probability mass.\n\nThe 4B model did something easy to miss when looking only at the final label.\n\nIf I only looked at:\n\n```\ndepartment → support\n```\n\nI might assume the model was confident.\n\nIt wasn't.\n\nThe probabilities were almost evenly split between support and billing.\n\nThat is exactly why I think the probability output is more useful than a single label.\n\nA simple application could treat this as an uncertain routing decision:\n\n```\nif max(probabilities.values()) < 0.80:\n    send_to_manual_review()\n```\n\nThe severity signal could potentially be handled differently because its probability distribution is much more concentrated.\n\nNow we have two real runs using the same state and questions:\n\n```\n|                           |   Kev-0.8B |      Kev-4B |\n| ------------------------- | ---------: | ----------: |\n| Department                |    support |     support |\n| Support probability       | **0.6763** |  **0.4155** |\n| Billing probability       |     0.1393 |  **0.4014** |\n| Urgency                   |     0.6559 |      0.5069 |\n| Severity score            |     1.3554 |  **1.7751** |\n| High severity probability |     0.5247 |  **0.8153** |\n| Reported latency          | 1,520.1 ms | 21,549.1 ms |\n```\n\nOne thing immediately stood out: the larger model did not simply become “more confident” about everything.\n\nIn fact, its department prediction was less concentrated, while its severity prediction was more concentrated.\n\nAnd the reported latency in this particular 4B run was substantially higher.\n\nI don't want to turn these two requests into a benchmark, though. They're individual local measurements, not a controlled latency study. We'll need repeated runs before concluding performance.\n\nThe more interesting takeaway for me was that changing the model can change not only the final decision, but also the shape of the probability distribution behind that decision.\n\nThat is exactly what I wanted to investigate by running the same workload across multiple Kev checkpoints.\n\nAt this point, we've gone from the basic idea behind Kev to actually running it locally.\n\nWe've looked at how the decision architecture works, how the different decision primitives are represented, how the pointer head produces probabilities, and then tested real checkpoints on the same support-ticket workload.\n\nSo far, I tested:\n\n```\nKev-0.5B\nKev-0.6B\nKev-0.8B\nKev-4B\n```\n\nAnd the interesting thing is that simply increasing the model size didn't produce one simple pattern.\n\nThe probability distributions changed.\n\nThe level of uncertainty changed.\n\nThe latency changed.\n\nEven when the final selected decision stayed the same, the confidence behind that decision could look very different.\n\nAnd that is exactly why I don't want to draw conclusions from only four models.\n\nThere are still two checkpoints left in the experiment:\n\n```\nKev-8B\nKev-9B\n```\n\nBoth are interesting for different reasons, especially because the 8B checkpoint belongs to the earlier Qwen3 generation while the 9B checkpoint belongs to the newer Qwen3.5 generation.\n\nI'll cover those experiments in the next article, using the same state, the same questions, and the same evaluation approach so that we can see what actually changes.\n\nWhat started as a simple question—\n\n“Can we build something like Jev openly?”\n\n—turned into a much more interesting exploration of how AI models can fit inside software.\n\nKev is not trying to be another general-purpose chatbot.\n\nIts interface is much narrower:\n\n```\nState\n  +\nTyped question\n       ↓\nDecision\n       ↓\nProbability\n       ↓\nApplication logic\n```\n\nThat sounds simple, but the design opens up a different way of thinking about AI applications.\n\nInstead of asking one large language model to generate text for every small decision, we can imagine specialized models sitting inside specific parts of a system:\n\n```\nSupport ticket\n      ↓\nDecision model\n      ↓\nRoute / Escalate / Review\n```\n\nAnd the open-source ecosystem makes this even more interesting.\n\nJev introduced the idea in a closed implementation, while projects like Laya, Kev, and others are experimenting with their own approaches to decision-first models.\n\nThey are not all implementing the same architecture, and they should not be treated as direct copies of Jev. What they do share is a broader idea:\n\nAI doesn't always need to speak. Sometimes it just needs to decide.\n\nThat is the part of this space I find most interesting.\n\nAnd we haven't finished the experiment yet.\n\nLike | Follow | Subscribe to the newsletter.\n\nConnect with me on:\n\nGitHub: [https://github.com/Ayush7614](https://github.com/Ayush7614)\n\nLinkedIn: [https://www.linkedin.com/in/ayush-kumar-984443191/](https://www.linkedin.com/in/ayush-kumar-984443191/)\n\nTwitter: [https://x.com/AYUSHKUMAR82274](https://x.com/AYUSHKUMAR82274)\n\nSubstack: [https://substack.com/@felixayush](https://substack.com/@felixayush)\n\nDev.to Blog: [https://dev.to/ayush7614](https://dev.to/ayush7614)\n\nPersonal Blog: [https://neural-verse-peach.vercel.app/](https://neural-verse-peach.vercel.app/)\n\nWebsite: [https://ayushbuilds-dev.vercel.app/](https://ayushbuilds-dev.vercel.app/)", "url": "https://wpnews.pro/news/kev-an-open-source-jev-alternative-i-ran-locally", "canonical_source": "https://dev.to/ayush7614/kev-an-open-source-jev-alternative-i-ran-locally-gek", "published_at": "2026-09-30 10:43:05+00:00", "updated_at": "2026-09-30 10:47:39.519861+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-agents", "ai-research"], "entities": ["Kev", "Jev", "Laya", "Qwen", "Hugging Face", "GitHub", "Apple Silicon", "jaredpalmer"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/kev-an-open-source-jev-alternative-i-ran-locally", "markdown": "https://wpnews.pro/news/kev-an-open-source-jev-alternative-i-ran-locally.md", "text": "https://wpnews.pro/news/kev-an-open-source-jev-alternative-i-ran-locally.txt", "jsonld": "https://wpnews.pro/news/kev-an-open-source-jev-alternative-i-ran-locally.jsonld"}}