cd /news/ai-tools/kev-an-open-source-jev-alternative-i… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-142432] src=dev.to β†— pub= topic=ai-tools verified=true sentiment=↑ positive

Kev: An Open-Source Jev Alternative I Ran Locally

A developer ran the open-source decision model Kev locally on Apple Silicon, testing the Kev-0.5B, 0.6B, 0.8B, and 4B checkpoints on a shared support-ticket workload and comparing their probability distributions, uncertainty, and latency. Kev, positioned as an open alternative to the closed Jev model, takes a state plus typed questions and returns structured decisions with probabilities rather than generated text. The developer reports that changing checkpoint size altered confidence, uncertainty, and latency rather than simply increasing confidence, with Kev-8B and Kev-9B left for a follow-up.

by read22 min views1 publishedSep 30, 2026

If you spend enough time around the AI and open-source community, you have probably noticed the noise around Jev.

Jev appeared with a somewhat different idea: instead of using an AI model mainly to generate text, what if the model was designed to make fast, typed decisions?

That immediately caught my attention.

But there was another interesting part of the story.

Jev itself wasn't released as an open-source model that developers could simply download, inspect, modify, and run however they wanted. That left a natural question for the open-source community:

Can we build something similar ourselves?

And not long after, projects started appearing around the same general idea.

Small decision models.

Open implementations.

Different architectures.

Different training approaches.

Different ways of producing probabilities.

Projects such as Laya and Kev are part of that growing conversation, along with several other experiments exploring the broader System One approach.

I had already spent time looking at Laya and writing about the idea behind Jev. This time, I wanted to take the same approach with Kev.

Not just:

β€œHere is another Jev alternative.”

I wanted to actually run it.

I wanted to understand what was happening under the hood, install the models locally, send them the same questions, look at the probability distributions, see how the model behaves as the size changes, and find out what actually works on consumer hardware.

That is what this article is about.

We will start with the architecture and the idea behind Kev, then move into a hands-on experiment with its different checkpoints.

And rather than treating all of these models as interchangeable copies of Jev, I'll treat them as what they are: independent open-source attempts at building decision-first AI systems.

So let's start with the basic question.

I’ve been exploring Jev and the growing open-source ecosystem around decision-first AI, and Kev caught my attention because it takes the idea and makes it runnable and customizable.

Instead of generating text, Kev takes a state + typed questions and returns structured decisions with probabilities:

State
  ↓
Typed Questions
  ↓
Kev
  ↓
Decision + Probability
  ↓
Application Logic

In this walkthrough, I explored the architecture behind Kev and ran Kev-0.5B, 0.6B, 0.8B, and 4B locally on Apple Silicon using the same support-ticket workload.

The results were interesting: changing the checkpoint didn't simply make the model β€œmore confident.” The actual probability distributions, uncertainty, and latency changed across models.

There are still more checkpoints to explore, especially Kev-8B and Kev-9B, which I'll test in the next article.

The bigger idea: AI doesn't always need to generate text. Sometimes, it just needs to make a decision.

| Requirement             | Details                                                   |
| ----------------------- | --------------------------------------------------------- |
| **Python**              | 3.12 or 3.13                                              |
| **Git**                 | Required to clone the Kev repository                      |
| **Virtual Environment** | Python `venv` recommended                                 |
| **Package Manager**     | `pip`                                                     |
| **OS**                  | macOS / Linux                                             |
| **Apple Silicon**       | MPS + MLX supported for compatible checkpoints            |
| **NVIDIA GPU**          | Supported for larger checkpoints                          |
| **Models Tested**       | Kev-0.5B, 0.6B, 0.8B, 4B, 8B, 9B                          |
| **Storage**             | Enough space for the selected model and Qwen base weights |
| **Internet**            | Required for the first model download from Hugging Face   |

Model Link Page: https://huggingface.co/jaredpalmer

GitHub: https://github.com/jaredpalmer/kev

Let's start with the simplest possible explanation.

A traditional LLM workflow often looks like this:

User input
    ↓
LLM
    ↓
Generated text
    ↓
Parse the text
    ↓
Application logic

Maybe we ask the model:

Which department should handle this support ticket?

And then tell it to return JSON:

{
  "department": "billing"
}

That looks structured, but the model is still fundamentally generating tokens.

Kev approaches the same problem differently.

The application provides:

The model then produces a decision and its probability distribution.

The high-level flow looks like this:

The three main decision types are:

noul
choice
score

You can think of them as:

noul   β†’ yes/no
choice β†’ select one option
score  β†’ assign a level

So instead of asking Kev to write a paragraph about a customer ticket, we can ask:

Which department should handle this?

Does this need urgent human attention?

How frustrated is the customer?

That is a much smaller interface.

Imagine a customer sends this:

My package arrived two weeks late, the shoes are the wrong size, and I was charged twice.

A conventional LLM could summarize the message and explain what should happen.

With Kev, we can define three decisions:

Department
β†’ returns/shipping/billing

Urgency
β†’ yes/no

Frustration
β†’ calm / frustrated / very angry

Conceptually, the model can return something like:

Department

returns     0.47
shipping    0.28
billing     0.25

Urgency

yes         0.93

Frustration

calm        0.00
frustrated  0.56
very angry  0.44

The application can then decide what to do.

That last step is important.

The model makes the decision. The application owns the action.

That distinction becomes very useful when building production systems.

This was probably the first question I had when I started looking at these projects.

We already have extremely capable LLMs.

So why create a separate model for decisions?

The answer is not that LLMs suddenly became useless.

It is that generation and decision-making are different interfaces.

A general LLM is designed to continue a sequence.

A decision model can instead expose the thing an application actually wants:

decision
probability
decision
probability
decision
probability

The architecture therefore starts looking less like:

and more like:

That is the core idea behind the entire experiment.

One of the first things you notice when looking at the project is that there are several Kev checkpoints.

The current family contains:

Kev-0.8B
Kev-4B
Kev-9B
Kev-27B

The repository also keeps older checkpoints:

Kev-0.5B
Kev-0.6B
Kev-8B

The project documentation distinguishes those older releases from the current model family. Kev-0.5B is the older Qwen2.5-based reference model; the 0.6B and 8B checkpoints belong to the previous Qwen3 generation. The current family has moved to Qwen3.5, while Kev-27B uses Qwen3.8.

That gives us a nice little history:

For this walkthrough, I want to go through the models up to Kev-9B.

That means we can look at:

| Model    | Generation | What we'll do   |
| -------- | ---------- | --------------- |
| Kev-0.5B | Qwen2.5    | Run and inspect |
| Kev-0.6B | Qwen3      | Run and inspect |
| Kev-0.8B | Qwen3.5    | Run and inspect |
| Kev-4B   | Qwen3.5    | Run and test    |
| Kev-8B   | Qwen3      | Run and inspect |
| Kev-9B   | Qwen3.5    | Run and test    |

I'm leaving Kev-27B out of the local hands-on part because its hardware requirements are in a completely different category. The repository lists 80 GB-class GPU hardware for that checkpoint.

This is probably the most interesting part of the project.

According to the repository, Kev checkpoints use a rank-16 LoRA adapter and a pointer head on top of a Qwen base model. The pointer head scores option representations against a decision representation, and a softmax converts those scores into probabilities.

A simplified version looks like this:

Suppose the pointer head produces these scores:

returns     1.72
shipping    1.21
billing     1.08

Softmax turns those scores into something like:

returns     0.47
shipping    0.28
billing     0.25

Now our application has something it can reason about directly.

It can say:

if returns_probability > 0.80:
    route_to_returns()
else:
    send_to_human()

The model is no longer responsible for deciding what the entire application should do.

It provides the signal.

This is another detail that makes Kev different from simply asking an LLM three questions in a prompt.

Imagine we have:

Question 1 β†’ Which department?
Question 2 β†’ Is it urgent?
Question 3 β†’ How frustrated is the customer?

The same state can be reused for those decisions while keeping the questions isolated.

Conceptually:

The Kev repository describes the questions as sharing the text but not reading one another. For the newer Qwen3.5/Qwen3.8 models, the implementation runs each question as its own row because the underlying Gated DeltaNet layers are recurrent and do not follow attention masks in the same way as attention-only models. The state can still be computed once and reused through caching.

That sounds complicated, but the practical idea is simple:

One piece of state can feed many independent decisions.

The previous generation is worth understanding because it explains how the design evolved.

For attention-only Qwen3 bases, the repository describes a sequence structure roughly like:

<state> ...state...

<q> instructions
    <opt> option 1 </opt>
    <opt> option 2 </opt>
    <opt> option 3 </opt>
    <decide>

<q> instructions
    <opt> option 1 </opt>
    <opt> option 2 </opt>
    <opt> option 3 </opt>
    <decide>

The attention mask prevents one question from reading another.

We can visualize that as:

For the current Qwen3.5/Qwen3.8-based models, the implementation takes a different route because of the recurrent components of those base models.

So there is an important distinction:

Older Qwen3 Kev
β†’ attention masking

Current Qwen3.5 / Qwen3.8 Kev
β†’ independent rows + shared cached state

That distinction comes directly from the current project implementation and is worth preserving rather than collapsing all Kev versions into one generic architecture description.

Now we can zoom in one level further.

The pointer head is the piece that turns the model's representations into option scores.

Very roughly:

This is why the output is naturally suited to a choice question.

Instead of asking the model to generate:

The most likely department is returns because...

we can compare the available options and derive a probability distribution over them.

The same overall idea is used for noul and score, with their own output interpretation.

Let's make this concrete.

A binary decision.

Is this customer asking for a refund?

yes β†’ 0.91

The value represents the probability of the positive outcome.

Choose one item from several options.

Which department should handle this?

returns     0.47
shipping    0.28
billing     0.25

Kev returns the selected choice along with the distribution.

Instead of picking one label, the model can place probability across ordered levels.

For example:

How frustrated is the customer?

0 β†’ Calm
1 β†’ Frustrated
2 β†’ Very angry

The result can contain:

0 β†’ 0.00
1 β†’ 0.56
2 β†’ 0.44

and the resulting score can be represented as a weighted level.

This is useful when the distinction is not simply yes/no or category A/category B.

This is one of the details I care about most.

A normal classifier might simply return:

billing

That gives the application very little information about uncertainty.

Kev can return something closer to:

billing    0.51
returns    0.29
shipping   0.20

Now the application can make its own decision.

But there is an important caveat here.

Probability is not the same thing as correctness.

A model can be highly confident and still be wrong.

That is why the Kev repository includes calibration, and why I would not build a production workflow that blindly says:

if confidence > 0.9:
    trust_the_model()

A threshold needs to be validated against the actual workload.

Now we can finally get away from the architecture diagrams and run Kev.

The current project recommends Python 3.12 or 3.13 and uv. The repository's .python-version currently points the normal uv workflow toward Python 3.13.

I prefer starting with a clean checkout rather than mixing it into another Python environment:

git clone https://github.com/jaredpalmer/kev.git
cd kev

I'd use Python 3.13:

python3.13 --version

You should see something like:

Python 3.13.x
python3.13 -m venv .venv

This creates:

kev/
β”œβ”€β”€ .venv/
β”œβ”€β”€ kev/
β”œβ”€β”€ README.md
β”œβ”€β”€ pyproject.toml
└── ...
source .venv/bin/activate

Your terminal should now look roughly like:

(.venv) (base) ayushkumar@Ayushs-Mac-mini-2 kev %
python --version
which python

The second command should point inside your project:

.../kev/.venv/bin/python

Since Kev defines its serving dependencies as a serve extra, we can install the project directly into this virtual environment:

python -m pip install --upgrade pip
pip install -e ".[serve]"

The serve extra includes FastAPI, the TypeSafe SDK, Uvicorn, and the Apple Silicon MLX backend when running on an ARM Mac.

python -c "import kev; print('Kev installed successfully')"

Then check the package:

pip show kev
python -m kev.serve \
  --run jaredpalmer/kev-0.8b \
  --port 8009

Then, in another terminal:

curl http://localhost:8009/v1/models

At this point, Kev was installed, and the local server was up.

The next thing I wanted to verify was not whether the process was alive, but which model and backend were actually being used.

I ran:

curl http://localhost:8009/v1/models

The response showed that my Mac was running:

Model:       jaredpalmer/kev-0.8b
Base:        Qwen/Qwen3.5-0.8B-Base
LoRA rank:   16
Device:      MPS
Backend:     MLX
Dtype:       bfloat16
Temperature: 2.351...

So the local setup looked like this:

The part I found interesting here was the backend.

Kev was not running through CUDA on my machine. Since this is an Apple Silicon Mac, it was using the MLX backend with MPS and bfloat16.

There was also a jev-latest entry in the response.

That does not mean TypeSafe's hosted Jev was running locally. In this setup, that name points to the same local Kev checkpoint. So for the rest of the article, I'll use kev-latest to avoid confusing the two.

The server was ready.

Now it was time to actually ask Kev a question.

Until now, we had only verified that the model loaded.

A model running successfully is one thing.

Getting it to make a useful decision is another.

So I created a small support-ticket example.

Here is the situation:

β€œMy laptop was delivered three days late, the screen is damaged, and I was charged twice.”

There are several different pieces of information in that single sentence.

Instead of asking an LLM to explain the problem, I want Kev to answer three specific questions:

1. Which team should handle this issue?
2. Does this require urgent human attention?
3. How serious is the issue?

That maps directly to Kev's three decision primitives:

choice β†’ department
noul   β†’ urgency
score  β†’ severity

So the request becomes:

This is where the API becomes interesting.

The server was running, so the next step was simple: give Kev an actual problem to solve.

I used a support-ticket scenario with three different types of questions:

State:
"My laptop was delivered three days late,
the screen is damaged, and I was charged twice."

Then I asked Kev:

1. Which team should handle this issue?
2. Does this require urgent human attention?
3. How serious is this customer issue?

That gives us one choice, one noul, and one score question in the same request.

Here is the request I sent:

curl -s http://localhost:8009/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "state": "My laptop was delivered three days late, the screen is damaged, and I was charged twice.",
    "model": "kev-latest",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Which team should handle this issue?",
        "criteria": {
          "support": "Hardware problems, damaged devices, technical issues",
          "shipping": "Delivery delays, tracking, lost packages",
          "billing": "Charges, invoices, payment problems"
        }
      },
      "urgent": {
        "type": "noul",
        "instructions": "Does this require urgent human attention?"
      },
      "severity": {
        "type": "score",
        "instructions": "How serious is this customer issue?",
        "criteria": [
          "Low",
          "Medium",
          "High"
        ]
      }
    }
  }'

And this time, instead of using a made-up response, let's look at what my local Kev-0.8B instance actually returned.

{
  "model": "kev-latest",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "support",
      "confidence": 0.5144,
      "probabilities": {
        "support": 0.6763,
        "shipping": 0.1844,
        "billing": 0.1393
      }
    },
    "urgent": {
      "type": "noul",
      "noul": 0.6559
    },
    "severity": {
      "type": "score",
      "score": 1.3554,
      "legend": {
        "0": "Low",
        "1": "Medium",
        "2": "High"
      },
      "probabilities": {
        "0": 0.1693,
        "1": 0.306,
        "2": 0.5247
      },
      "confidence": 0.0332
    }
  },
  "usage": {
    "input_tokens": 95,
    "output_tokens": 173
  },
  "latency_ms": 1520.1
}

This one request produced three different kinds of outputs:

This is a nice demonstration of why the System One-style interface is interesting.

One piece of state can produce several independent decision signals in one request.

The response also included:

latency_ms: 1520.1

So this particular request took approximately:

1.52 seconds

according to the API's reported latency.

The request contained:

95 input tokens
173 output tokens

There is an important detail here, though.

This is one local run, not a benchmark.

A single request doesn't tell us what the normal p50 or p95 latency looks like, and it certainly isn't enough to compare Kev with another model.

We'll collect more measurements later.

For now, this number simply tells us what happened during this particular experiment.

The most interesting thing about this first test wasn't actually the support result.

It was the fact that one model call gave the application three different signals:

Department β†’ support
Urgency    β†’ 0.6559
Severity   β†’ 1.3554

A conventional LLM could certainly produce the same information, but we'd normally be asking it to generate some structured response and then parse that response.

Here, the interface is built around the decisions themselves.

That changes the programming model.

Instead of:

"Tell me what you think."

we are effectively asking:

"Answer these specific decisions."

And that is much closer to how normal application logic works.

Now let's make the experiment slightly more practical.

Suppose our application has these rules:

High urgency
β†’ human escalation

Strong support classification
β†’ technical support

Everything uncertain
β†’ manual review

We can represent that with a simple flow:

This is where the separation between the model and the application becomes important.

Kev doesn't need to know what your business process is.

It provides the decision signals.

Your software decides whether those signals are enough to trigger an action.

Our first experiment also reinforces a rule that is easy to forget.

A probability is not a guarantee.

In this run:

support = 0.6763

doesn't mean:

β€œThe model is 67.63% correct.”

And:

urgent = 0.6559

doesn't mean:

β€œThere is definitely a 65.59% chance that a human should intervene.”

These values describe the model's output distribution. Whether a particular probability threshold is actually useful needs to be evaluated against the real workload.

This becomes especially important once we start automating actions.

Kev-0.8B has now passed the first basic test.

But one model isn't enough for what I want to explore.

The repository contains several generations and sizes, and I want to see how the same decision workload behaves across them.

So next I'm going to run the same experiment with:

Kev-0.5B
Kev-0.6B
Kev-0.8B
Kev-4B
Kev-8B
Kev-9B

We'll keep the state, questions, and criteria the same.

That gives us a much cleaner experiment:

Kev-0.5B is the earliest checkpoint in this family and is based on Qwen2.5-0.5B.

For this test, I used the same support-ticket request from the previous section.

Start it with:

python -m kev.serve \
  --run jaredpalmer/kev-0.5b \
  --port 8009

Then verify:

curl http://localhost:8009/v1/models

Once the server is up, send the same /v1/systemone request.

The local server loaded it successfully on my Mac:

Model:       jaredpalmer/kev-0.5b
Base:        Qwen/Qwen2.5-0.5B
Device:      MPS
Backend:     PyTorch
Dtype:       bfloat16
LoRA rank:   16
Temperature: 1.0

The /v1/models endpoint also reports a prefix cache of up to 4 states, with caching enabled for states of at least 384 tokens.

One small detail worth noting: kev-latest and jev-latest both point to the same local Kev-0.5B checkpoint. The latter is simply the server's compatibility alias; it is not TypeSafe's hosted Jev.

At this stage, the important thing is that the 0.5B model runs locally on Apple Silicon, giving us a lightweight baseline before moving to the newer checkpoints.

Next, I moved to Kev-0.6B, an older-generation Kev checkpoint built on Qwen3-0.6B-Base.

Run:

python -m kev.serve \
  --run jaredpalmer/kev-0.6b \
  --port 8009

Then:

curl http://localhost:8009/v1/models

The model loaded successfully on my Mac:

Model:       jaredpalmer/kev-0.6b
Base:        Qwen/Qwen3-0.6B-Base
Device:      MPS
Backend:     PyTorch
Dtype:       bfloat16
LoRA rank:   16
Temperature: 1.0

So the main change from Kev-0.5B is the underlying Qwen generation:

Kev-0.5B
Qwen2.5-0.5B

        ↓

Kev-0.6B
Qwen3-0.6B

The server again exposes the same kev-latest and jev-latest compatibility names, both pointing to the local Kev-0.6B checkpoint.

At this stage, the setup was working exactly as expected. The next step was to send the same support-ticket request we used for Kev-0.8B and see whether moving from Qwen2.5 to Qwen3 changes the actual decisions or probability distribution.

Quick comparison so far

Model   Base    Backend Device  Dtype
Kev-0.5B    Qwen2.5-0.5B    PyTorch MPS BF16
Kev-0.6B    Qwen3-0.6B  PyTorch MPS BF16
Kev-0.8B    Qwen3.5-0.8B    MLX MPS BF16

After testing the smaller checkpoints, I moved to Kev-4B, one of the current-generation models in the Kev family.

Kev-4B is built on Qwen3.5-4B-Base and, on my Apple Silicon machine, it loaded through the MLX backend.

The /v1/models response showed:

Model:       jaredpalmer/kev-4b
Base:        Qwen/Qwen3.5-4B-Base
LoRA rank:   16
Device:      MPS
Backend:     MLX
Dtype:       bfloat16
Temperature: 2.406...

The response came back with:

{
  "department": {
    "choice": "support",
    "confidence": 0.1233,
    "probabilities": {
      "support": 0.4155,
      "shipping": 0.1831,
      "billing": 0.4014
    }
  },
  "urgent": {
    "noul": 0.5069
  },
  "severity": {
    "score": 1.7751,
    "probabilities": {
      "0": 0.0402,
      "1": 0.1445,
      "2": 0.8153
    },
    "confidence": 0.6627
  }
}

The model reported:

Input tokens:  95
Output tokens: 174
Latency:       21549.1 ms

The interesting part was not simply that Kev-4B selected support.

The probability distribution was much closer between two options:

support     0.4155
billing     0.4014
shipping    0.1831

So although support was the selected option, the model was not particularly decisive about the department.

That is actually useful information.

The ticket contains three different issues:

Damaged laptop β†’ support
Late delivery  β†’ shipping
Double charge  β†’ billing

Kev-4B reflected that ambiguity in its probability distribution instead of completely ignoring the other possibilities.

The urgency decision was almost evenly split:

urgent β†’ 0.5069

Again, this is far from a strong signal.

The severity result was different:

Low      0.0402
Medium   0.1445
High     0.8153

Here the model was much more decisive.

The resulting score was:

1.7751

with High receiving the largest probability mass.

The 4B model did something easy to miss when looking only at the final label.

If I only looked at:

department β†’ support

I might assume the model was confident.

It wasn't.

The probabilities were almost evenly split between support and billing.

That is exactly why I think the probability output is more useful than a single label.

A simple application could treat this as an uncertain routing decision:

if max(probabilities.values()) < 0.80:
    send_to_manual_review()

The severity signal could potentially be handled differently because its probability distribution is much more concentrated.

Now we have two real runs using the same state and questions:

|                           |   Kev-0.8B |      Kev-4B |
| ------------------------- | ---------: | ----------: |
| Department                |    support |     support |
| Support probability       | **0.6763** |  **0.4155** |
| Billing probability       |     0.1393 |  **0.4014** |
| Urgency                   |     0.6559 |      0.5069 |
| Severity score            |     1.3554 |  **1.7751** |
| High severity probability |     0.5247 |  **0.8153** |
| Reported latency          | 1,520.1 ms | 21,549.1 ms |

One thing immediately stood out: the larger model did not simply become β€œmore confident” about everything.

In fact, its department prediction was less concentrated, while its severity prediction was more concentrated.

And the reported latency in this particular 4B run was substantially higher.

I don't want to turn these two requests into a benchmark, though. They're individual local measurements, not a controlled latency study. We'll need repeated runs before concluding performance.

The more interesting takeaway for me was that changing the model can change not only the final decision, but also the shape of the probability distribution behind that decision.

That is exactly what I wanted to investigate by running the same workload across multiple Kev checkpoints.

At this point, we've gone from the basic idea behind Kev to actually running it locally.

We've looked at how the decision architecture works, how the different decision primitives are represented, how the pointer head produces probabilities, and then tested real checkpoints on the same support-ticket workload.

So far, I tested:

Kev-0.5B
Kev-0.6B
Kev-0.8B
Kev-4B

And the interesting thing is that simply increasing the model size didn't produce one simple pattern.

The probability distributions changed.

The level of uncertainty changed.

The latency changed.

Even when the final selected decision stayed the same, the confidence behind that decision could look very different.

And that is exactly why I don't want to draw conclusions from only four models.

There are still two checkpoints left in the experiment:

Kev-8B
Kev-9B

Both are interesting for different reasons, especially because the 8B checkpoint belongs to the earlier Qwen3 generation while the 9B checkpoint belongs to the newer Qwen3.5 generation.

I'll cover those experiments in the next article, using the same state, the same questions, and the same evaluation approach so that we can see what actually changes.

What started as a simple questionβ€”

β€œCan we build something like Jev openly?”

β€”turned into a much more interesting exploration of how AI models can fit inside software.

Kev is not trying to be another general-purpose chatbot.

Its interface is much narrower:

State
  +
Typed question
       ↓
Decision
       ↓
Probability
       ↓
Application logic

That sounds simple, but the design opens up a different way of thinking about AI applications.

Instead of asking one large language model to generate text for every small decision, we can imagine specialized models sitting inside specific parts of a system:

Support ticket
      ↓
Decision model
      ↓
Route / Escalate / Review

And the open-source ecosystem makes this even more interesting.

Jev introduced the idea in a closed implementation, while projects like Laya, Kev, and others are experimenting with their own approaches to decision-first models.

They are not all implementing the same architecture, and they should not be treated as direct copies of Jev. What they do share is a broader idea:

AI doesn't always need to speak. Sometimes it just needs to decide.

That is the part of this space I find most interesting.

And we haven't finished the experiment yet.

Like | Follow | Subscribe to the newsletter.

Connect with me on:

GitHub: https://github.com/Ayush7614

LinkedIn: https://www.linkedin.com/in/ayush-kumar-984443191/

Twitter: https://x.com/AYUSHKUMAR82274

Substack: https://substack.com/@felixayush

Dev.to Blog: https://dev.to/ayush7614

Personal Blog: https://neural-verse-peach.vercel.app/

Website: https://ayushbuilds-dev.vercel.app/

── more in #ai-tools 4 stories Β· sorted by recency
── more on @kev 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/kev-an-open-source-j…] indexed:0 read:22min 2026-09-30 Β· β€”