cd /news/ai-tools/kev-and-laya-the-open-source-answer-… · home › topics › ai-tools › article
[ARTICLE · art-145654] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Kev and Laya: The Open Source Answer to TypeSafe's Jev

Two open-source projects, Kev and Laya, emerged within days of TypeSafe's closed-weight Jev launch, with Laya surpassing 19,000 GitHub stars and Kev matching Jev within a point on new sources (0.851 vs 0.857) while trailing badly on MMLU-Pro (0.675 vs 0.840). Kev, from Turborepo developer Jared Palmer, ships Apache-2.0 decision models on Qwen3.5/3.8 that match TypeSafe's /v1/systemone API contract, while Laya's research found its English checkpoint reported 0.952 confidence on Khmer while scoring 0.000 accuracy.

by read7 min views38 publishedOct 5, 2026

"12 million views for a JSON classifier? Yeah, we're in a bubble."

That tweet was Niels Rogge's, the Hugging Face guy. Half of X was losing its mind over a model that answers yes/no questions. The other half was annoyed anyone cared.

The launch was September 15, 2026. Two years in stealth. A founder from OpenAI.

Locked weights, waitlist access.

Jev does one thing: send it text plus typed questions, get probabilities back. Choice, score, yes/no. No sentences, no tokens, nothing to parse.

And the reaction was huge. The HN thread passed 1,900 points with 520 comments.

Then something more interesting happened. People started rebuilding it.

Within 24 hours the first clones appeared. AINews counted six in two days.

The awesome-jev list now tracks more than a dozen. Laya, the biggest, passed 19,000 stars and added 5,000 in one day.

I kept refreshing GitHub, wondering why a JSON classifier gets this much energy.

The answer says something about how we feel about hosted AI. If you've ever paid an LLM to write text you immediately parse into an if statement, this fight is about you.

Here's a question people always ask: what is a System One model?

The shape is simple. You send state, any text, email or JSON document, with questions attached. Each question declares its answer type.

{
  "state": "My payouts have failed three times.",
  "questions": {
    "queue": {"type": "choice", "instructions": "Which team handles this?"},
    "escalate": {"type": "noul", "instructions": "Needs urgent human attention?"}
  }
}

The model reads it once and answers everything in parallel. It returns probabilities, not prose.

Jev's founder is Diogo Almeida, who helped build the instruction work that became ChatGPT at OpenAI. Two years in stealth, and the pitch is sharp.

think of Jev as a frontier-intelligence function call.

No string generation means no hallucinated JSON, no parsing, no repair step. The model can't hallucinate. It also can't write a sentence.

Input runs $0.042 per million tokens, output is free, and responses land in 70 to 500 milliseconds.

A frontier chat model takes 3 to 329 seconds for the same job. That gap is the whole story.

What happens after a launch like this is predictable. Closed weights and no technical paper get read as homework.

A researcher, Archer Hume, probed the public API, watching latency scale with context and answers shift when questions reordered.

"Jev's Architecture Unmasked" came out of it: a causal transformer, probably sparse MoE, that encodes state once, runs every question branch in parallel, and reads probabilities straight off internal representations. He admits it's speculative.

It was also the clearest picture anyone had.

That post became a blueprint. The bubble takes were loud. But the builders were louder.

Most tutorials tell you to fine-tune a model for this. Kev shows the other way.

Kev comes from Jared Palmer, the Turborepo developer. It's a family of small decision models on Qwen3.5 and Qwen3.8, following the unmasked architecture. Apache-2.0, four sizes, from a 0.8B for a laptop to a 27B for a data centre GPU.

The smart move is the API. Kev matches TypeSafe's /v1/systemone contract, so you point TypeSafe's own Python SDK at your local server and nothing changes. That's how you take a closed product's users the polite way.

Then the evals. The part i respect most.

Kev-27B lands within a point of Jev on new sources, 0.851 against 0.857. And the README says it plainly: this isn't a controlled comparison, because we don't know what Jev was trained on.

It also lists where Kev loses. On MMLU-Pro it scores 0.675 against Jev's 0.840.

On day-precision date math, the small models trail.

That honesty is rarer than the weights.

The problem isn't what you think it is. It's confidence.

Laya's own research found the English checkpoint scored 0.000 accuracy on Khmer while reporting 0.952 confidence. A model that is certain and completely wrong.

If you branch production code on that, you ship silent failures. Laya fixes it with a router that detects the script before the forward pass and switches to a multilingual checkpoint.

Twenty-two alphabets, under half a millisecond.

Laya is the star of this wave. 421 million parameters on a ModernBERT backbone, pip install laya, runs on a CPU box, answers in roughly 21 to 33 milliseconds.

It shipped September 18, three days after Jev, from Convai Innovations. Its founder says he published the core idea first, in a March 2025 arXiv paper, then built the open version instead of staying bitter.

"Instead of staying bitter, I decided to take everything I learned, fix every architectural limitation of the old approach, and build a completely open, horizontal System 1 decision model family."

The history is genuinely contested, and one review called the David and Goliath framing marketing-adjacent. Fair.

So is the benchmark catch. The headline accuracy number for Laya comes from a checkpoint fine-tuned on the benchmark's own training split.

Zero-shot, the base model scores below the majority-class baseline. Out of the box it is a base to fine-tune, not a drop-in.

Weights Out-of-box accuracy Price
Jev Closed Highest $0.042/MTok input
Kev Apache-2.0 Within a point of Jev (27B) Free
Laya Apache-2.0 Near chance, fine-tune it Free

I used to think cloning a model took a lab. Then this weekend happened.

The interface is public, so the wave moved fast. The weights and data are not, so quality lagged.

SemIf trains nothing at all. It reads answer logits off a frozen Qwen3.5-4B and ran 21 decisions five times faster than generated JSON.

NanoJev is a 0.6B model for real-time loops and beat Jev on ViZDoom Basic, 128 out of 128 against 56. jevlike ships a trainer instead of a model.

Here's what the wave means:

On an independent 49-task benchmark, Jev still scores 0.966 macro accuracy. The best open entrant scores 0.704.

The clones win on speed and price. The scoreboard is not close yet.

The part i couldn't stop watching was Flappy Bird. Laya plays it live, about 30 decisions a second, by asking itself one question: where is the bird?

Asked which way to move, every checkpoint answered backwards. Asked where the bird is, clean graded answers.

P(below) of 0.95, 0.82, 0.06. The game flaps when the probability crosses half.

Same story in Tetris. 1,799 decisions in a minute, 52 lines cleared, no top-out.

It loses eventually, because the pieces fall faster and a life lasts about two minutes. And it cannot read numbers at all.

Give it two altitudes and no checkpoint can say which is lower. Do the arithmetic in code, hand the model the conclusion in words.

That's the weirdest part of this whole genre. The question is the craft.

Reword it and accuracy swings by 30 points. I spent an evening rephrasing prompts to watch the probability move.

My partner watched me do this and asked if it was work. I said yes, and i wasn't sure.

Most people should not run a decision model this week.

The open alternatives win on price, latency and control. Not accuracy. Not yet.

If you route a thousand tickets a day and your vendor doubles the price, Laya or Kev is your insurance.

If you route 50 tickets a day, use a folder and an if statement. Seriously. A 4GB model for your inbox is overkill.

You'll pay for the wiring with your weekend.

If you do pick one, plan to fine-tune it. Laya's base checkpoints score near chance without it.

Kev's own docs tell you where it loses, so read those before you trust a threshold. And the cheap irony: Jev is not expensive.

$0.042 per million tokens is basically nothing. Free is a win for privacy and control, not your invoice.

The real take: a model is not a moat. An eval set with your data in it is.

I keep coming back to the Khmer number. 0.000 accuracy, 0.952 confidence. A model that was certain and catastrophically wrong, because nobody checked what the words meant.

The open source wave matters because it makes checking possible. Weights you can audit. Datasets you can read.

Evals someone else can rerun. Kev and Laya are not better than Jev yet, and they will tell you that themselves, which is exactly the point.

i still don't know who invented the idea. But i know who lets me check their work. That's the one i'm building with.

── more in #ai-tools 4 stories · sorted by recency
── more on @typesafe 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/kev-and-laya-the-ope…] indexed:0 read:7min 2026-10-05 · —