"12 million views for a JSON classifier? Yeah, we're in a bubble."
That tweet was Niels Rogge's, the Hugging Face guy. Half of X was losing its mind over a model that answers yes/no questions. The other half was annoyed anyone cared.
The launch was September 15, 2026. Two years in stealth. A founder from OpenAI.
Locked weights, waitlist access.
Jev does one thing: send it text plus typed questions, get probabilities back. Choice, score, yes/no. No sentences, no tokens, nothing to parse.
And the reaction was huge. The HN thread passed 1,900 points with 520 comments.
Then something more interesting happened. People started rebuilding it.
Within 24 hours the first clones appeared. AINews counted six in two days.
The awesome-jev list now tracks more than a dozen. Laya, the biggest, passed 19,000 stars and added 5,000 in one day.
I kept refreshing GitHub, wondering why a JSON classifier gets this much energy.
The answer says something about how we feel about hosted AI. If you've ever paid an LLM to write text you immediately parse into an if statement, this fight is about you.
Here's a question people always ask: what is a System One model?
The shape is simple. You send state, any text, email or JSON document, with questions attached. Each question declares its answer type.
{
"state": "My payouts have failed three times.",
"questions": {
"queue": {"type": "choice", "instructions": "Which team handles this?"},
"escalate": {"type": "noul", "instructions": "Needs urgent human attention?"}
}
}
The model reads it once and answers everything in parallel. It returns probabilities, not prose.
Jev's founder is Diogo Almeida, who helped build the instruction work that became ChatGPT at OpenAI. Two years in stealth, and the pitch is sharp.
think of Jev as a frontier-intelligence function call.
No string generation means no hallucinated JSON, no parsing, no repair step. The model can't hallucinate. It also can't write a sentence.
Input runs $0.042 per million tokens, output is free, and responses land in 70 to 500 milliseconds.
A frontier chat model takes 3 to 329 seconds for the same job. That gap is the whole story.
What happens after a launch like this is predictable. Closed weights and no technical paper get read as homework.
A researcher, Archer Hume, probed the public API, watching latency scale with context and answers shift when questions reordered.
"Jev's Architecture Unmasked" came out of it: a causal transformer, probably sparse MoE, that encodes state once, runs every question branch in parallel, and reads probabilities straight off internal representations. He admits it's speculative.
It was also the clearest picture anyone had.
That post became a blueprint. The bubble takes were loud. But the builders were louder.
Most tutorials tell you to fine-tune a model for this. Kev shows the other way.
Kev comes from Jared Palmer, the Turborepo developer. It's a family of small decision models on Qwen3.5 and Qwen3.8, following the unmasked architecture. Apache-2.0, four sizes, from a 0.8B for a laptop to a 27B for a data centre GPU.
The smart move is the API. Kev matches TypeSafe's /v1/systemone contract, so you point TypeSafe's own Python SDK at your local server and nothing changes. That's how you take a closed product's users the polite way.
Then the evals. The part i respect most.
Kev-27B lands within a point of Jev on new sources, 0.851 against 0.857. And the README says it plainly: this isn't a controlled comparison, because we don't know what Jev was trained on.
It also lists where Kev loses. On MMLU-Pro it scores 0.675 against Jev's 0.840.
On day-precision date math, the small models trail.
That honesty is rarer than the weights.
The problem isn't what you think it is. It's confidence.
Laya's own research found the English checkpoint scored 0.000 accuracy on Khmer while reporting 0.952 confidence. A model that is certain and completely wrong.
If you branch production code on that, you ship silent failures. Laya fixes it with a router that detects the script before the forward pass and switches to a multilingual checkpoint.
Twenty-two alphabets, under half a millisecond.
Laya is the star of this wave. 421 million parameters on a ModernBERT backbone, pip install laya, runs on a CPU box, answers in roughly 21 to 33 milliseconds.
It shipped September 18, three days after Jev, from Convai Innovations. Its founder says he published the core idea first, in a March 2025 arXiv paper, then built the open version instead of staying bitter.
"Instead of staying bitter, I decided to take everything I learned, fix every architectural limitation of the old approach, and build a completely open, horizontal System 1 decision model family."
The history is genuinely contested, and one review called the David and Goliath framing marketing-adjacent. Fair.
So is the benchmark catch. The headline accuracy number for Laya comes from a checkpoint fine-tuned on the benchmark's own training split.
Zero-shot, the base model scores below the majority-class baseline. Out of the box it is a base to fine-tune, not a drop-in.
| Weights | Out-of-box accuracy | Price | |
|---|---|---|---|
| Jev | Closed | Highest | $0.042/MTok input |
| Kev | Apache-2.0 | Within a point of Jev (27B) | Free |
| Laya | Apache-2.0 | Near chance, fine-tune it | Free |
I used to think cloning a model took a lab. Then this weekend happened.
The interface is public, so the wave moved fast. The weights and data are not, so quality lagged.
SemIf trains nothing at all. It reads answer logits off a frozen Qwen3.5-4B and ran 21 decisions five times faster than generated JSON.
NanoJev is a 0.6B model for real-time loops and beat Jev on ViZDoom Basic, 128 out of 128 against 56. jevlike ships a trainer instead of a model.
Here's what the wave means:
On an independent 49-task benchmark, Jev still scores 0.966 macro accuracy. The best open entrant scores 0.704.
The clones win on speed and price. The scoreboard is not close yet.
The part i couldn't stop watching was Flappy Bird. Laya plays it live, about 30 decisions a second, by asking itself one question: where is the bird?
Asked which way to move, every checkpoint answered backwards. Asked where the bird is, clean graded answers.
P(below) of 0.95, 0.82, 0.06. The game flaps when the probability crosses half.
Same story in Tetris. 1,799 decisions in a minute, 52 lines cleared, no top-out.
It loses eventually, because the pieces fall faster and a life lasts about two minutes. And it cannot read numbers at all.
Give it two altitudes and no checkpoint can say which is lower. Do the arithmetic in code, hand the model the conclusion in words.
That's the weirdest part of this whole genre. The question is the craft.
Reword it and accuracy swings by 30 points. I spent an evening rephrasing prompts to watch the probability move.
My partner watched me do this and asked if it was work. I said yes, and i wasn't sure.
Most people should not run a decision model this week.
The open alternatives win on price, latency and control. Not accuracy. Not yet.
If you route a thousand tickets a day and your vendor doubles the price, Laya or Kev is your insurance.
If you route 50 tickets a day, use a folder and an if statement. Seriously. A 4GB model for your inbox is overkill.
You'll pay for the wiring with your weekend.
If you do pick one, plan to fine-tune it. Laya's base checkpoints score near chance without it.
Kev's own docs tell you where it loses, so read those before you trust a threshold. And the cheap irony: Jev is not expensive.
$0.042 per million tokens is basically nothing. Free is a win for privacy and control, not your invoice.
The real take: a model is not a moat. An eval set with your data in it is.
I keep coming back to the Khmer number. 0.000 accuracy, 0.952 confidence. A model that was certain and catastrophically wrong, because nobody checked what the words meant.
The open source wave matters because it makes checking possible. Weights you can audit. Datasets you can read.
Evals someone else can rerun. Kev and Laya are not better than Jev yet, and they will tell you that themselves, which is exactly the point.
i still don't know who invented the idea. But i know who lets me check their work. That's the one i'm building with.