{"slug": "kev-and-laya-the-open-source-answer-to-typesafe-s-jev", "title": "Kev and Laya: The Open Source Answer to TypeSafe's Jev", "summary": "Two open-source projects, Kev and Laya, emerged within days of TypeSafe's closed-weight Jev launch, with Laya surpassing 19,000 GitHub stars and Kev matching Jev within a point on new sources (0.851 vs 0.857) while trailing badly on MMLU-Pro (0.675 vs 0.840). Kev, from Turborepo developer Jared Palmer, ships Apache-2.0 decision models on Qwen3.5/3.8 that match TypeSafe's /v1/systemone API contract, while Laya's research found its English checkpoint reported 0.952 confidence on Khmer while scoring 0.000 accuracy.", "body_md": "\"12 million views for a JSON classifier? Yeah, we're in a bubble.\"\n\nThat tweet was Niels Rogge's, the Hugging Face guy. Half of X was losing its mind over a model that answers yes/no questions. The other half was annoyed anyone cared.\n\nThe launch was September 15, 2026. Two years in stealth. A founder from OpenAI.\n\nLocked weights, waitlist access.\n\nJev does one thing: send it text plus typed questions, get probabilities back. Choice, score, yes/no. No sentences, no tokens, nothing to parse.\n\nAnd the reaction was huge. The HN thread passed 1,900 points with 520 comments.\n\nThen something more interesting happened. People started rebuilding it.\n\nWithin 24 hours the first clones appeared. AINews counted six in two days.\n\nThe awesome-jev list now tracks more than a dozen. Laya, the biggest, passed 19,000 stars and added 5,000 in one day.\n\nI kept refreshing GitHub, wondering why a JSON classifier gets this much energy.\n\nThe answer says something about how we feel about hosted AI. If you've ever paid an LLM to write text you immediately parse into an if statement, this fight is about you.\n\nHere's a question people always ask: what is a System One model?\n\nThe shape is simple. You send state, any text, email or JSON document, with questions attached. Each question declares its answer type.\n\n```\n{\n  \"state\": \"My payouts have failed three times.\",\n  \"questions\": {\n    \"queue\": {\"type\": \"choice\", \"instructions\": \"Which team handles this?\"},\n    \"escalate\": {\"type\": \"noul\", \"instructions\": \"Needs urgent human attention?\"}\n  }\n}\n```\n\nThe model reads it once and answers everything in parallel. It returns probabilities, not prose.\n\nJev's founder is Diogo Almeida, who helped build the instruction work that became ChatGPT at OpenAI. Two years in stealth, and the pitch is sharp.\n\nthink of Jev as a frontier-intelligence function call.\n\nNo string generation means no hallucinated JSON, no parsing, no repair step. **The model can't hallucinate. It also can't write a sentence.**\n\nInput runs $0.042 per million tokens, output is free, and responses land in 70 to 500 milliseconds.\n\nA frontier chat model takes 3 to 329 seconds for the same job. That gap is the whole story.\n\nWhat happens after a launch like this is predictable. Closed weights and no technical paper get read as homework.\n\nA researcher, Archer Hume, probed the public API, watching latency scale with context and answers shift when questions reordered.\n\n\"Jev's Architecture Unmasked\" came out of it: a causal transformer, probably sparse MoE, that encodes state once, runs every question branch in parallel, and reads probabilities straight off internal representations. He admits it's speculative.\n\nIt was also the clearest picture anyone had.\n\nThat post became a blueprint. The bubble takes were loud. But the builders were louder.\n\nMost tutorials tell you to fine-tune a model for this. Kev shows the other way.\n\nKev comes from Jared Palmer, the Turborepo developer. It's a family of small decision models on Qwen3.5 and Qwen3.8, following the unmasked architecture. Apache-2.0, four sizes, from a 0.8B for a laptop to a 27B for a data centre GPU.\n\nThe smart move is the API. Kev matches TypeSafe's `/v1/systemone` contract, so you point TypeSafe's own Python SDK at your local server and nothing changes. That's how you take a closed product's users the polite way.\n\nThen the evals. The part i respect most.\n\nKev-27B lands within a point of Jev on new sources, 0.851 against 0.857. And the README says it plainly: this isn't a controlled comparison, because we don't know what Jev was trained on.\n\nIt also lists where Kev loses. On MMLU-Pro it scores 0.675 against Jev's 0.840.\n\nOn day-precision date math, the small models trail.\n\n**That honesty is rarer than the weights.**\n\nThe problem isn't what you think it is. It's confidence.\n\nLaya's own research found the English checkpoint scored 0.000 accuracy on Khmer while reporting 0.952 confidence. A model that is certain and completely wrong.\n\nIf you branch production code on that, you ship silent failures. Laya fixes it with a router that detects the script before the forward pass and switches to a multilingual checkpoint.\n\nTwenty-two alphabets, under half a millisecond.\n\nLaya is the star of this wave. 421 million parameters on a ModernBERT backbone, `pip install laya`, runs on a CPU box, answers in roughly 21 to 33 milliseconds.\n\nIt shipped September 18, three days after Jev, from Convai Innovations. Its founder says he published the core idea first, in a March 2025 arXiv paper, then built the open version instead of staying bitter.\n\n\"Instead of staying bitter, I decided to take everything I learned, fix every architectural limitation of the old approach, and build a completely open, horizontal System 1 decision model family.\"\n\nThe history is genuinely contested, and one review called the David and Goliath framing marketing-adjacent. Fair.\n\nSo is the benchmark catch. The headline accuracy number for Laya comes from a checkpoint fine-tuned on the benchmark's own training split.\n\nZero-shot, the base model scores below the majority-class baseline. Out of the box it is a base to fine-tune, not a drop-in.\n\n|  | Weights | Out-of-box accuracy | Price | \n|---|---|---|---|\n| Jev | Closed | Highest | $0.042/MTok input | \n| Kev | Apache-2.0 | Within a point of Jev (27B) | Free | \n| Laya | Apache-2.0 | Near chance, fine-tune it | Free | \n\nI used to think cloning a model took a lab. Then this weekend happened.\n\nThe interface is public, so the wave moved fast. The weights and data are not, so quality lagged.\n\nSemIf trains nothing at all. It reads answer logits off a frozen Qwen3.5-4B and ran 21 decisions five times faster than generated JSON.\n\nNanoJev is a 0.6B model for real-time loops and beat Jev on ViZDoom Basic, 128 out of 128 against 56. jevlike ships a trainer instead of a model.\n\nHere's what the wave means:\n\nOn an independent 49-task benchmark, Jev still scores 0.966 macro accuracy. The best open entrant scores 0.704.\n\nThe clones win on speed and price. The scoreboard is not close yet.\n\nThe part i couldn't stop watching was Flappy Bird. Laya plays it live, about 30 decisions a second, by asking itself one question: where is the bird?\n\nAsked which way to move, every checkpoint answered backwards. Asked where the bird is, clean graded answers.\n\nP(below) of 0.95, 0.82, 0.06. The game flaps when the probability crosses half.\n\nSame story in Tetris. 1,799 decisions in a minute, 52 lines cleared, no top-out.\n\nIt loses eventually, because the pieces fall faster and a life lasts about two minutes. And it cannot read numbers at all.\n\nGive it two altitudes and no checkpoint can say which is lower. Do the arithmetic in code, hand the model the conclusion in words.\n\nThat's the weirdest part of this whole genre. The question is the craft.\n\nReword it and accuracy swings by 30 points. I spent an evening rephrasing prompts to watch the probability move.\n\nMy partner watched me do this and asked if it was work. I said yes, and i wasn't sure.\n\nMost people should not run a decision model this week.\n\nThe open alternatives win on price, latency and control. Not accuracy. Not yet.\n\nIf you route a thousand tickets a day and your vendor doubles the price, Laya or Kev is your insurance.\n\nIf you route 50 tickets a day, use a folder and an if statement. Seriously. A 4GB model for your inbox is overkill.\n\nYou'll pay for the wiring with your weekend.\n\nIf you do pick one, plan to fine-tune it. Laya's base checkpoints score near chance without it.\n\nKev's own docs tell you where it loses, so read those before you trust a threshold. And the cheap irony: Jev is not expensive.\n\n$0.042 per million tokens is basically nothing. Free is a win for privacy and control, not your invoice.\n\nThe real take: a model is not a moat. An eval set with your data in it is.\n\nI keep coming back to the Khmer number. 0.000 accuracy, 0.952 confidence. A model that was certain and catastrophically wrong, because nobody checked what the words meant.\n\nThe open source wave matters because it makes checking possible. Weights you can audit. Datasets you can read.\n\nEvals someone else can rerun. Kev and Laya are not better than Jev yet, and they will tell you that themselves, which is exactly the point.\n\ni still don't know who invented the idea. But i know who lets me check their work. That's the one i'm building with.", "url": "https://wpnews.pro/news/kev-and-laya-the-open-source-answer-to-typesafe-s-jev", "canonical_source": "https://dev.to/dishant0406/kev-and-laya-the-open-source-answer-to-typesafes-jev-15o4", "published_at": "2026-10-05 21:03:54+00:00", "updated_at": "2026-10-05 21:18:00.515992+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-research", "ai-products", "developer-tools"], "entities": ["TypeSafe", "Jev", "Kev", "Laya", "Jared Palmer", "Diogo Almeida", "Niels Rogge", "Hugging Face"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/kev-and-laya-the-open-source-answer-to-typesafe-s-jev", "markdown": "https://wpnews.pro/news/kev-and-laya-the-open-source-answer-to-typesafe-s-jev.md", "text": "https://wpnews.pro/news/kev-and-laya-the-open-source-answer-to-typesafe-s-jev.txt", "jsonld": "https://wpnews.pro/news/kev-and-laya-the-open-source-answer-to-typesafe-s-jev.jsonld"}}