Decisions as a Service A developer built two open-source tools on top of Jev, a new LLM from TypeSafe that returns calibrated probabilities for yes/no decisions instead of generating text. One tool, bscheck, scores documents for jargon and other forms of vague language, while jev-sort uses Jev's choice primitive as a comparator to perform semantic sorting of lists. The developer reports Jev is cheap and fast and expects it to simplify business integrations such as labelling and triage. Jev https://typesafe.ai/ is a new type of LLM. It doesn’t produce text; it makes decisions and returns calibrated probabilities. For “flow chart” type applications, it’s the bit of intelligence in the middle. And not only that, but it’s also cheap as chips and super-fast. Let’s explore what we can do. Noul questions I don’t know about you, but there’s a lot more bullshit https://fffej.substack.com/p/bullshit-detectors flying around. What would be good is if we could somehow feed a document in, and get the BS highlighted. We just need to make decisions about whether the different types of BS, with questions like: Does the text use impressive-sounding terminology in place of concrete meaning? Exclude legitimate technical language with a clear meaning in context. This is perfect for Jev, we fire off a body like this: { "state": "We leverage synergistic paradigms to unlock holistic excellence.", "questions": { "jargon": { "type": "noul", "instructions": "Does the text use impressive-sounding terminology in place of concrete meaning? Exclude legitimate technical language with a clear meaning in context." } } } And get a calibrated response: { "answers": { "jargon": { "type": "noul", "noul": 0.98 } }, "usage": { "input tokens": 120, "output tokens": 20 } } “noul” is the term used by TypeSafe for a yes/no decision apparently, it’s short for Bernoulli . So, in this case, the state we sent ”we leverage synergistic …” scores 0.98 98% for the particular question we sent. The API allows you to send multiple questions at once, so building a bullshit detector is rather easy. See here https://github.com/fffej/bscheck/ . Sorting lists Sorting things is one of the most studied problems in computer science. For most sorting algorithms you just need a comparator. We can write a comparator in jev using the choice https://docs.typesafe.ai/primitives/choice primitive. But for comparators to be useful, they have to be consistent. To do this, we just give both questions to jev and do some cross-checking in code. "state": {}, "questions": { "pair 0 1": { "type": "choice", "instructions": { "question": "Which comes first?", "A": "chicken", "B": "egg" }, "criteria": { "A": "A comes first.", "B": "B comes first.", "tie": "Neither comes first; they are equal in order." } and vice-versa with the choice the other way around. We get a response with confidence out the other end too: "answers": { "pair 0 1": { "type": "choice", "choice": "B", "probabilities": {"A": 0.0, "B": 1.0, "tie": 0.0}, "confidence": 1.0 }, "pair 1 0": { "type": "choice", "choice": "A", "probabilities": {"A": 1.0, "B": 0.0, "tie": 0.0}, "confidence": 1.0 } } So now we can do a semantic sort by just using this comparator. sort five,three,four,2,one = "one", "2", "three", "four", "five" sort "The Great Emu War, The Signing of the Magna Carta, The Fall of Constantinople" = "The Signing of the Magna Carta", "The Fall of Constantinople", "The Great Emu War" And in cases where the comparator is unsure, we just return with an error. My favourite so far: uv run jev-sort "regret,tequila,2am kebab,confidence" "tequila", "confidence", "2am kebab", "regret" Full code here https://github.com/fffej/jev-sort/ . Conclusion? Jev is fun to use, and I can imagine it making business integrations such as labelling, triage and other such exciting tasks dead simple. I suspect it’ll be very successful