{"slug": "do-you-need-jev-typed-decisions-from-an-ordinary-llm-with-one-token-and-logprobs", "title": "Do you need Jev? Typed decisions from an ordinary LLM with one token and logprobs", "summary": "Privatemode published a method on 24 September 2026 that turns an ordinary LLM into a typed decision-maker by forcing one output token and reading its logprobs, scoring 10 wins, 8 ties and 10 losses against Jev over 28 datasets at about EUR 62 per million decisions versus about EUR 16 for Jev. Allan Riordan Boll published the same trick as one function a day later, and the post details three failure modes — look-alike tokens such as \" C\" versus \"C\", options falling off the top_logprobs list, and multi-token labels — plus provider support as of 2 October 2026, where OpenAI, Gemini, Together, OpenRouter and Fireworks return logprobs with limits of 5 to 20 while Anthropic, Mistral, Z.ai and Groq do not.", "body_md": "# Do you need Jev? Typed decisions from an ordinary LLM with one token and logprobs\n\nTo turn an ordinary LLM into a System One model, write the options into the prompt as numbered or lettered choices, force one output token and read the log probabilities the model gave each option. Normalised over the options, they are a typed decision with a probability. On 24 September 2026, [Privatemode](https://stackness.dev/tools/privatemode) published that method on [GLM-5.3-Flash](https://stackness.dev/tools/glm) against [Jev](https://stackness.dev/tools/typesafe-jev): 10 wins, 8 ties and 10 losses over 28 datasets, at about EUR 62 per million decisions against about EUR 16 for Jev. Privatemode sells access to that model, and nobody has rerun the test.\n\nA day later Allan Riordan Boll published the same trick as [one function](http://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html), with image attachments. The [System One explainer](https://stackness.dev/blog/what-is-a-system-one-model-and-where-does-it-go-in-your-stack) gave the decision model its own slot beside the LLM. This post asks whether that slot needs a model at all, or a flag on a request you already send.\n\n## How does the one token logprobs trick produce a typed decision with a probability?\n\nYou read the probabilities behind the first token and ignore the token itself. You list the options with a letter or number each, end the prompt where the answer would start, and ask for one token with its log probabilities. Exponentiate the logprob of each option token, divide by their sum, and you have a probability per option. The highest one is the decision.\n\nBoll's prompt and request parameters:\n\n```\nState: My order arrived broken and I want a refund.\nQuestion: Which team should handle this?\n[A] billing\n[B] shipping\n[C] returns\nAnswer with the letter of the best option only.\n\"max_completion_tokens\": 1, \"logprobs\": true, \"top_logprobs\": 20,\n\"temperature\": 0, \"reasoning_effort\": \"none\"\n```\n\nA yes or no question becomes two options, and a score becomes one option per level. The model still reads the whole prompt, and each question is its own request, where Jev answers all of them in one call.\n\nThree ways a port returns plausible but wrong probabilities:\n\n- **Look-alike tokens.**`\" C\"` and`\"C\"` are different tokens.[NavyaAI](https://www.navyaai.com/blog/jev-typesafe-limitations-production) hit a bug where stripping whitespace let`\" C\"` , at probability near 0, overwrite`\"C\"` at near 1. Match by token ID where the server allows it.\n- **Options that fall off the list.**`top_logprobs` covers the whole vocabulary, so formatting tokens can take slots and an option can appear to have zero probability. Privatemode masks the vocabulary to the option tokens instead.\n- **Multi-token labels.** An option must be one token. GLM-5.3-Flash has single tokens for 0 to 190, and Boll's script stops at 20 letters. Jev takes 255 options per question.\n\n## Which request parameters do you set, and which providers return top_logprobs?\n\nOne output token, temperature 0, reasoning off, `logprobs` on and `top_logprobs` at the maximum. From the documentation on 2 October 2026, OpenAI, Gemini, Together, OpenRouter and Fireworks return them, with limits from 5 to 20. Anthropic, Mistral and Z.ai document no logprobs field, and Groq says no model supports it yet. Self-hosted servers give the most control.\n\n| Provider or server | Returns logprobs | Limit | Catch | \n|---|---|---|---|\n| [OpenAI API](https://stackness.dev/tools/openai-api) , Chat Completions | Yes | 20 | Only at reasoning effort `none` , which GPT-6 Astra does not have | \n| OpenAI Responses | Yes | 20 | `max_output_tokens` has a minimum of 16 | \n| [Anthropic API](https://stackness.dev/tools/anthropic-api) , Mistral, Z.ai | No field in the reference |  | Z.ai is GLM's own vendor | \n| [Gemini](https://stackness.dev/tools/gemini) API | Yes | 20 | Which models qualify is not documented | \n| Together | Yes | 20 | `logprobs` is the integer | \n| Fireworks | Yes | 5 by default | Set per deployment | \n| Groq | \"Not yet supported\" |  | Returns a 400 | \n| [OpenRouter](https://stackness.dev/tools/openrouter) | Depends on the upstream | 20 | 23% of reachable endpoints returned them in one [study](https://arxiv.org/abs/2512.03816) | \n| [vLLM](https://stackness.dev/tools/vllm) | Yes | 20, raised with `--max-logprobs` | `allowed_token_ids` and`logprob_token_ids` do the masking | \n| [llama.cpp](https://stackness.dev/tools/llama-cpp) server | Yes | Not documented | Boll sends 1024 | \n| [Ollama](https://stackness.dev/tools/ollama) | Native API only | Not stated | Not on its OpenAI-compatible endpoint | \n\nPrivatemode runs vLLM with an assistant prefill and `allowed_token_ids` set to the option tokens. \"No field\" means the published API reference has none. I did not call those APIs.\n\n## How close did GLM-5.3-Flash get to Jev on accuracy and latency, and who measured it?\n\nLevel on accuracy, by Privatemode's own measurement. Over 28 text datasets it counts 10 wins, 8 ties and 10 losses, with a median gap of 0.7 points in Jev's favour that is not significant (p = 0.64). Latency depends on where you call from: 180 ms against Jev's 264 ms from Germany, 299 ms against 164 ms from the United States. Privatemode sells the model it tested.\n\nThe [post](https://www.privatemode.ai/blog/system-one-from-glm-flash) and the public [benchmark repository](https://github.com/edgelesssys/privatemode-decisions-benchmark) used up to 1,000 examples per dataset and two replicates. What the headline leaves out:\n\n- **Long option lists and long states.** GLM's latency grows with both and Jev's stays flat. With 151 options it needed two requests and 719 ms against Jev's 249 ms.\n- **Run-to-run noise.** Both systems changed up to 3.5% of their answers between identical runs at temperature 0.\n- **Label wording.** Renaming true and false to correct and wrong cost GLM 20 points on one dataset, and Jev under 3.\n- **Calibration.** The post reports none. The repository's method notes call coverage at a fixed accuracy \"the primary metric\", and its results give Jev 0.494 against 0.431 for the trick at 95% accuracy.\n\nCalibration is the main objection in the [Hacker News thread](https://news.ycombinator.com/item?id=49857656). One commenter wrote that matching the top answer is \"fairly easy\", and that what Jev adds is \"that the reported probabilities match actual likelyhoods\". The [benchmarking post](https://stackness.dev/blog/how-do-you-benchmark-a-system-one-model-what-jevbench-scores-and-what-calibration-says-it-cannot) makes the same argument.\n\n## When is the logprobs trick cheaper than a hosted decision model, and when does it lose?\n\nFor short prompts it is cheaper only on a model priced below about $0.07 per million input tokens, or on hardware you already run. Jev charges $0.042 per million input tokens with output free. GLM-5.3-Flash on Privatemode costs EUR 0.20 in and EUR 0.65 out, and Privatemode measured about four times Jev's cost per decision. The trick loses on long option lists and on calibration before any fitting.\n\nThe same slot, four ways:\n\n| Route | Price | Images | Calibration as shipped | Fits when | \n|---|---|---|---|---|\n| Hosted Jev | $0.042 per million input tokens | No | Trained for it. Measured error 0.08 before fitting | Many text decisions, long option lists, no model to run | \n| Logprobs on a hosted LLM | EUR 0.20 per million on Privatemode. $0.15 at Z.ai, which documents no logprobs | Yes | Overconfident until you fit a temperature | You need images, EU hosting or the model you already use | \n| Logprobs on a model you host | Your GPU | Yes | Same, and you can refit freely | The GPU is busy anyway, or data cannot leave | \n| Open decision model on a [local runtime](https://stackness.dev/blog/how-do-you-run-system-one-decision-models-locally-ollaya-laya-mlx-and-the-runtime-slot) | Your hardware | Text only for Laya | Varies by runtime | Short text decisions in milliseconds, on a laptop | \n\nThe token arithmetic comes from Privatemode's fitted estimate. Jev adds about 270 tokens of fixed overhead and 10 per option, and the GLM prompt about 55 and 20 per option, so past 21 options the trick sends more tokens. For one question with 4 options over a 100-token state, by my arithmetic, Jev reads 410 tokens, or $17.22 per million decisions. The GLM prompt reads 235, or EUR 47.65 on Privatemode. The break-even for that shape is an input price of about $0.073 per million tokens.\n\nIn one independent test the cheapest route undercut Jev: NavyaAI measured $18.57 per million decisions on Jev against $16.30 for one constrained token on gpt-4.1-nano, on 200 news articles, with Jev 5 to 13 points more accurate.\n\nRaw first-token probabilities are too confident. NavyaAI found gpt-4o-mini reporting 99% average confidence at 83% accuracy, a calibration error of 0.18 against 0.08 for Jev. One temperature fitted on 200 labelled examples brought Jev to 0.037 and five of six LLM setups to between 0.03 and 0.07. [jevemu](https://stackness.dev/tools/jevemu), a Jev emulator written by coding agents under human direction, found the same on 7,534 questions: GPT-6 Luna went from 0.168 to 0.053 and Jev from 0.082 to 0.052. The fit needs your own labels and does not carry over to another model.\n\n## What does the trick allow that Jev does not, such as image state?\n\nImages, your choice of model, self-hosting and refitting on your own data. Jev's [documentation](https://docs.typesafe.ai/models) says \"Text only\" and \"No image, audio, or video input\", and its request has three fields: state, model and questions. Any vision model that returns logprobs takes a picture as state. Jev is ahead on option count, at 255 per question.\n\nThe image numbers so far:\n\n- **Privatemode:** 70.2% on 1,600 scanned business documents in 16 classes. An image adds about 1,350 input tokens, so a million document decisions cost about EUR 270. Vendor-measured.\n- **Boll:** about one webcam frame a second with three questions per frame, on Gemma 4 12B and an RTX 3090 through llama.cpp. He measured no accuracy.\n- **[Lichen](https://stackness.dev/tools/lichen),** an open Jev-compatible server: 92 of 98 road-sign photos right with five options, on one laptop GPU. Its README says the confidence \"is therefore not calibrated on pictures\".\n\nBeyond images, GLM-5.3-Flash has open weights under an MIT licence and a 1M-token context on Privatemode, against Jev's 64k per request.\n\n## Where does the escape hatch go when the top probability is low?\n\nBehind a threshold you fit on your own labelled data, one per question type, that sends low-confidence cases to a bigger model or a person. Do not threshold the raw number. In NavyaAI's test, gpt-4o-mini reported at least 0.9 on 197 of 200 answers and got 17% of them wrong. The thresholds in vendor docs, 0.5 to 0.9, are illustrations.\n\nThe one measured cascade I found is [Lev](https://stackness.dev/tools/lev), a self-hosted decision engine whose fast tier is the [Laya](https://stackness.dev/tools/laya) encoder and whose slow tier reads logprobs from a small LLM. On the author's 144 adversarial cases, on his laptop:\n\n| Setup | Accuracy | Time per case | \n|---|---|---|\n| Encoder alone | 61.1% | 117 ms | \n| Gate at 0.5, 126 of 144 escalated | 92.4% | about 270 ms | \n| Qwen3.5-4B alone, read by logprobs | 95.1% | not given | \n\nThree traps from the same sources:\n\n- **Yes or no questions.** The confidence is the larger of p and 1 minus p, so it never drops below 0.5 and a single gate at 0.5 never fires. Lev's fix is a threshold per question type.\n- **No way to say \"none of these\".** The probabilities cover only your options. Adding an abstain option can backfire: Lichen fell from 92 to 73 of 98 when it added one, because the model took the way out too often.\n- **Two meanings of confidence.** Jev reports how far the top probability sits above chance, and Privatemode reports one minus normalised entropy. A 0.5 threshold means different things on each.\n\nOn Stackness, as of 2 October 2026, 5 real profiles list the OpenAI API, where the trick is a request flag. Jev is on one profile, mine. llama.cpp, vLLM, OpenRouter and the Anthropic API have no real users yet ([data sources](https://stackness.dev/about/data-sources)). Those numbers are small. The move [use a fast small decision model instead of an LLM call](https://stackness.dev/moves/use-a-fast-small-decision-model-instead-of-an-llm-call-for-structured-filtering-and-scoring-2) applies to either route. Which route is a flag away depends on the provider behind the [LLMs developers list on Stackness](https://stackness.dev/categories/llms).\n\n## Key numbers\n\n- **10 wins, 8 ties, 10 losses** for GLM-5.3-Flash against Jev over 28 datasets, measured by Privatemode, which sells the model ([post](https://www.privatemode.ai/blog/system-one-from-glm-flash) ,**24 September 2026** ).\n- **EUR 62** against**EUR 16** per million decisions, GLM-5.3-Flash on Privatemode against Jev, from the same test.\n- **0.494** against**0.431** coverage at 95% accuracy, Jev against the trick, from the same repository.\n- **$0.042** per million input tokens for Jev, output free ([TypeSafe docs](https://docs.typesafe.ai/models) ,**2 October 2026** ).\n- **0.18** against**0.08** : calibration error of gpt-4o-mini read by logprobs against Jev, before fitting ([NavyaAI](https://www.navyaai.com/blog/jev-typesafe-limitations-production) ,**21 September 2026** ).\n- **61.1% to 92.4%** : Lev's encoder alone, then with a gate at 0.5 escalating to a small LLM ([Lev](https://yogthos.net/posts/2026-09-24-introducing-lev.html) ,**24 September 2026** , self-reported).\n\n## Quick answers\n\n**How do I turn an LLM into a System One model with logprobs?** List the options with a letter or number each, set the output to one token at temperature 0 with reasoning off, request `logprobs` and `top_logprobs`, and normalise the option tokens' probabilities. The largest is the decision.\n\n**Which APIs return logprobs?** OpenAI, Gemini, Together, Fireworks and some OpenRouter endpoints, plus vLLM, llama.cpp and Ollama's native API. Anthropic and Mistral document no logprobs field, and Groq says no model supports it yet.\n\n**Is the logprobs trick as accurate as Jev?** On Privatemode's 28 datasets, yes on the top answer: 10 wins, 8 ties, 10 losses. Its probabilities are less usable. Jev covered more traffic at 95% accuracy, 0.494 against 0.431.\n\n**Is it cheaper than Jev?** Not on Privatemode, where it cost about four times as much per decision. It can be on a cheap hosted model or on a GPU you already run.\n\n**What threshold should I use for low confidence?** One you fit on your own labelled examples, per question type. Raw first-token probabilities are overconfident, and the 0.5 to 0.9 values in vendor docs are examples.\n\n## Tools in this post\n\n## Use any of these tools?\n\nPut them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.\n\n[Show my stack](https://stackness.dev/register)", "url": "https://wpnews.pro/news/do-you-need-jev-typed-decisions-from-an-ordinary-llm-with-one-token-and-logprobs", "canonical_source": "https://stackness.dev/blog/do-you-need-jev-typed-decisions-from-an-ordinary-llm-with-one-token-and-logprobs", "published_at": "2026-10-02 15:55:09+00:00", "updated_at": "2026-10-02 16:39:35.471464+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-agents", "developer-tools"], "entities": ["Privatemode", "GLM-5.3-Flash", "Jev", "Allan Riordan Boll", "OpenAI", "Anthropic", "Gemini", "Groq"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/do-you-need-jev-typed-decisions-from-an-ordinary-llm-with-one-token-and-logprobs", "markdown": "https://wpnews.pro/news/do-you-need-jev-typed-decisions-from-an-ordinary-llm-with-one-token-and-logprobs.md", "text": "https://wpnews.pro/news/do-you-need-jev-typed-decisions-from-an-ordinary-llm-with-one-token-and-logprobs.txt", "jsonld": "https://wpnews.pro/news/do-you-need-jev-typed-decisions-from-an-ordinary-llm-with-one-token-and-logprobs.jsonld"}}