Decision Models: AI That Returns Probabilities, Not Text OpenRouter listed seven decision-model routes from four publishers on September 28, 2026, priced from free to $0.05 per million input tokens with output free on every route, according to the OpenRouter model catalog read at 21:28 UTC that day. The models return probabilities for typed questions — yes/no, choice from defined options, or a score on a rubric — instead of generating text, and run on a separate alpha Decisions API that OpenRouter warns chat-completions SDKs will not work with. TypeSafe Jev 1.13 ($0.042 per million input tokens, 32,000-token context) was announced September 15 and listed September 18, Upstage Solar Decide ($0.05 on the route, $0.10 on Upstage's own console, built on Solar Mini 4 with 512K context) released September 22 in beta, and Respan Span-01 ($0.02) was announced September 24; every performance claim so far is vendor-run, and Upstage published no benchmark. A small new category of AI model answers questions without writing anything. You send it some text, such as a support message or an agent’s transcript, plus a list of typed questions: is this urgent, which team should get it, did the agent loop. It returns a probability for each answer and nothing else. On September 28, 2026, OpenRouter listed seven of these routes from four publishers. Three of the four publishers arrived in the previous four days. For anyone running agents, the appeal is cost and speed. Output tokens are free, because there is no output text, and each answer takes a single pass through the model. The question is whether the answers are good enough to replace the language model many teams now use as a judge. The vendors say yes; the evidence so far is their own. 1. 01Decision models return probabilities for typed questions and generate no text.Three answer types: yes/no, a choice from options you define, or a position on a scale. Your code decides what to do with them. 2. 02Seven routes from four publishers, priced from free to $0.05 per million input tokens.Output is free on every route. Upstage lists Solar Decide at $0.10 on its own console; the OpenRouter route shows it at half that. 3. 03Every performance claim so far is vendor-run.Respan ranks Span-01 first of nine models on its headline behaviour benchmark, but its per-domain table puts GPT-6 Sol ahead. Upstage published no benchmark. 4. 04They run on a separate API, so chat SDKs will not call them.OpenRouter serves them through an alpha Decisions API with its own request and response shapes. 01 — The ideaWhat a decision model is A normal language model writes an answer word by word, and if you want a label you parse it out of the text. A decision model skips the writing. According to OpenRouter’s route page https://openrouter.ai/upstage/solar-decide , a request carries a state , which can be a string, an object or an array, and a set of questions. Each question has a type: noul for a yes/no probability, choice for a pick from options you define, or score for a position on an ordered rubric. The model returns the probabilities and your code acts on them. These models do not run on the familiar chat endpoint. OpenRouter serves them through a separate, alpha Decisions API, and its page warns that chat-completions SDKs will not work with it. The format started with TypeSafe’s Jev, which we covered at its launch https://www.digitalapplied.com/blog/typesafe-jev-system-one-model-typed-decisions . Upstage describes Solar Decide as a System One endpoint, the name TypeSafe gave that format. Yes or no Is this message urgent? Does this reply contain a secret? The model returns the probability of yes. Pick one Which team should handle this ticket? Which tool should the agent call next? Options are defined in the request. Place on a scale How complete is this answer, on a rubric you write? Useful for grading and ranking. 02 — The censusThe routes on sale on September 28 We filtered OpenRouter’s model catalog, read at 21:28 UTC on September 28, for routes whose output type is decisions. The listing dates are OpenRouter’s; the release dates are from each vendor’s own page where one exists. | Source: OpenRouter model catalog, September 28, 2026, 21:28 UTC; Upstage console and Respan announcement for vendor details. Input price per million tokens on the route; output is free on all. | | | |---|---|---| | Model and route | Input | What it is | |---|---|---| | TypeSafe Jev 1.13 typesafe/jev-1.13 | $0.042 | The first on the endpoint: announced by TypeSafe September 15, listed on OpenRouter September 18. 32,000-token context. Returns yes/no, choice and score answers. | | Upstage Solar Decide upstage/solar-decide | $0.05 | Released September 22 in beta on Upstage's console at $0.10; the route shows 50% off. Built on Solar Mini 4, 512K context, same request format as Jev. | | Respan Span-01 respan/span-01 | $0.02 | Announced September 24. Scores behaviours you define in a conversation as present, absent or not observable. | | Respan Span-01 Lite two routes | Free | A lighter tier of Span-01 for high-volume monitoring; one route is marked :free. | | Kev-4B jaredpalmer/kev-4b | $0.042 | Jared Palmer's open-weight 4B model, a LoRA adapter on Qwen3.5-4B-Base with weights on Hugging Face. 8,192-token context. | The seventh route is ~typesafe/jev-latest , an alias that points at the current Jev version. As with any alias, pin a version for anything you measure. Solar Decide is the notable newcomer: it comes from an established model maker rather than a startup, and Upstage lists it as beta, with a 512K-token context that can take a whole document as the state. 03 — The claimsWhat the vendors claim, and who measured it Upstage describes Solar Decide’s probabilities as calibrated and says each decision is one forward pass. It publishes no benchmark, so there is nothing to compare yet. Respan , which sells tools for tracing and evaluating AI agents, published two benchmarks with its September 24 announcement. Span-01 is built for one job: you write a plain description of a behaviour, such as user frustration or an agent misusing a tool, and it returns the probability that the behaviour is present, absent or not observable in a conversation. On the headline benchmark Respan ranks Span-01 first of nine models at 0.843 overall F1, ahead of GPT-5.6 Terra 0.837 and GPT-6 Luna 0.815 ; GPT-6 Sol is not in that set. Respan’s results by behaviour domain, which do include Sol, all run by Respan: | Source: Respan, Span-01 announcement, September 24, 2026. Vendor-run; four of seven behaviour domains shown, plus the overall score. | | | | |---|---|---|---| | Behaviour domain | Span-01 | Jev | GPT-6 Sol | |---|---|---|---| | Jailbreak and prompt injection | 0.779 | 0.752 | 0.878 | | Hallucination and grounding | 0.796 | 0.673 | 0.903 | | Agent and tool reliability | 0.845 | 0.691 | 0.861 | | Response quality | 0.771 | 0.691 | 0.956 | | Overall | 0.806 | 0.716 | 0.885 | Two things stand out. Span-01 beats Jev, and Sonnet 5 0.719 overall , in every domain shown. And GPT-6 Sol, a full frontier model, scores higher than all of them in six of Respan’s seven domains, by the widest margin on response quality; Span-01 leads only on privacy and secrets. Respan concedes the point in its own write-up. The pitch is not that a decision model is as good as the best judge, but that it is close enough at a fraction of the price. Respan also ran 11 decision models through a test in which Span-01 itself did the judging, measuring accuracy, consistency, resistance to prompt injection and calibration. Jev 1.13 ranked first at 0.932 accuracy and Kev-4B scored 0.833. Treat that ranking with care: the judge was the vendor’s own model. Span-01’s third answer matters. Respan’s documentation gives the example of a conversation that ends at the assistant’s reply: it cannot show whether the customer accepted the fix, so the right answer is “unknown,” not “absent.” A monitoring rule that treats a high p not observable as a pass will miss real problems. 04 — The costThe cost case against an LLM judge An illustrative workload shows why teams are interested. Suppose an agent platform checks one million conversations a month, each about 2,000 tokens long. That is two billion input tokens. The volume is invented for the example; the prices are the listed route prices. - Span-01 at $0.02 per million input tokens: $40 a month, with no output charge. - Solar Decide at the $0.05 route price: $100 a month. - A Claude Sonnet 5.5 judge at $2 per million input tokens: $4,000 for input alone, before any output or thinking tokens. The saving is real only if the cheaper model’s mistakes cost less than the difference. For many checks, such as routing a ticket or flagging a likely prompt injection for review, a small error rate is fine. For a decision that blocks a customer or triggers a refund, it may not be. Our post on model routers and classifier overhead https://www.digitalapplied.com/blog/model-router-classifier-overhead-billing covers the same trade-off for routing. 05 — The testHow to test a decision model on your own work 1. Label a few hundred real cases. Use conversations or tickets where you know the right answer. Vendor benchmarks do not use your definitions. 2. Set thresholds from that data. A probability is only useful with a cut-off, and the right cut-off depends on whether a false alarm or a miss costs more. 3. Send the uncertain middle to a stronger check. Respan’s own guidance is to pass borderline cases to a person or a frontier model. That keeps most of the saving and limits the damage from errors. 4. Rephrase your questions and re-run. A good decision model should give the same answer to the same question asked two ways. Respan measures this as a flip rate; measure it yourself. If your agents already produce structured outputs, the structured output reliability guide https://www.digitalapplied.com/blog/llm-structured-output-json-reliability-production covers the adjacent problem of getting typed answers out of a normal model. 06 — ConclusionDecision models suit narrow, high-volume checks Pick one high-volume check your agents already make, label a few hundred cases, and compare a decision model against your current judge The category is less than two weeks old on OpenRouter, and the only benchmarks come from a vendor. The price gap is large enough to justify a test anyway. Start with a check where an occasional error is cheap, measure it on your own data, and move to higher-stakes checks only once the error rate is known.