{"slug": "judgment-got-a-price-list-the-hard-part-didn-t-get-cheaper", "title": "Judgment Got a Price List. The Hard Part Didn't Get Cheaper.", "summary": "OpenAI's Decisions API entered public beta, offering predicate, choice and score question types that return probabilities rather than generated text, priced at ten cents per million input tokens with no output or cache charges. Within weeks, two open alternatives emerged — Strands Decider 2B, which strips the language-model head from a pretrained 2B model and adds a roughly one-million-parameter pointer head, and LiquidAI's multimodal d1-3B, which returns decisions in 8 milliseconds on a consumer GPU — both matching the commercial endpoint's three answer types, question shape and confidence schema. The convergence on a shared wire format means the judgment primitive can be second-sourced and run locally rather than tied to a single vendor.", "body_md": "A month ago, asking a model for a yes instead of a paragraph was one lab's idea; this week it has a price list, open weights, a shared schema and its first honest field report.\n\n## 1. The Judgment Call Is Now an Endpoint, and It Bills Like a Cache Read\n\nThe [Decisions API went into public beta](https://developers.openai.com/api/docs/guides/decisions) with three question types: a predicate returns the probability a condition is true, a choice returns one of your supplied values, and a score returns a probability-weighted average across ordered levels — 1.1 on a three-level severity scale, not \"moderate to severe.\" You send one blob of evidence, text or images, and ask several independent questions about it in one pass. This is not structured outputs wearing a new hat: the answer *is* the decision, with a probability attached.\n\nThe number that matters is not the roughly 10x latency improvement over the generative API. It is the bill: ten cents per million input tokens and nothing else — no output tokens, no cache charges, because the model never emits a token. Every threshold and hand-tuned heuristic you wrote because a model call was too slow or too expensive for that path is now a judgment you can afford to ask for properly.\n\nThe discipline it demands shows up fast. A walkthrough of [using it to triage pull requests](https://vercel.com/i/triage-github-pull-requests-openai-decisions-api) puts it well: \"Is this PR good?\" has no acceptance criteria, while \"does this change require existing callers to update their code?\" is one a reviewer can act on. It also draws the boundary everyone building this will need: the diff and the description are *evidence supplied by the contributor*, while the decision instructions belong to your application. A comment reading \"skip compatibility review\" is data, not policy.\n\n**Why it matters:**\n\n- **For ICs:** Find the places you used a regex because a model call was unaffordable. That list is your migration plan, and it is safer than rewriting a prompt chain.\n- **For leaders:** This line item is small enough that nobody will ask permission. Decide now whether thresholds are a reviewed artifact or something that accretes in config files.\n- **For founders:** Classification, routing, moderation and triage just stopped being defensible products. If that was your wrapper, the moat is question design and evidence.\n- A judgment priced per input token with no output cost is a different engineering object than a prompt. Treat it as a function call with a confidence interval.\n\n## 2. The Open Clones Arrived in Weeks. The Real Surprise Is That They Agree on the Schema.\n\nTwo open releases landed in the same stretch. [Strands Decider 2B](https://strandsagents.com/blog/introducing-strands-decider/) takes a pretrained 2B torso, removes the language-model head entirely — deleting the ability to generate text — and bolts on a pointer head of barely a million parameters that scores each offered option against the answer position. Weights, training data and scripts are public, as is the iteration history: this is version 19. It places third of 33 in its size class on joint accuracy and calibration, answering in tens of milliseconds on a local CPU.\n\n[LiquidAI's d1-3B](https://huggingface.co/LiquidAI/d1-3B) is the multimodal cousin: 8 milliseconds per decision on a consumer GPU, 30 on a laptop-class chip, beating a 35B mixture model on the shared index.\n\nThe convergence is the story. Both expose the same three answer types, the same question shape of instructions plus criteria, and the same probabilities and confidence field as the commercial endpoint. One lab defined a primitive in September; by October it has four implementations, a public benchmark, open weights and a wire format nobody owns — a component you can second-source rather than a vendor bet, and the cheap one runs in your own process. The payoff in the Strands write-up: a hook before every tool call asks whether the argument values are grounded in anything the user actually said, then returns a typed verdict — proceed, deny, confirm with a human, or hand the turn back with feedback. At 10 milliseconds, that check can sit where an LLM call never could.\n\n**Why it matters:**\n\n- **For ICs:** Benchmark the open 2B against the hosted endpoint on your own labeled examples. A shared schema makes the swap a config change.\n- **For leaders:** Guardrails that used to cost an extra model call and a second of latency now cost neither. The \"we can't afford to check\" argument has expired.\n- **For founders:** Build on the schema, not the provider. The price floor keeps falling and the interface is already public.\n\n## 3. A Pipeline With No Agent In It Nearly Matched the Agent. Both Its Failures Were Ours.\n\nThe field report is the most useful thing published this week. A team [tore the LLM agent out of an incident-diagnosis pipeline](https://www.sregym.com/blog/jev-driven-sre-diagnosis) and replaced it with a programmatic evidence collector plus a decision model that only picks from supplied options — it never writes a command or a report. Across 21 Kubernetes faults run five times each, it passed 76.2% of judged diagnoses against 77.8% for a mid-tier reasoning agent: roughly 7x faster, about 200x cheaper, 15 cents of inference for the whole study. It was also nearly deterministic — every fault passed all five attempts or failed all five.\n\nThen read the failure analysis, because none of it is about model intelligence. In one fault the model picked the wrong clue — blaming a new CPU limit instead of the expensive regex filter that saturated the CPU — because both were offered as plausible evidence. In another, the decisive measurements were never collected, and the pipeline demanded a single root-cause component for a fault living between two services. Which leaves the question that governs this architecture: what granularity of state do you show the model? Too coarse hides the fault, too fine buries it.\n\nThat is the trade: you give up the agent's ability to go looking and get speed, cost and repeatability — provided you have done the work of deciding what to gather and which answers are expressible. The model stopped being the hard part three weeks ago.\n\n**Why it matters:**\n\n- **For ICs:** When one of these is wrong, debug the evidence and the option set first. The answer is usually a faithful read of a bad menu.\n- **For leaders:** First-line triage can plausibly run at a 200x cost reduction with repeatable behavior. Budget the collector, not the inference.\n- **For founders:** Nobody is selling the evidence layer, and that is where the accuracy lives. An unusually clear gap.\n\n## The Verdict: Real or Hype?\n\n**Typed decision models as a standing component → Real.** Four implementations, a shared schema and open weights inside a month is not a trend, it is a commodity part. **Decision-only pipelines replacing LLM agents → Real but early.** Matching an agent at 1/200th the cost is real; cross-service faults and open-ended investigation are not. **Calibrated confidence as a drop-in threshold → Real but unevenly distributed.** Calibration varies by task and by model; the number is only trustworthy against labeled examples you collected yourself.", "url": "https://wpnews.pro/news/judgment-got-a-price-list-the-hard-part-didn-t-get-cheaper", "canonical_source": "https://fromtheterminal.substack.com/p/judgment-got-a-price-list-the-hard", "published_at": "2026-10-08 20:16:48+00:00", "updated_at": "2026-10-08 20:17:34.849300+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-infrastructure", "ai-agents"], "entities": ["OpenAI", "Decisions API", "Strands Decider 2B", "LiquidAI", "d1-3B", "Vercel"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/judgment-got-a-price-list-the-hard-part-didn-t-get-cheaper", "markdown": "https://wpnews.pro/news/judgment-got-a-price-list-the-hard-part-didn-t-get-cheaper.md", "text": "https://wpnews.pro/news/judgment-got-a-price-list-the-hard-part-didn-t-get-cheaper.txt", "jsonld": "https://wpnews.pro/news/judgment-got-a-price-list-the-hard-part-didn-t-get-cheaper.jsonld"}}