{"slug": "openai-decisions-api-stop-using-llms-for-yes-no-questions", "title": "OpenAI Decisions API: Stop Using LLMs for Yes/No Questions", "summary": "OpenAI launched the Decisions API in public beta on October 6, a dedicated endpoint that returns typed answers instead of generated text and routes a customer complaint in about 150ms at $0.10 per million input tokens with no output charges. The API supports three structured output types — predicate (a 0.0–1.0 probability), choice (one item from a predefined list with a confidence score), and score (an ordered scale) — and OpenAI says using the Responses API for pure routing wastes roughly five times the cost and adds 800–1,500ms of latency, since GPT-6 Luna through the Responses API costs approximately $0.50 per million input tokens. The launch follows TypeSafe AI's $870 million raise for Jev, a model purpose-built to return typed decisions.", "body_md": "OpenAI launched the Decisions API in public beta on October 6 — a dedicated endpoint that returns typed answers, not generated text. Feed it a customer complaint and a list of departments, and it routes the ticket in ~150ms at $0.10 per million input tokens with zero output charges. The argument is straightforward: stop using a language model for questions that have a fixed set of answers.\n\n## Three Types of Decisions\n\nThe [Decisions API](https://developers.openai.com/api/docs/guides/decisions) supports three structured output types, each mapped to a concrete classification problem:\n\n- **predicate** — returns a probability (0.0–1.0) that a statement is true. Use it to detect spam, flag PII, or check whether a message violates a policy.\n- **choice** — picks one item from a predefined list and returns a confidence score. Use it to route support tickets, label intent, or assign categories.\n- **score** — places an input on an ordered scale you define. Use it to rate ticket urgency, severity, or quality.\n\nNone of these return prose. You define the options; the model picks from them. That constraint is the point.\n\n## What It Looks Like in Code\n\nHere is a choice question routing a customer complaint:\n\n``` python\nimport openai\n\nclient = openai.OpenAI()\n\nresponse = client.decisions.create(\n    model=\"gpt-6-luna\",\n    input=\"I was charged twice for my subscription last month\",\n    questions=[\n        {\n            \"type\": \"choice\",\n            \"question\": \"Which department should handle this complaint?\",\n            \"choices\": [\n                {\"id\": \"billing\", \"label\": \"Billing\", \"description\": \"Payments, invoices, and refunds\"},\n                {\"id\": \"technical\", \"label\": \"Technical Support\", \"description\": \"Problems using the product\"},\n                {\"id\": \"shipping\", \"label\": \"Shipping\", \"description\": \"Delivery and tracking\"},\n                {\"id\": \"other\", \"label\": \"Other\", \"description\": \"Requests outside these categories\"}\n            ]\n        }\n    ]\n)\n\ndepartment = response.decisions[0].choice.id        # \"billing\"\nconfidence = response.decisions[0].choice.probability  # 0.96\n```\n\nAnd a predicate checking for spam:\n\n```\nresponse = client.decisions.create(\n    model=\"gpt-6-luna\",\n    input=comment_text,\n    questions=[\n        {\n            \"type\": \"predicate\",\n            \"statement\": \"This comment is promotional spam or contains unsolicited links\"\n        }\n    ]\n)\n\nis_spam = response.decisions[0].probability > 0.8\n```\n\nMultiple independent questions can share one request — you can check for spam, route to a department, and score urgency in a single call. Dependent decisions (where the answer to one question affects the next) require separate requests.\n\n## When to Use This Instead of the Responses API\n\nThe Responses API is not going away. The Decisions API is a narrower tool for a specific job. Here is the actual decision tree:\n\n- **Fixed set of answers (classify, route, score)** → Decisions API\n- **Arbitrary JSON extraction (pull order details from a message)** → Responses API with Structured Outputs\n- **Function call with parameters (look up an account)** → Responses API with function calling\n- **Reasoning chain required** → Responses API or an o-series model\n\nUsing a Responses API call for pure routing wastes roughly five times the cost and adds 800–1,500ms of latency you do not need. The Decisions API runs the same [Luna model announced at DevDay 2026](https://openai.com/index/devday-2026-recap/) but skips the generation step entirely.\n\n## The Price Case\n\nThe API charges $0.10 per million input tokens and nothing for output tokens. GPT-6 Luna through the Responses API costs approximately $0.50 per million input tokens. For teams running thousands of classifications per minute, that is a five-to-one cost difference on the same underlying model, plus the latency win.\n\nIt is worth comparing this to the other typed-decision story from this week: TypeSafe AI raised $870 million for Jev, a model purpose-built to return typed business decisions with a reasoning trace. They are not directly competing. Jev is for complex multi-step decisions with explainability requirements. The Decisions API is for fast, high-volume classification where you do not need a reasoning chain.\n\n## What Is Missing in Beta\n\nThe API is genuinely useful today, but a few gaps matter before building production workflows around it:\n\n- **One model only.** gpt-6-luna is the only option. More models are expected at general availability.\n- **Images must be base64.** Hosted HTTP/HTTPS image URLs are not supported — you must encode images inline.\n- **No documented rate limits.** OpenAI has not published RPM or TPM caps. Assume standard Luna limits apply until GA.\n- **GA timeline.** Public beta as of October 6, 2026. General availability expected within weeks.\n\nThe [OpenAI community thread](https://community.openai.com/t/decisions-api-is-now-available-in-public-beta/1403877) has the most current developer feedback and workarounds. For zero-shot classification with a controlled option set — and particularly for agent routing pipelines — the Decisions API is already the right call. The main reason to wait for GA is rate limit clarity.", "url": "https://wpnews.pro/news/openai-decisions-api-stop-using-llms-for-yes-no-questions", "canonical_source": "https://byteiota.com/openai-decisions-api-stop-using-llms-for-yes-no-questions/", "published_at": "2026-10-10 01:09:43+00:00", "updated_at": "2026-10-10 01:57:51.339474+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-tools", "developer-tools"], "entities": ["OpenAI", "Decisions API", "Responses API", "GPT-6 Luna", "TypeSafe AI", "Jev"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openai-decisions-api-stop-using-llms-for-yes-no-questions", "markdown": "https://wpnews.pro/news/openai-decisions-api-stop-using-llms-for-yes-no-questions.md", "text": "https://wpnews.pro/news/openai-decisions-api-stop-using-llms-for-yes-no-questions.txt", "jsonld": "https://wpnews.pro/news/openai-decisions-api-stop-using-llms-for-yes-no-questions.jsonld"}}