OpenAI launched the Decisions API in public beta on October 6 β a dedicated endpoint that returns typed answers, not generated text. Feed it a customer complaint and a list of departments, and it routes the ticket in ~150ms at $0.10 per million input tokens with zero output charges. The argument is straightforward: stop using a language model for questions that have a fixed set of answers.
Three Types of Decisions #
The Decisions API supports three structured output types, each mapped to a concrete classification problem:
- predicate β returns a probability (0.0β1.0) that a statement is true. Use it to detect spam, flag PII, or check whether a message violates a policy.
- choice β picks one item from a predefined list and returns a confidence score. Use it to route support tickets, label intent, or assign categories.
- score β places an input on an ordered scale you define. Use it to rate ticket urgency, severity, or quality.
None of these return prose. You define the options; the model picks from them. That constraint is the point.
What It Looks Like in Code #
Here is a choice question routing a customer complaint:
import openai
client = openai.OpenAI()
response = client.decisions.create(
model="gpt-6-luna",
input="I was charged twice for my subscription last month",
questions=[
{
"type": "choice",
"question": "Which department should handle this complaint?",
"choices": [
{"id": "billing", "label": "Billing", "description": "Payments, invoices, and refunds"},
{"id": "technical", "label": "Technical Support", "description": "Problems using the product"},
{"id": "shipping", "label": "Shipping", "description": "Delivery and tracking"},
{"id": "other", "label": "Other", "description": "Requests outside these categories"}
]
}
]
)
department = response.decisions[0].choice.id # "billing"
confidence = response.decisions[0].choice.probability # 0.96
And a predicate checking for spam:
response = client.decisions.create(
model="gpt-6-luna",
input=comment_text,
questions=[
{
"type": "predicate",
"statement": "This comment is promotional spam or contains unsolicited links"
}
]
)
is_spam = response.decisions[0].probability > 0.8
Multiple independent questions can share one request β you can check for spam, route to a department, and score urgency in a single call. Dependent decisions (where the answer to one question affects the next) require separate requests.
When to Use This Instead of the Responses API #
The Responses API is not going away. The Decisions API is a narrower tool for a specific job. Here is the actual decision tree:
- Fixed set of answers (classify, route, score) β Decisions API
- Arbitrary JSON extraction (pull order details from a message) β Responses API with Structured Outputs
- Function call with parameters (look up an account) β Responses API with function calling
- Reasoning chain required β Responses API or an o-series model
Using a Responses API call for pure routing wastes roughly five times the cost and adds 800β1,500ms of latency you do not need. The Decisions API runs the same Luna model announced at DevDay 2026 but skips the generation step entirely.
The Price Case #
The API charges $0.10 per million input tokens and nothing for output tokens. GPT-6 Luna through the Responses API costs approximately $0.50 per million input tokens. For teams running thousands of classifications per minute, that is a five-to-one cost difference on the same underlying model, plus the latency win.
It is worth comparing this to the other typed-decision story from this week: TypeSafe AI raised $870 million for Jev, a model purpose-built to return typed business decisions with a reasoning trace. They are not directly competing. Jev is for complex multi-step decisions with explainability requirements. The Decisions API is for fast, high-volume classification where you do not need a reasoning chain.
What Is Missing in Beta #
The API is genuinely useful today, but a few gaps matter before building production workflows around it:
- One model only. gpt-6-luna is the only option. More models are expected at general availability.
- Images must be base64. Hosted HTTP/HTTPS image URLs are not supported β you must encode images inline.
- No documented rate limits. OpenAI has not published RPM or TPM caps. Assume standard Luna limits apply until GA.
- GA timeline. Public beta as of October 6, 2026. General availability expected within weeks.
The OpenAI community thread has the most current developer feedback and workarounds. For zero-shot classification with a controlled option set β and particularly for agent routing pipelines β the Decisions API is already the right call. The main reason to wait for GA is rate limit clarity.