# OpenAI Decisions API: Stop Using LLMs for Yes/No Questions

> Source: <https://byteiota.com/openai-decisions-api-stop-using-llms-for-yes-no-questions/>
> Published: 2026-10-10 01:09:43+00:00

OpenAI launched the Decisions API in public beta on October 6 — a dedicated endpoint that returns typed answers, not generated text. Feed it a customer complaint and a list of departments, and it routes the ticket in ~150ms at $0.10 per million input tokens with zero output charges. The argument is straightforward: stop using a language model for questions that have a fixed set of answers.

## Three Types of Decisions

The [Decisions API](https://developers.openai.com/api/docs/guides/decisions) supports three structured output types, each mapped to a concrete classification problem:

- **predicate** — returns a probability (0.0–1.0) that a statement is true. Use it to detect spam, flag PII, or check whether a message violates a policy.
- **choice** — picks one item from a predefined list and returns a confidence score. Use it to route support tickets, label intent, or assign categories.
- **score** — places an input on an ordered scale you define. Use it to rate ticket urgency, severity, or quality.

None of these return prose. You define the options; the model picks from them. That constraint is the point.

## What It Looks Like in Code

Here is a choice question routing a customer complaint:

``` python
import openai

client = openai.OpenAI()

response = client.decisions.create(
    model="gpt-6-luna",
    input="I was charged twice for my subscription last month",
    questions=[
        {
            "type": "choice",
            "question": "Which department should handle this complaint?",
            "choices": [
                {"id": "billing", "label": "Billing", "description": "Payments, invoices, and refunds"},
                {"id": "technical", "label": "Technical Support", "description": "Problems using the product"},
                {"id": "shipping", "label": "Shipping", "description": "Delivery and tracking"},
                {"id": "other", "label": "Other", "description": "Requests outside these categories"}
            ]
        }
    ]
)

department = response.decisions[0].choice.id        # "billing"
confidence = response.decisions[0].choice.probability  # 0.96
```

And a predicate checking for spam:

```
response = client.decisions.create(
    model="gpt-6-luna",
    input=comment_text,
    questions=[
        {
            "type": "predicate",
            "statement": "This comment is promotional spam or contains unsolicited links"
        }
    ]
)

is_spam = response.decisions[0].probability > 0.8
```

Multiple independent questions can share one request — you can check for spam, route to a department, and score urgency in a single call. Dependent decisions (where the answer to one question affects the next) require separate requests.

## When to Use This Instead of the Responses API

The Responses API is not going away. The Decisions API is a narrower tool for a specific job. Here is the actual decision tree:

- **Fixed set of answers (classify, route, score)** → Decisions API
- **Arbitrary JSON extraction (pull order details from a message)** → Responses API with Structured Outputs
- **Function call with parameters (look up an account)** → Responses API with function calling
- **Reasoning chain required** → Responses API or an o-series model

Using a Responses API call for pure routing wastes roughly five times the cost and adds 800–1,500ms of latency you do not need. The Decisions API runs the same [Luna model announced at DevDay 2026](https://openai.com/index/devday-2026-recap/) but skips the generation step entirely.

## The Price Case

The API charges $0.10 per million input tokens and nothing for output tokens. GPT-6 Luna through the Responses API costs approximately $0.50 per million input tokens. For teams running thousands of classifications per minute, that is a five-to-one cost difference on the same underlying model, plus the latency win.

It is worth comparing this to the other typed-decision story from this week: TypeSafe AI raised $870 million for Jev, a model purpose-built to return typed business decisions with a reasoning trace. They are not directly competing. Jev is for complex multi-step decisions with explainability requirements. The Decisions API is for fast, high-volume classification where you do not need a reasoning chain.

## What Is Missing in Beta

The API is genuinely useful today, but a few gaps matter before building production workflows around it:

- **One model only.** gpt-6-luna is the only option. More models are expected at general availability.
- **Images must be base64.** Hosted HTTP/HTTPS image URLs are not supported — you must encode images inline.
- **No documented rate limits.** OpenAI has not published RPM or TPM caps. Assume standard Luna limits apply until GA.
- **GA timeline.** Public beta as of October 6, 2026. General availability expected within weeks.

The [OpenAI community thread](https://community.openai.com/t/decisions-api-is-now-available-in-public-beta/1403877) has the most current developer feedback and workarounds. For zero-shot classification with a controlled option set — and particularly for agent routing pipelines — the Decisions API is already the right call. The main reason to wait for GA is rate limit clarity.
