A practical way to use bounded AI decisions for routing and triage — while keeping authority, policy, and recovery in your application.
Fast choices are useful. Fast choices with unowned authority are not.
A model that answers in a few milliseconds can still make a costly mistake. It can route an urgent support ticket to the wrong queue, tell an agent to retry a charge, or send a sensitive request down a path that should have required a person.
That is the tension behind the OpenAI Decisions API. OpenAI’s changelog lists the Decisions API in beta with gpt-6-luna, and the company’s DevDay coverage describes a product for narrow, repetitive choices. That is a useful primitive. It is not a permission slip to turn model output into a production command.
The smart rollout is to treat a decision response as a recommendation inside a contract. Your application owns the options, proof, policy, effect, audit record, and recovery path. The model helps select a lane. It does not become your policy engine.
Most generative AI interfaces are built to write. You ask for an explanation, an answer, or a draft, then a human reads it. A decision interface has a smaller job: choose among a finite set of allowed answers. That can be powerful for support routing, document triage, agent next-step selection, quality checks, or deciding whether a workflow should .
The important word is allowed. A good decision request does not ask, “What should we do?” It asks, “Given this evidence, which of these specific safe paths fits best?”
That boundary improves speed, makes the result easier to test, and removes several kinds of unexpected behavior. But it does not make the decision correct. Structured data proves shape. It does not prove that the evidence was sufficient, the labels were fair, or the selected option was appropriate.
Use a decision model for: classification, prioritization, routing, bounded next steps, and triage. Do not use it as the sole authority for: money movement, access grants, destructive changes, medical or legal determinations, or anything you cannot explain and reverse.
Your first use case should be boring in the best possible way. Pick a frequent decision with a known trusted outcome and a harmless fallback. A support team might already tag tickets as billing, technical, account, or needs-review. An internal developer tool might decide whether a request is ready for a test run, needs a missing artifact, or should be handed to an engineer.
Avoid starting with “choose the right action for the agent.” That question is too broad. Start with “select one of three review lanes for this change request.” Small scope turns disagreement into a measurable signal rather than a surprise in production.
Use four filters before you build:
The decision contract is the most valuable artifact in this project. It is a small, versioned definition that says what evidence may be used, what options mean, what uncertainty looks like, and what each option is allowed to trigger. Put it in your codebase. Review it like an API change.
Here is a provider-neutral TypeScript shape. It intentionally does not assume an unpublished request field or SDK method.
type Route = "billing" | "technical" | "account" | "human_review";
type DecisionContract = { id: "support-route-v1"; evidence: { message: string; customerPlan: "free" | "paid" | "unknown"; priorOpenCases: number; }; options: readonly Route[]; abstain: "human_review"; policy: { automaticEffects: readonly ["billing", "technical", "account"]; forbiddenWhen: readonly string[]; };};
type Decision = { contractId: string; chosen: Route; providerSignal?: number; evidenceHash: string; receivedAt: string;};
Notice what is missing: a model name hard-coded into business logic, a raw free-form prompt as the source of truth, and permission to perform an effect. Those belong in separate layers.
Keep option definitions short and mutually useful. If two people cannot consistently separate billing from account, the model cannot rescue the taxonomy. Add a clear human_review option rather than forcing a false choice.
A production decision path should preserve an escalation lane and a reversible action boundary.
This is the rule that keeps a clever demo from becoming an incident. The model may recommend a route. Only deterministic application policy may authorize the next effect.
For example, a decision can recommend technical. Your policy layer then checks that the ticket has required fields, the destination integration is healthy, the user is eligible for automated routing, and the action has not already happened. If any check fails, the system sends the item to review.
function authorize(decision: Decision, ticket: Ticket): Effect { if (decision.chosen === "human_review") return { kind: "queue_review" }; if (!ticket.hasRequiredFields) return { kind: "queue_review" }; if (ticket.isHighRisk || ticket.isDuplicate) return { kind: "queue_review" }; return { kind: "route", destination: decision.chosen };}
The benefit is architectural as well as safe. You can swap model providers, change a prompt, or adjust a beta integration without changing the rights your system grants. You can also prove after the fact why an effect occurred: the decision was one input; deterministic checks were the authority.
Decision quality usually falls when you paste an entire conversation, agent trace, or knowledge-base dump into a fast choice call. More text can hide the signal, leak unrelated data, and make each result harder to reproduce.
Create a minimal evidence packet instead. Include only fields a reviewer would legitimately use for the same decision. Normalize dates, redact secrets, mark unknown values explicitly, and preserve source identifiers where a reviewer needs to inspect the record later.
For a support route, the packet might include the customer’s message, current product area, plan, and whether a payment event exists. It should not include a full internal profile, unrelated conversations, or hidden instructions copied from a web page. If the decision needs those details, that is a sign the decision is too broad or should remain human-led.
Hash the serialized packet and store the contract version beside the outcome. That turns a mysterious “the model did it” report into something you can replay against the same inputs.
Do not make the model live on the day it starts returning valid responses. Run it beside the existing workflow first. In shadow mode, the model receives real or replayed traffic and records a recommendation, but it cannot change anything. Compare it with the route selected by your current rules or a trained reviewer.
Build a small labeled evaluation set before that rollout. It should contain ordinary examples, ambiguous examples, incomplete evidence, adversarial wording, language variation, and cases where every automatic option would be wrong. Version the set with the contract.
Measure more than raw agreement:
If the provider returns confidence or probabilities, treat them as a hypothesis. A value of 0.95 means nothing until you compare that band with outcomes on your data. Confidence is evidence for a threshold; it is not authorization by itself.
Preview features are exciting precisely because their contract can change. Keep the OpenAI Decisions API behind a narrow adapter so a changed beta field, availability restriction, or pricing rule does not leak across your product.
interface DecisionProvider { decide(input: DecisionInput): Promise<Decision>;}
class DecisionService { constructor(private provider: DecisionProvider) {}
js
async recommend(input: DecisionInput) { const result = await this.provider.decide(input); return validateAgainstContract(input.contract, result); }}
The adapter translates your stable contract into a provider request and translates the response back. In tests, replace it with a fixture provider. In an outage, route to a deterministic rule, a human queue, or a compatible implementation. Do not let a beta response shape become your domain model.
This also prevents a common search-result trap. Multiple independent services now use “Decisions API” in their names. Their examples may be useful for learning the category, but they are not proof of the OpenAI beta’s public API contract. Check the official changelog and released reference documentation before copying an endpoint, parameter, price, or latency claim into production.
Not every branching problem deserves a model call. If a customer is on a blocked list, a field is missing, a ticket is already closed, or a request violates an explicit ownership rule, use ordinary code. Deterministic checks are cheaper, easier to audit, and exact by design.
A bounded model decision earns its place where language, context, or competing signals require judgment. “Does this customer describe a billing issue or a configuration failure?” can be a useful model question. “Is the message longer than 5,000 characters?” is not. Blending the two carelessly creates work you cannot explain: a model appears to be deciding a policy that should have lived in a testable rule.
Make the split explicit in your contract review. List the hard rules that must run before the model, the semantic judgment the model may make, and the hard rules that run after it. This three-part design gives engineers a simple debugging sequence: check the input policy, inspect the bounded recommendation, then inspect the authorization path. It also means a model regression cannot silently change your non-negotiable business rules.
Audit the recommendation, the policy checks, and the final effect as separate events.
Once shadow comparisons look good, promote one narrow, reversible action. That might be adding a queue label, suggesting a draft assignment, or opening a review task. It is not granting access, publishing content, deleting data, or changing a billing record.
Every automated effect should have an idempotency key, a trace ID, a clear owner, and a reversal procedure. Log the contract ID, input hash, chosen option, policy checks, resulting effect, reviewer override, and final business outcome. Those records are how you find drift after a product change, a model update, or a new customer segment.
Set triggers in advance. Examples include a jump in human overrides, an unexpected increase in abstentions, a destination outage, a contract version mismatch, or a critical false route. The point is not to make automation look perfect. It is to make a safe faster than a prolonged debate.
This may sound slower than wiring a new endpoint to a switch statement. In practice, it is faster than untangling a product rule, model prompt, provider quirk, and customer-impacting effect after they have been fused together.
The OpenAI Decisions API is interesting because many AI products do not need another paragraph. They need a fast, bounded choice that lets a broader system move. The teams that benefit most will not be the ones that hand it the most authority. They will be the ones that give it the clearest job.
Build a contract. Keep an abstain option. Test against trusted outcomes. Separate recommendation from authorization. Make the first effect reversible. Do that, and a beta decision API can become a useful component in a dependable workflow instead of a new place for invisible policy to hide.
Use it for narrow, finite choices such as routing, classification, triage, and a bounded next step. The application should still own permissions, business policy, and the final effect.
No. Structured output confirms that a response has the expected shape. It does not confirm that the evidence, labels, decision, or confidence level is correct. Evaluate outcomes on a representative labeled set.
Only after calibration on your own cases. Compare confidence bands with observed error, segment results by option and input type, and set thresholds based on the harm of a mistake. Keep an escalation lane.
Put the provider behind a small adapter that accepts your stable decision contract and returns your stable domain result. Keep beta request fields, model identifiers, and pricing assumptions inside that adapter.
Keep a human in the loop for irreversible or high-impact outcomes, including money movement, privileged access, deletion, legal or medical decisions, disciplinary actions, and decisions where you cannot clearly explain or reverse the outcome.
Track false automation rate: the share of decisions that would have taken an automatic route but should have gone to review. Pair it with error by option, override rate, and downstream correction rate.
OpenAI Decisions API Rollout: Build Fast AI Routing Without Letting It Make the Final Call was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.