FoxTPNL / Wikimedia Commons (CC BY 4.0) The new API uses a specialized GPT-6 Luna model to pick one answer from a developer-defined list, with early demos showing roughly 10x lower latency
OpenAI wants some of its models to talk less and decide more.
At DevDay 2026 on September 29, the company introduced the Decisions API, a tool built for quick, inexpensive choices rather than long-form answers. Picture a multiple-choice exam for AI. The developer writes the options, and the model circles one.
How the Decisions API works #
The engine is a specialized version of GPT-6 Luna. Instead of generating open-ended responses, it returns a single answer from a finite set of options that the user defines ahead of time.
The API accepts both text and image context. That means a developer can feed it a screenshot, a photo or a block of text and ask it to sort the input into a predefined bucket.
OpenAI pointed to several practical jobs for the tool. These include content classification, routing incoming requests such as support tickets, and choosing the next action an AI agent should take.
Speed is the standout number. Early demonstrations showed end-to-end response times of around 150 milliseconds, compared with ~1.6 seconds for standard Luna calls.
That works out to a latency improvement of approximately 10 times. For an agent chaining many small decisions together, those seconds stack up like cars at a broken traffic light.
Access is narrow for now. The Decisions API launched in limited preview for select customers, with broader availability expected in the coming days.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Several key details have not been published yet. OpenAI has not released pricing, request and response formats, confidence scores or accuracy benchmarks for the new API.
A growing market for small, sharp models #
TypeSafe AI’s Jev model is one example of a similar offering already in the space. OpenAI’s launch reads as both a competitive response to products like Jev and a signal that the segment is worth taking seriously.
The timing also fits the broader theme of this year’s DevDay. The event put heavy emphasis on improving AI agent infrastructure and capabilities, and a fast decision layer slots neatly into that agenda.
What this means for developers, rivals and the AI market #
For developers building agents, the Decisions API hints at a cleaner architecture. A heavyweight model can handle the reasoning, while a lightweight decision model directs traffic between steps.
If the latency figures hold up outside demos, a pipeline that waits ~1.6 seconds at every branch point behaves very differently from one that waits around 150 milliseconds.
Cost is the other half of the equation, and it remains unknown. OpenAI pitched the API as inexpensive, but without published pricing, teams cannot yet model whether switching their routing logic makes financial sense.
The missing confidence scores matter more than they might seem. In decision systems, knowing how sure a model is often determines whether an answer gets acted on automatically or escalated to a human or a larger model.
A constrained output is not automatically a correct output. A model forced to choose from three options will always choose one, even when none of them fits well, so accuracy benchmarks will be the real test.
Investors and builders watching the AI infrastructure race will likely track a few signals closely. Those include the pricing announcement, the first accuracy numbers, and whether the broader release lands in the coming days as expected.
Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our