Microsoft VP Achint Srivastava, a Pi Labs co-founder, says Microsoft-Decision-1 scores fixed choices; Foundry lists input tokens at $0.042 per million.
By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
· Published
Primary source: [Microsoft Command Line](https://commandline.microsoft.com/microsoft-decision-1-model-foundry/)
Why it matters #
Microsoft is packaging AI evaluation and routing as a low-cost model call inside Foundry. If developers trust its fixed-choice scores in production, the check between an agent's reasoning and its next action becomes a service Microsoft can sell across workflows.
Microsoft has put Microsoft-Decision-1, a model that scores fixed choices instead of generating text, in Microsoft Foundry. The launch brings Achint Srivastava, a Microsoft vice president who previously co-founded AI evaluation startup Pi Labs, into a growing effort to give software and AI agents a fast, bounded way to make operational calls.
In his October 9th announcement, Srivastava describes Decision-1 as a tool for routing, classification, prioritization, verification and workflow control. Given a situation and a defined set of choices, it returns a probability for each option. Developers can use those scores to route a support ticket, check an agent's proposed action, or send an uncertain case for human review.
The product fits the work Srivastava did before Microsoft. He and David Karam founded Pi Labs, which built tools to evaluate and improve AI applications; Accel's profile lists Microsoft as Pi Labs' acquirer. Decision-1 applies a related concern - assessing AI output against explicit criteria - inside Microsoft's own model platform. The available sources establish Srivastava's roles and Pi Labs' focus, but do not establish that Pi Labs' technology or team built Decision-1.
A model for the step between prompt and action
Decision-1 is designed for cases where an application already knows the possible answers and needs a model to select or score them. Microsoft's Foundry listing says the model accepts text inputs of up to 32,768 tokens and returns JSON. It supports yes-or-no questions, multiple choice, ratings, rubric grading, relevance judgments and an explicit "cannot tell" option. It does not generate explanations, handle open-ended questions or accept images, audio or video.
That narrower job matters in automated workflows. An agent can ask whether a proposed tool call meets a rubric, then proceed, retry or escalate based on the score. Microsoft's engineers describe the cost of serial checks in the announcement: adding 100 milliseconds to each of 20 dependent decisions adds two seconds to a workflow. A small, dedicated model could cut that delay and the expense of sending each check through a general-purpose language model.
Microsoft lists input tokens at $0.042 per million, with output tokens free. The price puts a number on Microsoft's claim that decision checks can run cheaply. The greater operational question is whether the score is dependable on the customer-specific decisions developers actually need to automate, and whether it is calibrated well enough to set escalation thresholds safely.
The benchmark claim is Microsoft's
Microsoft says Decision-1 had the highest accuracy in its comparison of 36 benchmarks and nearly 150,000 questions, with benchmarks kept blind from training. Microsoft also reports median latency 35 times faster than GPT-6 Sol and 4.5 times faster than Quyet-1.0-Large. These are Microsoft's results; its announcement does not provide the full datasets or latency test conditions needed to reproduce the comparison independently.
Microsoft also describes internal tests: Xbox Research used the model to sort more than 10,000 feedback items, while Copilot's team used it to assess AI responses. Those examples show the intended role inside Microsoft, but the results are also Microsoft-reported. The Foundry catalog warns that Decision-1 should not be the sole automated decision-maker for consequential matters involving employment, credit, housing, healthcare or legal rights. Integrating applications still need to set thresholds, human review and safeguards.
The release comes as decision models move from a specialist idea toward a platform category. Cloudflare introduced its Clef models on October 1st, with open weights and deployment through Workers AI. OpenAI's Decisions API is in public beta, returning typed answers for fixed questions. Microsoft's offering starts with Foundry availability and a plan to bring the model to OpenRouter; its announcement says the OpenRouter access is coming soon, without giving a date.
RuntimeWire recently covered Microsoft's ThinkingBox agent benchmark, which evaluates simulated business workflows against their final records. Decision-1 targets a different layer of the same engineering problem: how software checks intermediate choices while a workflow is running. As agents take more steps without waiting for a person at each turn, those small decisions become part of the product's control system.
Microsoft says Decision-1 is based on Alibaba's open-weight Qwen3.5-9B, post-trained with public datasets and synthetic data. Microsoft says it plans to rebase the model on other systems, including Microsoft MAI and OpenAI models. That makes the launch a test of a product format as much as a single model: whether customers will treat bounded scoring as a separate, reusable service inside their applications.
Srivastava's path from building AI evaluation tools at Pi Labs to announcing a decision model at Microsoft gives the launch a founder's throughline: software needs a practical way to judge model outputs, not just generate them. The commercial case will rest on whether the scores remain useful beyond Microsoft's own test suites and internal workflows.