Microsoft releases an agent benchmark where the database gets the final vote Microsoft and Hugging Face released ThinkingBox on October 3rd, a benchmark that grades AI agents on the final state they leave in business systems rather than on their transcripts or tool calls. The benchmark, written by Microsoft applied scientist Tuhin Kundu with Tommy Guy credited as co-author and reviewer, contains 507 simulated workflows spanning retail, travel and hospitality, auto insurance, neobank support, and consulting IT and HR, each run 20 times; 477 tasks are graded on state alone and 30 also use response rubrics. The authors report pass@1, pass@20, and observed 20/20 to distinguish a lucky single success from repeatable work. Microsoft releases an agent benchmark where the database gets the final vote ThinkingBox checks 507 simulated business workflows against their final records, with 20 runs per task to show the gap between a lucky success and repeatable work. By Ryan Merket https://runtimewire.com/author/ryan-merket ยท Published Primary source: Hugging Face Newsroom https://huggingface.co/blog/microsoft/thinkingbox Why it matters Agent demos often reward convincing transcripts and valid tool calls. ThinkingBox makes teams confront the deployment question: does the agent reliably leave the right records behind? Microsoft and Hugging Face released ThinkingBox on October 3rd, a benchmark that grades AI agents on what they actually change in business systems. The joint Hugging Face post https://huggingface.co/blog/microsoft/thinkingbox?ref=runtimewire , written by Microsoft applied scientist Tuhin Kundu https://huggingface.co/tuhink?ref=runtimewire , introduces 507 simulated workflows and a test that runs each one 20 times. The results show how often a plausible-looking agent leaves the wrong record behind. The post also credits Tommy Guy https://www.alumni.appstate.edu/s/1727/c20/Interior.aspx?cc=1&gid=2&pgid=5742&scontid=0&sessionid=40209ce0-6bf2-487d-9a29-9d62aa7c0339&sid=1727&sparam=Search&ref=runtimewire as a co-author and reviewer; it identifies him as founder of Enderis AI and a former Microsoft employee. Appalachian State's alumni profile says Guy double-majored in philosophy and religion and mathematics, graduated summa cum laude, and later earned master's degrees in mathematics and computer science at Wake Forest. In describing his Microsoft work, the university says he serves as a point of contact between AI developers and product managers. ThinkingBox tests a related question: what does "the agent succeeded" mean when the database says otherwise? The record is the test The example in the post is a customer waiting 15 days for a $745 kitchen appliance delayed at a Nashville distribution center. An agent makes nine tool calls, checks the refund policy, opens a support ticket and documents the case. It then marks the ticket resolved even though the carrier exception is still open and the required status is "on hold." Its final message says the query is resolved without answering the customer's concern. The example is not a live customer incident. The authors say the public benchmark's tasks are synthetic reconstructions, modeled on enterprise agent workflows. The test definition https://github.com/microsoft/thinkingbox-data/blob/thinkingbox-bench-v1.0/dataset/test case/sandbox external retail/sandbox external retail group1.py?ref=runtimewire L984-L1223 specifies the failure: one field, the ticket status, has the wrong value. The agent's fluent explanation and successful tool calls cannot repair that state. ThinkingBox separates the sandbox from the benchmark. The sandbox creates isolated sessions for tool-using agents; ThinkingBox-Bench https://github.com/microsoft/thinkingbox-data?ref=runtimewire supplies 507 tasks spanning retail, travel and hospitality, auto insurance, neobank support, and consulting IT and HR. Each task starts from a clean backend. Executable checks compare the final state and side effects with the expected outcome. In 477 tasks, grading relies on state alone; 30 also use response rubrics for requirements that cannot be reduced to a database field, according to the v1.0 release documentation https://github.com/microsoft/thinkingbox-data/blob/main/releases/thinkingbox bench v1/README.md?ref=runtimewire . One pass is not a service level The authors report https://huggingface.co/blog/microsoft/thinkingbox?ref=runtimewire three measures: pass@1, the success rate on individual attempts; pass@20, whether a task succeeds at least once in 20 tries; and observed 20/20, the literal number of tasks that passed all 20 recorded runs. Observed 20/20 is the closest match for work that has to be right every time. In the reported results https://huggingface.co/blog/microsoft/thinkingbox?ref=runtimewire , Claude Opus 5.5 https://runtimewire.com/models/anthropic/claude-opus-5.5:batch scored 67.16% pass@1, while open-weight Kimi-K3 https://runtimewire.com/models/moonshotai/kimi-k3 scored 57.37%. But Kimi-K3 completed 476 of 507 tasks at least once and passed all 20 attempts on only 68 tasks, or 13.41%. Claude Opus 5.5 passed every run on 241 tasks, or 47.53%. Claude Opus 5 https://runtimewire.com/models/azure/claude-opus-5 passed the same number consistently despite a lower single-attempt score. A model can look broadly capable while remaining unreliable on a particular workflow. Benchmarks based on conversation quality or valid tool syntax can count a carefully executed failure as progress. For agents touching customer accounts, bookings, claims or employee records, the risk lies in the state change and any collateral effects, not in whether the run ended without an error message. The authors' common-set ablation https://huggingface.co/blog/microsoft/thinkingbox?ref=runtimewire covered 121,680 valid trials across 12 models. Of the 79,853 attempts that failed the executable checks, 67.24% still ended cleanly, had invoked a state-changing tool and reported no final tool error. Among failures, graders found incorrect field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%; categories overlap. Those numbers describe the authors' test environment and models, not the failure rate of agents in production. Open code, bounded claims The release is available through Hugging Face's OpenEnv interface https://github.com/huggingface/OpenEnv/tree/main/envs/thinkingbox env?ref=runtimewire , with the framework and benchmark data published separately on GitHub. The documented ThinkingBox-Bench v1.0 setup requires Python 3.12, uv, Typesense 30.1, MCP tool servers and proxy services, plus configured agent, user-simulator and judge model endpoints, according to the release documentation https://github.com/microsoft/thinkingbox-data/blob/main/releases/thinkingbox bench v1/README.md?ref=runtimewire . For reproducibility, the project specifies the thinkingbox-data release tag and recommends recording the exact ThinkingBox and thinkingbox-data commits, configuration, model deployments and inference parameters. The documented procedure does not specify a bundle hash. The research paper, submitted to arXiv on August 20th and revised through October 1st https://arxiv.org/abs/2608.19741?ref=runtimewire , predates the October 3rd availability announcement. The release makes an existing research project easier to inspect and run; it does not show that the benchmark predicts performance against live company databases. Its tasks use synthetic records, policies and isolated tool sessions. The authors also say they have not measured whether proposed production safeguards, such as targeted retries or human approval for hard-to-reverse changes, improve scores. A benchmark can reveal where a model or agent stack breaks under defined conditions, but a synthetic test cannot certify that a company's own integrations, policies and messy records will behave the same way. ThinkingBox gives teams a way to test that gap instead of treating a successful transcript as proof of completed work. Teams can run repeated, state-based checks against the workflows they plan to automate.