The Agent Said It Was Done. The Database Disagreed. Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than on their generated text, running 507 stateful business workflows 20 times each against various LLM models. In a common-set ablation of 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks, and 67.24% of those failures still terminated cleanly, invoked a state-changing tool, and reported no final tool error, while executable checks found wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36%. The benchmark is available through Hugging Face and runnable via OpenEnv, with the paper at arxiv.org/abs/2608.19741. Viewer • Updated • 513 • 198 • 14 https://huggingface.co/datasets/microsoft/ThinkingBox-Bench The Agent Said It Was Done. The Database Disagreed. Enterprise Article https://huggingface.co/blog Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row. It is now available through Hugging Face. Figure 1: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind. From our ThinkingBox paper https://arxiv.org/abs/2608.19741 . This is a joint blog by Microsoft and Hugging Face, special thanks to Tommy Guy founder at Enderis AI, previously Microsoft , Sergio Paniego from Hugging Face and our former interns Zhuochun Li University of Pittsburgh , Ali Keramati UC Irvine , Youngmin Ko Northwestern for co-authoring/reviewing efforts. A customer writes in. Her $745 kitchen appliance has been stuck in a courier "exception" at a Nashville distribution center, fifteen days past its estimated delivery date. The AI agent does careful work. Nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly; her account segment genuinely does not qualify for late-delivery compensation. Then it closes the ticket as resolved and replies “ Since your query is resolved, is there anything I may assist you with? ” Two things are wrong. The carrier exception is still open, so the required end state was on hold , pending resolution. And the customer never got a real answer to what she actually asked. An AI grader checking tool calls would see nine well-formed ones. The grader checking whether the agent wrote to the database would see that too. The database is what disagrees . That gap is what ThinkingBox measures. Across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects. This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv https://github.com/huggingface/OpenEnv/tree/main/envs/thinkingbox env . You can run this one yourself: the example above is adapted from a benchmark task sandbox external retail group1.py:test case ST003 006 https://github.com/microsoft/thinkingbox-data/blob/thinkingbox-bench-v1.0/dataset/test case/sandbox external retail/sandbox external retail group1.py L984-L1223 , and the executable check that fails is a single field: the ticket's status is solved where the required end state is hold. The full trace is in Appendix D.4, Case 3 of our paper https://arxiv.org/pdf/2608.19741 . Contents Want to try it before reading the results? Skip to section Run it yourself run-it-yourself . A tool call is not an outcome Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. Only the records it leaves behind settle the question. The gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those state-check findings overlap. A trajectory is a claim. Database state is the evidence. Repetition is the trust test. One success is not reliability An agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent. So every task runs 20 independent times , each from an identical clean backend, and we report three different things: Table 1: The three numbers we report, and the question each one answers. | Metric | What it measures | What it answers | |---|---|---| | pass@1 | Share of all attempts that succeeded | How does it usually do? | | pass@20 | Share of tasks solved at least once in 20 tries | Can it ever do this? Breadth. | | Observed 20/20 | Tasks that actually passed all 20 recorded attempts | Can it always be correct? | We use observed 20/20 in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20. No estimator, no smoothing. Starting with the familiar view. The table below reports pass@1, the single-attempt score estimate, broken out by domain. This is the number most leaderboards publish, and on its own it reads like an ordinary capability ranking. Table 2: ThinkingBox-Bench pass@1 % by domain. Each model is evaluated on every task for 20 repeated trials. Bold marks the group leader; underline marks the runner-up. The standard errors for the single attempt score estimates are provided in Table 4 in our ThinkingBox paper https://arxiv.org/abs/2608.19741 . | Model | Retail 98 | Auto insurance 100 | Travel 104 | Neobank 104 | Consulting 101 | Overall, task-weighted 507 | |---|---|---|---|---|---|---| | Proprietary models | | | | | | | | Claude Opus 5.5 | 80.97 | 68.40 | 54.28 | 71.25 | 61.58 | 67.16 | | Claude Opus 5 | 80.71 | 65.80 | 49.95 | 70.62 | 66.19 | 66.50 | | GPT-5.4 | 76.33 | 62.65 | 68.12 | 65.34 | 54.60 | 65.36 | | GPT-5.6 Sol | 67.65 | 65.30 | 60.34 | 59.09 | 57.52 | 61.91 | | Claude Sonnet 4.6 | 72.35 | 54.40 | 58.94 | 56.39 | 54.31 | 59.19 | | GPT-6 Astra | 71.73 | 46.55 | 55.87 | 60.87 | 56.83 | 58.31 | | GPT-5.2 | 70.20 | 22.40 | 53.70 | 51.15 | 34.06 | 46.28 | | Claude Opus 4.6 | 68.62 | 8.30 | 21.11 | 35.67 | 27.82 | 32.09 | | o3-pro | 37.70 | 2.95 | 17.31 | 24.28 | 14.60 | 19.31 | | Grok-4.3 | 43.93 | 2.60 | 15.14 | 1.78 | 9.55 | 14.38 | | Open-weight models | | | | | | | | Kimi-K3 | 82.24 | 50.80 | 61.83 | 41.35 | 51.63 | 57.37 | | Qwen3.8-27B | 64.03 | 47.85 | 53.41 | 47.88 | 45.69 | 51.70 | | DeepSeek-V4-Pro | 68.21 | 29.65 | 43.13 | 44.86 | 31.04 | 43.26 | | Kimi-K2.6 | 53.72 | 24.50 | 39.52 | 33.65 | 37.33 | 37.66 | | GLM-5.1 | 58.67 | 25.70 | 35.43 | 13.27 | 34.06 | 33.19 | | Qwen3.6-27B | 43.11 | 29.00 | 46.39 | 27.84 | 18.37 | 32.94 | | Qwen3.5-9B | 19.90 | 0.70 | 4.71 | 1.15 | 2.33 | 5.65 | | Mistral-Large-3 | 11.28 | 1.30 | 8.99 | 1.15 | 0.74 | 4.66 | Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5. Kimi-K3 is the strongest open-weights model , within a point of GPT-6-Astra. Domain matters just as much : Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance. One good run tells you a model can do the work. It does not tell you whether it will do it again. So run every task 20 times and ask how much of that score survives. Figure 2: How much of each model's single-attempt score survives 20 repeats. Only three hold on to most of their pass@1 scores: GPT-6 Astra retains 78% of its single-attempt rate, and Claude Opus 5.5 and Claude Opus 5 each retain 71%. At the other end, GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each keep about 8%. The gap between what a model can do once and what it does every time is the whole story. Can you depend on the model behind your agent? Figure 3: Breadth and consistency pull apart. Twelve of the eighteen models are shown; six below 33% pass@1 are omitted for legibility. Kimi-K3 has the broadest coverage of any model we tested. It solves 93.89% of the benchmark at least once: 476 of 507 tasks. Only 31 tasks defeat it entirely, the lowest count in the field. On retail workflows it leads outright at 82.24% pass@1, ahead of every proprietary model. Kimi-K3 is also among the least consistent. Just 68 of 507 tasks, 13.41%, succeed in all 20 attempts. Claude Opus 5 inverts this. It solves fewer tasks at least once 79.09%; 106 defeat it entirely but completes 47.53% of the benchmark on every single attempt. A newer model does not fix this. Claude Opus 5.5 scores higher than Claude Opus 5 on every-attempt average, 67.16% against 66.50%, and solves more tasks at least once. It passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought no additional dependability at all. - Kimi-K3 solves 75 more tasks at least once than Opus 5. - Opus 5 solves 173 more tasks consistently than Kimi-K3. If you are choosing a model for work that touches real records, pass@20 is the wrong column to look at. What consistency costs Capability comparisons usually stop at the score. For anyone deploying, the relevant question is what a successful unit of work costs. We measure that as cost per successful task attempt. We say task attempt because every benchmark task is run repeatedly and cost is incurred per attempt, so pass@1 is the matching quality denominator. We took each model's recorded token usage from its full 507 × 20 campaign and priced it at undiscounted list rates available on OpenRouter https://openrouter.ai/