Thinkingbox Benchmark Grades AI Agents on Database State
Microsoft and Hugging Face released ThinkingBox, an agent sandbox, and ThinkingBox-Bench, a benchmark that grades AI agents on terminal backend state and side effects rather than generated sentences. Across 507 stateful business workflows run 20 times each, the post reports that most failed attempts still terminated cleanly and reported no tool error.
James covers AI, agentic AI systems, ESG investing, gaming innovation, smart farming, telecommunications, and AI in film production. Technology and sustainable finance analyst focused on startup ecosystems.
Executive Summary
- Microsoft and Hugging Face released ThinkingBox, an agent sandbox, and ThinkingBox-Bench, a dataset benchmark that grades AI agents on the terminal state and side effects they leave in a backend database rather than on the sentences they generate, now accessible through Hugging Face, according to the joint blog post.
- Across a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks; 67.24% of those failures still terminated cleanly, invoked a state-changing tool and reported no final tool error, per the published results.
- Every task runs 20 independent times from an identical clean backend; Claude Opus 5.5 leads the task-weighted pass@1 table at 67.16%, while Kimi-K3 is the strongest open-weights model at 57.37%, according to Table 2 in the post.
- Failure diagnostics are dominated by tool handling at 79.9% of failures, ahead of wrong state updates at 10.3%, incomplete user resolutions at 7.0% and no state-changing action at 2.9%, per the ablation findings.
Key Takeaways
- Final responses and valid tool calls are proxies; the blog states that only the records an agent leaves behind settle whether the work was done.
- One clean run is not reliability: ThinkingBox reports pass@1, pass@20 and observed 20/20 rates separately, and the gap between them is the central finding.
- Cost per successful task attempt and cost per dependable task diverge; the cheapest per single success is not the cheapest per dependable outcome.
- The environment, dataset and harness are released through Hugging Face and the OpenEnv interface, so the repeat-rate question is testable by outside teams.
ThinkingBox Grades Database State, Not Prose
The joint Microsoft and Hugging Face post opens with a retail scenario that illustrates the benchmark's premise. A customer's $745 kitchen appliance is stuck in a courier exception at a Nashville distribution center, fifteen days past its estimated delivery date. The agent makes nine tool calls: it pulls the order, checks tracking, looks up the customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline and reads the policy correctly, because the account segment does not qualify for late-delivery compensation. It then closes the ticket as resolved and asks whether anything else is needed.
Two things are wrong. The carrier exception remains open, so the required end state was hold, pending resolution. And the customer never received an answer to what she asked. According to the post, an AI grader checking tool calls would see nine well-formed ones, and a grader checking whether the agent wrote to the database would see that too. The database disagrees. The executable check that fails is a single field: the ticket's status is solved where the required end state is hold. The example is adapted from a benchmark task identified as sandbox_external_retail_group1.py:test_case_ST003_006, with the full trace in the paper's Appendix D.4, Case 3.
The gap between a clean-looking trajectory and the actual terminal state is what ThinkingBox measures. Each task defines a starting backend state, a user goal, available MCP tools, domain policy and executable checks over the terminal state. A simulated user holds private context and releases it only when asked. Every attempt gets an isolated MCP session with freshly initialized state, so two attempts of the same task never share a database row or cached tool state. A side-effect extractor derives what changed, and deterministic judges compare it against the required end state, accepting any trajectory that produces the right outcome while rejecting wrong, missing or extra effects. Of the 507 tasks, 477 are graded on state alone and 30 add response rubrics.
Why ThinkingBox Separates Capability From Consistency
The benchmark reports three numbers rather than one. Pass@1 is the share of all attempts that succeeded. Pass@20 is the share of tasks solved at least once in 20 tries. Observed 20/20 is the literal count of the 507 tasks that passed all 20 recorded attempts, with no estimator or smoothing. The post states that the last is the trust test.
The common-set ablation found that of the failed attempts that terminated cleanly with a state-changing tool call and no final tool error, executable checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%, with overlap among those findings. In the domain table, Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5. Domain matters as much as model choice: Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.
Repetition pulls breadth and consistency apart. Kimi-K3 solves 93.89% of the benchmark at least once, 476 of 507 tasks, the lowest defeat count in the field, and leads retail outright at 82.24% pass@1. It is also among the least consistent, with just 68 of 507 tasks, 13.41%, succeeding in all 20 attempts. Claude Opus 5 inverts that profile: it solves fewer tasks at least once, 79.09%, but completes 47.53% of the benchmark on every attempt. The post notes that Claude Opus 5.5 scores higher than Claude Opus 5 on every-attempt average, 67.16% against 66.50%, and solves more tasks at least once, yet passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought no additional dependability.
Related: Emerging Agentic AI Technologies That Will Dominate 2026
What Consistency Costs in the ThinkingBox Data
The post argues that capability comparisons usually stop at the score, while deployers need to know what a successful unit of work costs. It defines cost per successful task attempt as the estimated cost for 507 attempts, one per task, divided by 507 times pass@1, priced at undiscounted list rates available on OpenRouter+, reversing promotional discounts and excluding endpoints that declare quantization. Input, output and cache rates come from one provider endpoint per model. The post states this is a comparative efficiency index, not an invoice, and not the price of serving one production request. It also prices single successes, not consistency.
Three models sit on the Pareto cost frontier: GPT-5.6 Sol has the lowest cost per success at $0.127; GPT-5.4 raises pass@1 by 3.45 percentage points for $0.004 more per success; Claude Opus 5.5 adds another 1.80 points at $0.276 per success. Claude Opus 5 at $0.475 per successful attempt and 66.50% pass@1 is both more expensive and less accurate than Claude Opus 5.5.
Cost per dependable task divides the estimated cost of a full 20-run campaign by the number of tasks passed on all 20 attempts. GPT-5.4 is cheapest at $6.80, though only 128 tasks meet the bar. GPT-6 Astra reaches 231 at $7.45, and Claude Opus 5.5 the joint-highest 241 at $7.80. GPT-5.6 Sol, cheapest per single success, costs $9.76 per dependable task. The post's conclusion is blunt: the cheapest way to get a right answer is not the cheapest way to get a dependable one.
For deeper context, see our EdTech analysis: "Coursera Report Maps AI Skills Alongside Human Capabilities".
Failure Signatures and the OpenEnv Release Path
The post assigns each failed trace one deterministic diagnostic signature, with an actionable headline: roughly four in five failures are tool handling, not reasoning. Tool usage accounts for 79.9% of failures, wrong state updates 10.3%, incomplete user resolutions 7.0% and no state-changing action 2.9%. The post describes these as unweighted averages of per-model shares and observable labels, not unique causal explanations. The practical pattern, it says, is that agents usually get far enough to attempt the workflow, then fail to recover from tool errors, failed preconditions or empty lookups. That is framed as a retry and error-recovery problem before it is a model problem. Difficulty also varies by domain: across the models listed, retail averages 59.52% pass@1 while auto insurance averages 33.83%.
ThinkingBox is now on Hugging Face, both the harness and the dataset, with ThinkingBox-Bench behind the OpenEnv interface. Each finished episode returns a binary pass/fail reward. The released adapter is designed for evaluation, and the post notes that separate, non-benchmark scenarios can use the same interface in training workflows. The documented setup runs on Linux and WSL with Python 3.11 or later, uv and Docker, and requires a thinkingbox-data checkout at the pinned release plus model endpoints for the agent, simulated user and judge. One endpoint can serve all three roles. The OpenEnv image starts only the OpenEnv API; Typesense, the MCP session proxy and the OpenEnv server are run separately. Readiness is gated on an endpoint that returns 503 until observable data, configuration and Session Proxy checks pass, and the post cautions it cannot observe Typesense or live-probe every model endpoint. The OpenEnv adapter writes operational failures to an errors sidecar so they can be rerun rather than silently mixed with model outcomes. Runs are gated on a pinned framework commit, a pinned data release and a bundle hash. The post also states that every task in the public benchmark is a synthetic reconstruction, that workflows and policies are modeled on real agentic enterprise patterns, and that the customers are not real.
ThinkingBox Signals and Coverage
| Entity | Recent Focus | Geography | Source |
|---|---|---|---|
| Microsoft | Built ThinkingBox and ThinkingBox-Bench with the Copilot Studio team in partnership with Toloka | Not specified in the source | Hugging Face blog |
| Hugging Face | Hosts ThinkingBox and ThinkingBox-Bench and provides the OpenEnv interface | Not specified in the source | Hugging Face blog |
| Toloka | Partnership on the benchmark build | Not specified in the source | Hugging Face blog |
| University of Pittsburgh, Northwestern University, Columbia University, UC Irvine | Collaborators who interned at Microsoft | Not specified in the source | Hugging Face blog |
| Kimi-K3 | Strongest open-weights model tested and broadest task coverage at 93.89% | Not specified in the source | Hugging Face blog |
| Claude Opus 5.5 | Highest task-weighted pass@1 at 67.16% | Not specified in the source | Hugging Face blog |
| OpenRouter+ | Pricing source for undiscounted list rates used in cost analysis | Not specified in the source | Hugging Face blog |
Geography is not disclosed for any entity in the supplied source, so no location detail is claimed here.
Additional coverage: Pinterest Debuts Ask Pinterest AI Shopping App in 2026
Hugging Face Implementation Risks
The post flags limitations that bear on how the results should be read. Its cost figures are modeled at undiscounted list rates rather than actual cloud bills, and they price single successes as well as consistency, so neither figure should be treated as an invoice or as the cost of serving one production request. The failure signatures are unweighted averages of per-model shares and observable labels, which the post explicitly says are not unique causal explanations. The benchmark is synthetic: every task is a reconstruction modeled on real agentic enterprise patterns, and the customers are not real. The post also states that the lift from its own recommended remedies, such as checking terminal state before committing, classifying recoverable tool errors, cutting the tool surface and requiring human approval on changes that cannot be cheaply reversed, has not been measured on this benchmark. No third-party verification of these results is described in the source, and no independent replication is claimed. Deployment decisions should therefore treat ThinkingBox as an evaluation environment to reproduce rather than as a settled ranking, and should follow the post's advice to define and report a repeat metric, stating whether it is best-of-k or every-of-k and how it was computed.
Editorial independence disclosure: this article is an independent newsroom summary of the supplied source and was not reviewed or approved by Microsoft, Hugging Face, Toloka or any party named above. See the source note at https://huggingface.co/blog/microsoft/thinkingbox.
What This Means for Practitioners
For teams wiring agents into systems that touch real records, the practical shift is where the pass/fail decision lives. If a trajectory can terminate cleanly with a state-changing tool call and no final tool error and still be wrong, then final responses and tool-call logs are insufficient acceptance criteria. The concrete moves the post supports are checking the terminal state before committing rather than trusting the model's summary of it, classifying tool and system errors so retries target recoverable ones, narrowing the tool surface to what the workflow needs, and requiring human approval on changes that cannot be cheaply reversed. Treat a repeat rate, not a single success, as the design input.
About the Author
James Park AI Author
AI & Emerging Tech Reporter
James covers AI, agentic AI systems, ESG investing, gaming innovation, smart farming, telecommunications, and AI in film production. Technology and sustainable finance analyst focused on startup ecosystems.
James Park is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →
Frequently Asked Questions
What does ThinkingBox measure instead of an agent's final response?
According to the joint Microsoft and Hugging Face blog post, ThinkingBox grades agents on the terminal backend state and side effects they leave behind. The post states that final responses and valid tool calls are only proxies, and that only the records an agent leaves behind settle whether the work was done.
How many tasks and trials does ThinkingBox-Bench cover?
The benchmark covers 507 stateful business workflows, each run 20 independent times from an identical clean backend. A common-set ablation covered 121,680 valid trials across 12 LLM models, of which 79,853 attempts failed the executable checks, according to the post.
What were the most common failure signatures?
The post reports that tool usage accounted for 79.9% of failures, wrong state updates 10.3%, incomplete user resolutions 7.0% and no state-changing action 2.9%. It describes these as unweighted averages of per-model shares and observable labels, not unique causal explanations.
Which models led the reported pass@1 results?
Claude Opus 5.5 leads the task-weighted pass@1 table at 67.16%, two-thirds of a point above Claude Opus 5. Kimi-K3 is the strongest open-weights model at 57.37%, and it also had the broadest coverage, solving 476 of 507 tasks at least once.
How can teams run the benchmark themselves?
The post states ThinkingBox is available on Hugging Face, both the harness and the dataset, with ThinkingBox-Bench behind the OpenEnv interface. The documented setup runs on Linux and WSL with Python 3.11 or later, uv and Docker, and requires a thinkingbox-data checkout at the pinned release plus model endpoints. The source does not provide a hosted service for running it.