Microsoft and Hugging Face have jointly released ThinkingBox, an agent evaluation framework, along with its companion benchmark ThinkingBox-Bench. It scores agents based on what actually happened in the backend database after the work was done — not on the text where the agent itself reports "task complete."
The benchmark comprises 507 stateful business-process tasks across five industries: 98 in retail, 100 in auto insurance, 104 in travel, 104 in digital banking, and 101 in consulting. Each task runs independently 20 times against a clean backend, and executable check scripts verify the database's final state and any side effects — whether the fields that should have changed actually changed correctly, and whether the agent accidentally touched records it shouldn't have.
The team summed up the problem in one line on their blog:
"The Agent Said It Was Done. The Database Disagreed."
Results Across 18 Models
Eighteen models were tested, spanning both closed and open-weight systems. By the "passed all 20 runs" measure, Claude Opus 5.5 and Claude Opus 5 tied for first place, each reliably solving 241 tasks — 47.53%. GPT-6 Astra came in third with 231 tasks, or 45.56%. In other words, even the best-performing models fail to get it right every single time on more than half of the tasks.
The widest coverage belongs to Kimi-K3: of the 507 tasks, it got 476 right at least once, a 93.89% hit rate. But only 68 of those held up across all 20 runs. The same model looks almost omnicapable by pass@20, yet falls toward the back of the pack on reliability — a sevenfold gap between the two numbers.
The breakdown of failure causes gets more specific. Among 79,853 attempts that "ran normally but produced the wrong result," 77.61% involved an incorrectly written field value, and 43.30% caused an unintended side effect (the two can overlap). The team's assessment: roughly 80% of failures trace back to the tool-calling step — wrong parameters, wrong call order — while errors in the reasoning itself account for a much smaller share.
The team also reports a "cost per reliable task" figure — how much it costs on average to get a task right across all 20 runs: $6.80 for GPT-5.4, $7.45 for GPT-6 Astra, and $7.80 for Claude Opus 5.5. The older-generation GPT-5.4 actually comes out ahead of the newer models on this metric.
How It Compares to Existing Agent Benchmarks
ThinkingBox isn't the first to measure reliability through repeated runs. Sierra's τ-bench, released in 2024, introduced the pass^k metric, requiring a model to succeed on the same task k times in a row, covering retail and airline scenarios at the time. ThinkingBox differs in scale and scoring: it expands to 507 tasks across five industries and treats side effects as a standalone failure condition. Under τ-bench's rules, an agent that touches one extra, unrelated record wouldn't necessarily lose points.
This site previously covered Berkeley's Vero benchmark, which has agents complete 43 software projects — the best score so far is 27 finished. That benchmark measures "how big a job it can finish in one go." ThinkingBox points the opposite direction: the individual tasks aren't hard, but the question is whether handing the same job to an agent 20 times produces even one failure. In enterprise workflows, that second kind of failure is often the more dangerous one — getting a policy amount or a transfer field wrong once costs far more to fix than simply not finishing.
How to Use It
The code is open-sourced under the MIT license, the benchmark data uses CDLA-Permissive-2.0, and the OpenEnv environment is BSD-3-Clause — it can be run directly through the OpenEnv interface on Hugging Face. A local deployment needs Docker and Typesense to manage state, tools are served through an MCP server, and three model endpoints need to be configured: the agent under test, the simulated user, and the judge model.
Both the simulated user and the judge are played by models, which introduces its own margin of error. The blog doesn't separately disclose how closely the judge model's calls agree with human annotation — until that number is available, a one- or two-percentage-point gap between models probably shouldn't be read into too heavily. The project also involved interns from the University of Pittsburgh, UC Irvine, Northwestern University, and Columbia University.
Sources: Hugging Face Blog, Microsoft Research, CocoLoop, Sierra τ-bench paper; per-model pass counts, failure-type ratios, and per-task costs follow the figures published by the ThinkingBox team.