Back to articles
Evaluation & Benchmarks

An AI Agent Can Say “Done” While the Database Says Otherwise

3 min read

Introduction

An AI agent can use the right tools, produce a polished response, and still fail the business task. ThinkingBox, introduced by Microsoft and Hugging Face, shifts evaluation from what an agent says to what it actually leaves in the system. The benchmark runs agents in isolated MCP tool sessions and checks the resulting backend state and side effects.

The hidden failure behind a successful trace

The source describes a retail case involving a costly kitchen appliance delayed for 15 days after a carrier exception. The agent made nine apparently sensible tool calls: it retrieved the order, checked tracking, reviewed the customer profile and refund policy, confirmed that no ticket existed, opened one, and documented the timeline. It also correctly determined that the customer did not qualify for late-delivery compensation.

Yet the run still failed. The carrier exception remained unresolved, so the ticket should have ended in a hold state. Instead, the agent marked it solved. Its final response also failed to provide a meaningful answer to the customer. A grader focused on tool calls might accept the trajectory; an executable database check would reject the resulting state.

Key takeaways

  • Outcomes matter more than traces. ThinkingBox checks whether fields have the required values, whether required side effects occurred, and whether unintended changes were introduced.
  • One success is not reliability. The benchmark covers 507 stateful business tasks. Each task starts from a clean backend and runs independently 20 times, with separate reporting for pass@1, pass@20, and observed 20/20 success.
  • Clean termination can hide failure. In a common-set ablation with 121,680 valid trials, 79,853 failed executable checks. Of those failures, 67.24% still ended cleanly, used a state-changing tool, and reported no final tool error.
  • Coverage and consistency diverge. Kimi-K3 succeeded at least once on 476 of 507 tasks, but only 68 tasks passed all 20 trials. Claude Opus 5 covered fewer tasks at least once, yet passed every trial on 241 tasks.

Why it matters

The benchmark exposes a problem in how enterprise agents are often evaluated. A strong pass@1 score shows that a model can complete a workflow, but it does not show that the model will write the correct value on the next attempt. In customer service, insurance, banking, or travel systems, an incorrect status, a missing update, or an extra side effect can be more damaging than an imperfect sentence.

Production evaluation should therefore include executable backend assertions and report several dimensions at once: how many tasks a model can solve at least once, how often a single attempt succeeds, and how many tasks it completes reliably across repeated runs. ThinkingBox is important less as another leaderboard than as a reminder that an agent trajectory is a claim, while the resulting system state is the evidence.

Source: Hugging Face Blog

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
PhysVista Tests VLM Physical Intelligence Through a Perception–Reasoning–Assessment Loop
Evaluation & Benchmarks
cctest.ai

PhysVista Tests VLM Physical Intelligence Through a Perception–Reasoning–Assessment Loop

PhysVista introduces a benchmark that evaluates whether vision-language models understand physical consistency rather than merely recognizing visual content. It combines physical state perception, dynamics reasoning, and plausibility assessment across real-world and AI-generated videos.

Read more
CCTest · Blog
OpenTumorBoard Tests Whether AI Can Reason With a Cancer Care Team
Evaluation & Benchmarks
cctest.ai

OpenTumorBoard Tests Whether AI Can Reason With a Cancer Care Team

OpenTumorBoard turns public multidisciplinary tumor board recordings into a benchmark for evaluating models on specialist answers and full clinical discussions. Its results show that even advanced general and medical models still struggle to reproduce expert responses and board-level consensus.

Read more