UndoBench: Why AI Agents Must Be Tested on Recovery, Not Just Completion
Introduction
For an AI agent operating enterprise software, completing a workflow under normal conditions is only the first test. The more consequential question is what happens when a tool call times out, an acknowledgment is lost, or the system cannot tell whether a mutation has already occurred. A blind retry may make the workflow appear successful while creating duplicate payments, records, or other external side effects.
UndoBench is designed to measure that gap. Instead of treating a task’s final pass or fail result as a single measure of agent quality, it separates nominal task competence from operational recovery capability.
Key findings
- Counterfactual pairing isolates the effect of faults. The benchmark covers 36 base workflows and 36 fault scenarios across eight enterprise domains. Normal and fault-injected executions are paired under identical seeds, while wire-level effect histories and environment-state oracles are used to determine what actually happened.
- Completion is not the same as safe recovery. In the frozen lost-acknowledgment study, 12 held-out test workflows produced 5,760 executions, or 2,880 paired trials. Nominal competence reached 83.54%, but conditional recovery success rate fell to 46.72%. Naive retry generated duplicate external effects in 53.33% of trials.
- The recovery phase matters. Before any mutation, the evaluated methods performed similarly and capable runs did not create duplicate effects. During partial mutation, naive retry, per-call idempotency, and zero-privilege journaling all collapsed on the tested composite workflows. After commit but before acknowledgment, verification and server-side idempotency provided substantially safer behavior.
- Recovery has an operational price. Faulted runs showed 29.3% higher latency, 10.5% more completion tokens, and 55.8% more tool calls. A run that eventually succeeds may still be too expensive or too operationally risky for production.
Why it matters
UndoBench’s broader contribution is methodological. Conventional agent benchmarks often inspect only the final state, which can hide unsafe intermediate actions. For payment, ticketing, inventory, access-control, or data-entry workflows, a duplicate action may be more damaging than a visible failure.
The findings also suggest that recovery should not be left entirely to the model. Production systems need verifiable state queries, effect histories, server-side idempotency keys, and audit mechanisms that work within the agent’s permission boundary. Evaluation should report not only whether a faulted run eventually completes, but also whether it creates duplicates and how much extra time, token usage, and tool activity it requires.
For model and framework developers, a high nominal success rate is therefore insufficient evidence of reliable tool use. Fault injection and phase-specific recovery tests should become part of pre-deployment evaluation, especially for workflows with irreversible external consequences.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...