Back to articles
AI Agents

Learning from Failure: AED Scales Agent Error Diagnosis to 50,000 Pairs

3 min read

Introduction

For an LLM agent, a failed rollout contains more information than a zero reward. The model’s observations, selected actions, and the environment’s responses can reveal where the trajectory went wrong and what should have happened next. The practical challenge is to identify the decision worth revising and test a concrete alternative instead of treating failure as a single negative label.

The Agent Error Dataset, or AED, is designed around this problem. The dataset contains 50,228 error-diagnosis pairs collected from 9,961 source tasks. Its coverage spans 33 text-based environments, 19 harness families, and 23 policy models. Rather than storing only a diagnosis, the release retains source traces and execution metadata. This allows researchers to study failures across settings and perform a new diagnosis without rerunning the original rollout.

How the pipeline works

The accompanying Agentic Error-to-Training pipeline follows five broad stages. It first gathers naturally occurring failures. A diagnostic system then identifies the problematic decision and proposes a correction using the available observations, actions, and environment responses. Those diagnoses are checked against recorded evidence. When an environment supports replay, the proposed correction is compared with a retry of the original action from the same checkpoint under matched execution conditions. The resulting data is organized into separate views for diagnosis training and actor recovery.

This separation is important. A model can know that a trajectory failed without knowing why it failed, and it can identify a mistake without producing a useful next action. AED therefore treats explanation and recovery as related but distinct learning targets.

Main findings

  • Across 3,062 matched replay pairs, first-proposal corrections increased verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points.
  • On a separately frozen diagnosis release, fine-tuning Qwen3-8B on 1,656 source tasks raised exact-step agreement with internal teacher labels from 47.2% to 63.6% on a 943-case holdout, averaged across three seeds.
  • The strongest prompted reference in that comparison reached 54.7%, while mean agreement improved at each of four increasing training-set sizes.
  • In a single-seed actor-training comparison, action-only repair training scored 6.67 percentage points higher than success-only training on WebShop-lite.

Why it matters

AED’s contribution is not simply its sample count. It proposes a reusable loop that links a failure trace to a diagnosis, a candidate replacement action, and execution-based verification. This makes unsuccessful experience more actionable for post-training and may be particularly useful for browsing, tool use, and other tasks where a small decision can determine the rest of a trajectory.

The findings also need to be read within their experimental scope. Replay-based evaluation is available only in some environments. Diagnosis agreement is measured against the paper’s internal teacher labels, and the WebShop-lite actor comparison uses a single seed. It remains open whether the approach will transfer reliably to unseen environments, longer-horizon tasks, and different tool ecosystems.

Even with those caveats, AED points toward a more failure-aware approach to agent training: instead of collecting only more successful traces, systems can preserve and test what went wrong, then turn that evidence into targeted recovery behavior.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles