TraceDance Turns Agent Deployment Traces into Behavioral Benchmarks
Introduction
An agent can finish a task and still behave badly along the way. It may use a tool unnecessarily, overlook an important constraint, or make an unsafe choice at a critical step. Conventional benchmarks are usually designed around fixed tasks and expected outcomes, so they often miss the narrow behavioral failures that emerge only after deployment. TraceDance, presented by researchers including a ByteDance team, proposes a way to turn real deployment records into targeted tests for those failures.
How it works
TraceDance takes deployment traces from coding and general tool-use agents, together with a user-specified undesirable behavior. Its benchmark construction process combines three ideas:
- Anchor-and-Confirm: Programmable retrieval first identifies potentially relevant fragments from a large trace collection. A Flash LLM then reviews candidates individually, helping reduce the false positives that broad keyword or embedding search can produce.
- Anchor Synthesis Loop: If the requested behavior is not yet described precisely, the system generates a behavioral specification and revises it using evidence from retrieval and confirmation. The target therefore becomes more operational and testable over multiple iterations.
- Decision-point continuation: Rather than replaying the full environment or requiring a single reference answer, the benchmark stops at a recorded decision point. The evaluated model generates its next turn, which is judged using a rubric tailored to the behavior under investigation.
This shifts the evaluation question from “Did the agent complete the task?” to “Did it act appropriately while completing the task?” A successful final outcome does not automatically erase problematic intermediate decisions.
Results
The experiments drew on 252,557 sessions and produced 107 benchmarks with 4,125 instances. The system fulfilled 95.3% of requested benchmark-building targets. In sampled cases, human annotators confirmed that the requested behavior was actually present in 84% of instances. The automated grader’s agreement on pass/fail judgments was comparable to the agreement between human annotators, suggesting that automated scoring can support the workflow, although it should not be treated as infallible.
Across nine frontier LLMs, the mean pass rate was only 26.7%. The result indicates that broad language and task competence does not guarantee reliable behavior at specific agent decision points. The figure should be read as performance on these targeted behavioral tests, not as a universal ranking of overall model quality.
Why it matters
TraceDance suggests a practical loop from deployment issue to focused benchmark and then to model improvement. Developers can build regression tests around observed failures instead of waiting for a static benchmark suite to be updated. This could make agent evaluation more responsive and provide a testing component for recursive self-improvement workflows.
There are also clear caveats. Benchmark quality depends on the coverage of the source traces, the clarity of the requested behavior, and the reliability of the automated grader. The 84% human confirmation rate shows that retrieval and specification synthesis are not perfectly precise. High-stakes use will require broader human review, validation across tasks, and continued analysis of evaluation bias.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...