Back to articles
AI Agents

X-Tree Turns Reusable Experience into a Vocabulary for AI Agents

3 min read

Introduction

Many multi-step agents do not fail because they cannot perform an individual action. They fail because a procedure learned in one task is not reliably reused in another. Standard supervised fine-tuning and reinforcement learning generally flatten a trajectory into an action stream, assigning similar importance to individual tokens or actions. As a result, routines such as filling two fields and submitting a form, or navigating to a location to retrieve an object, may be relearned repeatedly across tasks.

X-Tree addresses this problem by recovering the hierarchy already hidden in existing trajectories and making it part of training. The authors call the resulting structure a reusable eXperience tree.

Mining skills like a tokenizer

Rather than asking an LLM to write skill descriptions, X-Tree uses a data-driven merging process inspired by subword tokenizers. Raw actions are first canonicalized into comparable symbols. An action such as entering a date or clicking a button retains its type, while page-specific element IDs and values are represented as slots. Similar behavior on different pages can therefore be recognized as the same kind of operation.

The system then scores adjacent spans according to their reuse value. The score considers how often a span recurs, how much longer the merged unit is, and how often it appears in successful episodes. A merge is retained only when the reduction in the corpus justifies its cost. Repeated merges build a hierarchy: short actions become reusable subroutines, which can then compose into higher-level procedures.

For example, filling two fields and submitting may form one node. Combined with preceding navigation, it can become a broader routine for filtering a table by a date range. Because the process is based on normalization and counting, the resulting tree is deterministic and auditable, with no additional LLM calls.

Three ways to use the tree

The paper integrates X-Tree into three training settings:

  • Offline reinforcement learning: each tree node becomes a training instance, exposing the model to program-level experience rather than only raw action sequences.
  • Online RLVR: the agent receives an adaptive skill bonus that reinforces valuable and reusable behavior during interaction.
  • On-policy self-distillation: the tree serves as privileged context for the self-teacher, helping the student extract higher-level structure from trajectories.

This differs from systems that place LLM-written skills only in the prompt. Such skills may improve the current context but do not necessarily enter the model’s weights or generalize once retrieval is unavailable. X-Tree instead turns recurring experience into a training signal.

Results and broader significance

Across WebArena, ScienceWorld, and WebShop, and at three model scales, X-Tree outperformed standard recipes under matched data and budget. The reported gains reached 4.5 percentage points in WebArena success rate, 5.8 points in ScienceWorld, and 4.1 points in WebShop success. Matched analyses attributed the improvements to the tree structure and its three training integrations, rather than simply to adding more context.

The broader contribution is a shift from manually prompted skills to skills discovered from agent data. When trajectories are expensive to collect, program-level segmentation can make each episode more useful and provide a common interface for offline data curation, reward shaping, and capability transfer. The approach still depends on sensible action canonicalization and sufficiently informative successful trajectories. Procedures that are tightly bound to page context may remain difficult to abstract, but X-Tree offers a practical route for turning repeated behavior into trainable structure.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Why LLM Agents Keep Following Instructions Users Have Withdrawn
AI Agents
cctest.ai
AI Agents

Why LLM Agents Keep Following Instructions Users Have Withdrawn

A new study turns intent drift into a measurable failure mode: superseded or withdrawn user requirements can continue to shape an agent’s answer or tool action. Its IntentFlux benchmark and StateForge method show that explicit state maintenance helps, but does not eliminate the gap between evolving dialogues and direct single-turn tasks.

Read more