ASCENT Enables Long-Horizon Agents to Learn During Deployment
Introduction
A long-horizon agent rarely solves a task in one generation. It may need to reason, call tools, inspect the environment, and take actions across many turns before receiving a final success or failure signal. In deployment, such tasks often arrive as a stream of related problems. That stream creates a potentially valuable source of experience: trajectories produced on earlier tasks could help the agent perform better on later ones.
The difficulty is deciding what should actually be learned from one attempt. In-context adaptation systems commonly store reflections, memories, or skills as text. Their effectiveness then depends on retrieving the right item and on a frozen policy correctly following it. Directly imitating every token from a single rollout is also risky. A trajectory may contain invalid actions, inefficient choices, or errors that should not be reinforced. The result can be policy instability rather than useful adaptation.
How ASCENT works
ASCENT, short for Agentic Self-distillation for Cross-task EvolutioN at Test-time, addresses this problem with a hindsight-based self-distillation procedure:
- One-pass execution: The agent processes each task once as it moves through the task stream. The completed trajectory and its terminal verification result are the only learning signals available for persistent updates.
- A privileged teacher view: A frozen copy of the model’s initial state reads the verified trajectory. Because it can see the outcome and the complete execution history, it has information that was unavailable to the acting policy at each earlier step. It uses that hindsight to produce next-token distributions along the trajectory.
- Persistent fast-weight updates: Those distributions are distilled into LoRA parameters. The adapted weights remain available for subsequent tasks, while the system avoids treating every generated action from the original attempt as equally correct.
The paper also removes turns associated with invalid actions, creating a cleaner form of privileged experience. This is not conventional teacher-student training with a larger external model. The teacher signal comes from a stable, frozen version of the same model, supplied with information that becomes available only after the trajectory has been verified.
Results and implications
The authors report that ASCENT raises exact success by more than 22 points over the base model on ALFWorld and WebShop, at both model scales evaluated. It also reduces the number of turns per episode, and the improvements transfer to held-out scenes. These findings suggest that test-time training can draw useful supervision from an agent’s own verified experience without requiring an external reference solution or a stronger teacher.
The broader contribution is a shift from storing experience as retrievable text to modifying the executable policy itself. For systems that repeatedly handle related tasks, this could make adaptation less dependent on memory retrieval. At the same time, the setting has important boundaries: it assumes a terminal verification signal, and learning from a single trajectory can still be sensitive to noise and failure modes. Extending stable self-distillation to weaker, delayed, or more open-ended feedback will be essential before online agent training can become routine in deployment.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...