Marathoner: Training AI Agents for Ultra-Long-Horizon Execution
Introduction
Many AI agents can call tools, write code, and complete short multi-step workflows. Their performance becomes less reliable, however, when a task requires hours of sustained planning and execution. They may drift from the original objective, repeat failed actions, or stop making meaningful progress near the end. Marathoner, proposed by a research team from Ant Group, targets this specific gap: the ability to keep pursuing a difficult goal across an ultra-long execution horizon.
What the method does
Marathoner is not presented as a new model architecture. It is a post-training pipeline designed around long-running tasks.
- Task synthesis from substantial software changes. The researchers use major GitHub release pull requests containing more than 1,000 lines of new code as a primary source for constructing challenging task-level data. These examples provide longer dependency chains and more realistic failure modes than simple coding prompts.
- Multi-task chaining. Several generated tasks are connected into one larger assignment. This creates trajectories with greater depth and aims to approximate frontier-level difficulty.
- Rejection-sampling fine-tuning. A strong teacher model, together with multiple execution harnesses, produces candidate trajectories on the synthesized tasks. Selected successful trajectories are then used to supervise the base model.
- Reinforcement learning in sandboxes. The cold-started model performs tasks in independent environments during rollout. This exposes it to actual tool use, changing intermediate states, and the need to recover from errors rather than merely imitate a static answer.
- A later-stage bonus. The Later Stage Bonus Reward explicitly encourages useful actions in the latter part of an execution. It is intended to address a common long-horizon failure pattern in which an agent is active early but gradually becomes inactive or ineffective.
Results and implications
The paper evaluates Marathoner on five benchmarks containing ultra-long-horizon tasks and reports consistent, substantial improvements over the base model. It also claims that the system surpasses a strong proprietary model in the reported comparisons. Additional analysis says Marathoner can work for more than 10 hours and make over 1,000 tool calls on highly challenging tasks. The supplied material does not provide benchmark names, detailed scores, or the full configuration of the comparison models, so those claims require verification against the complete paper and reproducibility materials.
The broader contribution is an attempt to make persistence a trainable and measurable agent capability. Large software changes offer a practical source of tasks with extended dependencies and concrete end states. Sandbox rollouts move training beyond text-only imitation, while the late-stage reward focuses directly on degradation that appears after many steps.
Longer trajectories do not automatically imply better understanding or reliable autonomy. They can also amplify early mistakes, increase inference cost, and create new requirements for state tracking and final-result validation. Marathoner should therefore be viewed as a training direction rather than definitive evidence of general-purpose autonomous intelligence. Its performance on non-programming tasks, open environments, and stricter reliability tests remains an important open question.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...