QwenGyre Targets GPU Idling and Trajectory Redundancy in Long-Horizon Agent RL
Introduction
As agents move beyond short responses toward repository-level coding and other multi-stage tasks, the main difficulty in online reinforcement learning is no longer only model throughput. A single rollout may last for hours, include hundreds of model–environment interactions, and produce close to one million tokens. Because different executions finish at very different times, parts of a training cluster can remain idle while other rollouts are still running. At the same time, branching task behavior can generate many paths that repeat essentially the same work.
Qwen’s paper presents QwenGyre as an end-to-end framework designed for this extreme-long, or xlong, setting. Its focus is not simply to make one inference faster, but to coordinate the execution and learning pipeline so that both compute and trajectory data are used more efficiently.
What the framework does
- Elastic resource allocation. QwenGyre dynamically reallocates GPUs between rollout generation and training. The claimed design allows live executions to continue while resources are moved, reducing the impact of uneven rollout durations on utilization.
- Branch-aware history reconstruction. Its trajectory processor reconstructs branching histories instead of treating every resulting path as an isolated sample. This gives the training pipeline a structured view of how different executions diverged.
- Partial-progress scoring. The system scores progress before a long task is fully completed. Such intermediate signals can provide learning information from unfinished executions rather than relying only on a final outcome.
- Trajectory deduplication. Redundant paths are identified and removed to keep the amount of training data under control. This is particularly relevant when non-linear task interaction produces many similar branches.
Reported results and implications
The paper reports an experiment using the Qwen 3.8 2.4T model with approximately 700K tokens per rollout. After 48 steps, QwenGyre raised the NL2RepoBench result from 52.5% to 58.5%, a six-percentage-point absolute improvement. Across evaluations involving different training-data domains, the authors report up to a 1.85× speedup over Colocate and up to a 1.78× speedup over Async.
The broader lesson is that long-horizon agent training is also a systems problem. Even a powerful model can be held back by stragglers, idle accelerators, and an uncontrolled number of nearly duplicate trajectories. QwenGyre brings scheduling, branch reconstruction, progress estimation, and deduplication into one training workflow, offering a possible infrastructure direction for coding agents and other applications that require sustained interaction.
The supplied material does not provide component-level ablations, detailed scheduling overheads, or evidence across a wider range of models and tasks. The reported gains should therefore be read as results under the paper’s stated settings, rather than a guarantee for every xlong-horizon workload.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...