GRAFT Lets RLVR Models Learn from Peer Reasoning Trajectories
Introduction
Reinforcement Learning with Verifiable Rewards (RLVR) trains a model by checking whether its generated answer is correct. Methods such as GRPO typically sample several responses for a prompt, compare their verifiable rewards, and use the resulting group-level advantages to update the policy. This design is effective when at least one useful response appears in the group. It becomes much less useful when the entire group fails: there is no successful sample to provide a meaningful reward contrast, even though additional rollouts might eventually uncover one.
The paper Learning Beyond What You Sample proposes GRAFT, or Gated Replacement of Answer-Failed groups with peer Trajectories. Its central observation is that heterogeneous reasoning models do not fail on exactly the same prompts. A response that one model misses may already have been found by another. Rather than treating each model’s rollout pool as an isolated source of experience, GRAFT turns these differences into an opportunity for mutual learning without requiring a permanently designated teacher.
Key ideas
- Replace uninformative groups: When a model produces only failed answers for a group, GRAFT can substitute trajectories generated by a peer model on the same prompts.
- Share both positive and negative evidence: The framework transfers successful and unsuccessful peer responses, together with advantages computed under the peer’s own rollout group. This preserves comparative information instead of reducing exchange to copying correct answers.
- Control the off-policy gap: Peer trajectories come from a different policy, so using them naively could produce unstable updates. GRAFT applies sequence-level compatibility weighting and clips token-level importance ratios to limit the effect of large policy mismatches.
- Enable trajectory reuse: The reported results indicate that stored peer trajectories retain most of the improvement, suggesting that exchange need not always happen through simultaneous co-training.
Results and implications
The authors evaluate GRAFT with three heterogeneous model pairs on five mathematical reasoning benchmarks. Against GRPO under the same per-model rollout budget, both models improve consistently. The reported average gain is 2.1 points, with up to a 4.5-point improvement in model-level average performance. When previously stored peer trajectories are used instead of simultaneous exchange, the method still improves over GRPO by 1.8 points on average.
The broader lesson is that an RLVR system’s useful experience need not be limited to what one policy happens to sample. A collection of models with different strengths can act as a complementary exploration pool, potentially reducing the waste caused by all-fail groups. The work also highlights an important constraint: cross-model sharing is not automatically safe. Differences between policies must be measured and controlled, and the quality of the update depends on compatibility weighting and importance-ratio clipping.
GRAFT therefore offers a practical direction for trajectory caching, cooperative RLVR, and teacher-free multi-model training. Its contribution is less about simply increasing the number of rollouts than about making previously discovered evidence available to the model that failed to find it.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...