Back to articles
Reinforcement Learning

Scaling On-Policy Distillation: How Small Teachers Can Train Larger Reasoning Models

3 min read

Introduction

Reinforcement learning can unlock substantial reasoning ability in large language models, but transferring that ability to another model is not simply a matter of copying outputs. The result depends on the relationship between teacher and student, their initial capabilities, and how supervision is delivered. The paper Scaling Properties of Same-Family On-Policy Distillation studies this question through on-policy distillation (OPD), where the student generates its own responses and is then trained with token-level reverse-KL supervision from a teacher policy.

The authors examine weak-to-strong, same-base, and strong-to-weak teacher–student configurations within the same model family. They also provide an interactive project page with trajectory replays and a peak-performance grid, as well as 311 training checkpoints spanning Qwen2.5 models from 0.5B to 14B parameters. These resources make the scaling trends easier to inspect and reproduce.

Main findings

  • Early OPD follows a useful-transfer regime. The study measures training progress through the square root of reverse KL divergence between the student policy and its initialization. During the early phase, held-out gold score rises approximately linearly with this distance, suggesting a relatively regular window in which policy updates translate into capability gains.
  • Weak teachers can produce stronger students. In every observed weak-to-strong pairing, the student’s peak gold score exceeded the teacher’s own score. The teacher therefore acts more like a directional source of useful behavior than a hard ceiling on the student’s performance.
  • Teacher scaling has a practical limit. Increasing teacher size improves peak gold score only until the teacher is roughly comparable to the student in scale. Beyond that point, additional teacher parameters do not automatically produce proportional gains for the student.
  • A teacher’s score is not its whole supervision value. When teachers have matched gold scores, smaller teachers can transfer more effectively. This means that evaluation performance alone is insufficient for selecting a teacher; the structure of its policy and its relation to the student also matter.
  • The study models several scaling effects. Power laws are fitted for peak gold score and the slope of the useful-transfer regime as functions of student size, teacher size, and teacher score. The authors also investigate OPD variants, bootstrapping in weak-to-strong training, and the amount of on-policy supervision.

Why it matters

The paper reframes distillation as a resource-allocation problem. Instead of assuming that the largest available teacher is always optimal, practitioners may need to balance teacher size, student size, teacher quality, and the efficiency of the supervision signal. A compact RL specialist could be trained first and then used to guide a much larger model, potentially reducing the cost of building a strong student.

The results also caution against using a single benchmark score to estimate a teacher’s usefulness. Two teachers with similar scores may induce different learning dynamics, and the more compact one may sometimes provide a better training direction. At the same time, the findings should not be generalized without qualification: the experiments focus on one model family and a defined OPD setup. Whether the reported power laws hold across architectures, tasks, reward systems, and cross-family transfer remains an open question.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles