OPD² Reframes On-Policy Distillation Around What Reasoning Tuning Adds
On-Policy Delta Distillation, or OPD², replaces direct imitation of a teacher’s full output distribution with a signal based on the difference between a reasoning-tuned teacher and its base model. The paper reports consistent gains across math, science, and code-reasoning benchmarks.
Read more