Back to articles
Reinforcement Learning

Cross-Tokenizer Distillation Needs Reliable Supervision, Not Maximum Coverage

3 min read

On-policy distillation (OPD) trains a student on trajectories produced by the student itself, while a teacher supplies probability-based feedback. This setup can provide supervision on the states the student actually visits. However, when the teacher and student use different tokenizers, their predictions cannot be compared directly. Alignment must be solved at both the sequence level and the vocabulary level.

The paper asks whether increasing alignment coverage is necessarily the right way to address this problem. The authors evaluate three heterogeneous teacher–student pairs on mathematical reasoning and code generation. Their results challenge the assumption that more supervised positions automatically translate into better distillation.

The main findings are:

  • Strict one-to-one alignment groups already cover most tokens generated by the student, despite substantial differences between the two vocabularies.
  • On responses sampled from students before distillation, the shared vocabulary at strictly aligned positions preserves nearly all of the teacher’s and student’s probability mass on average.
  • Restricting the reverse KL objective to a student-selected top-16 subset of the shared vocabulary reaches accuracy comparable to full shared-vocabulary OPD and outperforms the evaluated cross-tokenizer baselines.
  • Adding mean squared error supervision on span log-probabilities for mismatch groups achieves broader, eventually complete coverage, but lowers accuracy.

The authors also examine why extra supervision can be harmful. At checkpoints trained with only the strict loss, gradients from the mismatch-span objective show weak directional agreement with strict-loss gradients and can even point in opposing directions. Their magnitude also grows relative to the strict gradients. In other words, the added objective may fill a formal coverage gap while pushing optimization away from the direction supported by the more reliable signal.

This leads to a broader design lesson for cross-tokenizer OPD. Coverage should not be treated as the only quality metric. A useful supervision signal must also preserve meaningful probability mass, remain numerically stable, and produce gradients that cooperate with the rest of the training objective. A compact candidate set at positions with clear alignment may therefore be more effective than forcing a loss onto every unmatched span.

The result is not a claim that mismatch supervision is universally harmful. The experiments cover a limited set of model pairings and tasks, so the conclusions should be interpreted as diagnostic rather than universal. Still, the study offers a practical evaluation checklist: measure alignment coverage, inspect probability mass, and compare gradient direction and scale before adding broader losses. For OPD systems with heterogeneous tokenizers, improving the reliability of supervision may matter more than maximizing its nominal extent.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
TRACE Aligns FP4 Training and Rollouts for MoE Reinforcement Learning
Reinforcement Learning
cctest.ai

TRACE Aligns FP4 Training and Rollouts for MoE Reinforcement Learning

TRACE addresses the mismatch between training and rollout quantization in reinforcement learning for MoE language models. It uses rollout-side quantization outcomes to guide training-side FP4 rounding and reports up to 5.4× rollout speedup while preserving performance close to BF16 rollouts.

Read more