Why LLM Reinforcement Learning Drifts Between Training and Inference
Introduction
Reinforcement learning with verifiable rewards (RLVR) often separates generation from optimization. An inference engine samples responses, while a training engine evaluates the sampled tokens and computes policy gradients. In theory, both engines should represent the same policy. In practice, they may assign different probabilities even when they use identical model weights. Numerical precision, hardware, parallel execution, and implementation details can all contribute to this training-inference mismatch.
The paper Rethinking Training-Inference Mismatch in LLM Reinforcement Learning proposes Calibrated Importance Sampling, or CIS, to account for the discrepancy. Rather than applying a fixed cap to importance ratios, the method first studies how the mismatch appears in probability space and then makes the correction depend on token confidence.
Key ideas
- The mismatch is a systems-level problem. Rollouts come from the inference engine, but gradients are calculated by the training engine. If their token probabilities differ, importance ratios can become unusually large or small, making policy updates unstable or inaccurate.
- Log-odds provide a useful description. The authors’ empirical characterization treats the mismatch as an additive displacement in log-odds. This displacement is linked to the per-logit perturbation before the softmax, and its distribution is approximately invariant to the confidence of the selected token.
- Clipping should depend on confidence. CIS truncates sufficiently large positive displacements at one constant threshold in displacement space. When translated back into importance-ratio space, this produces a stricter cap for high-confidence tokens and a looser effective cap for low-confidence tokens. The goal is to avoid discarding too much useful signal from uncertain tokens.
- Variance is traded for controlled bias. The theoretical analysis shows that CIS replaces the potentially unbounded second moment governing the error of exact importance sampling with a term bounded by a constant. The price is a bias related to the excess that is truncated. Diagnostic results also suggest that clipping small importance weights upward can hurt held-out accuracy, so stabilization is not automatically beneficial in every direction.
Results and implications
The method was evaluated on three mixture-of-experts models and five mathematical reasoning benchmarks. CIS obtained the highest five-benchmark average among the evaluated baselines for all three models. Additional diagnostics found that CIS introduced less truncation bias on low-confidence tokens than ordinary truncated importance sampling.
The broader lesson is that infrastructure details can become algorithmic issues in LLM reinforcement learning. A small probability discrepancy between engines may be amplified by importance ratios and eventually change the direction or magnitude of policy updates. This makes coordination between training frameworks, inference runtimes, numerical formats, and parallelization strategies an important part of RLVR design.
CIS also offers a practical analytical lens: instead of treating every ratio as equally risky, it links correction strength to the confidence of the token. That may be useful for systems where rollout generation and optimization are deliberately decoupled. At the same time, the reported evidence is limited to the models and mathematical tasks studied in the paper. Broader tasks, engine combinations, and training settings will be needed to determine how consistently the method transfers.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...