TRACE Aligns FP4 Training and Rollouts for MoE Reinforcement Learning
Reinforcement learning has become a central part of post-training large language models, but its cost is not limited to updating model parameters. The model must repeatedly generate rollout samples, making compute, memory usage, and KV-cache capacity major bottlenecks. Lowering rollout precision to FP4 is an attractive way to reduce those costs. Yet simply quantizing a BF16-trained policy after training does not necessarily preserve the behavior needed by the reinforcement-learning loop.
The issue is a path mismatch
Existing FP4 reinforcement-learning approaches generally optimize quantization quality on the training path and the rollout path separately. Each path may have a small local quantization error, but the two paths can still produce meaningfully different results. This matters in reinforcement learning because the policy continuously generates the data used for further optimization. If the training computation reflects one quantized model while rollout generation uses another effective computation path, the policy can be trained on behavior that does not match the behavior used to collect new samples.
TRACE, short for Train-Rollout Quantization Alignment via Compact Guidance, focuses on this discrepancy. The framework is designed for mixture-of-experts language models and treats alignment between the two execution paths as a first-class quantization objective.
How TRACE works
- Rollout-guided quantization-aware training: TRACE feeds quantization outcomes from the rollout side back into training. These outcomes guide the FP4 rounding choices made during training, encouraging the training computation to follow the low-precision behavior that will actually be used for generation.
- FP4 across the generation pipeline: The framework supports FP4 weights and activations, while also enabling FP4 KV-cache rollouts. This extends the optimization beyond model parameters to the cache that can consume substantial memory during long generation.
- Compact quantization guidance: Passing rollout information into training can create storage and communication overhead. TRACE reduces that cost by selectively retaining mantissa and scale information from deeper layers instead of keeping all quantization details.
Results and significance
The paper evaluates TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon reinforcement-learning tasks. Its results indicate that joint FP4 quantization of weights, activations, and KV caches can achieve reinforcement-learning performance comparable to BF16 rollouts. Compared with post-hoc FP4 quantization applied to BF16-trained policies, TRACE also delivers stronger final FP4 performance. The reported maximum rollout speedup is 5.4×.
The broader contribution is a shift in how low-bit reinforcement learning is framed. Rather than asking only how closely quantization can reproduce a high-precision model, TRACE asks whether the training process can adapt to the precise low-precision execution path used during rollout. That distinction is particularly relevant to MoE systems, where repeated generation and cache movement can make rollout efficiency a practical bottleneck.
The available material provides aggregate claims rather than detailed breakdowns for every model, task, or hardware configuration. Therefore, the 5.4× figure should be read as the maximum reported speedup, not a guaranteed gain in every deployment. Even so, TRACE supports the idea that execution-path alignment may be as important as minimizing standalone quantization error in low-precision RL.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...