Diffusion Reward Models Move RLHF Beyond a Single Preference Score
Introduction
Reward models are central to reinforcement learning from human feedback. They translate judgments about helpfulness, safety, or quality into signals that can optimize a language model. Most existing designs, however, compress each prompt–response pair into one scalar reward, or assume that the reward follows a predefined parametric distribution. This makes training convenient, but it can hide genuine disagreement: the same answer may be judged positively for different reasons, and several evaluations may be reasonable at the same time.
A paper from the Tsinghua NLP Group proposes Diffusion Reward Model, or DRM, to model that uncertainty directly. Instead of estimating only a point value, DRM formulates reward modeling as conditional density estimation over p(r|x,y), where the input prompt and response condition a distribution of possible rewards.
How DRM works
- Diffusion-based reward generation: A frozen large language model encodes the prompt and response. A lightweight Diffusion Transformer then starts from Gaussian noise and denoises it into a reward vector.
- No fixed output family: Unlike a head tied to a Gaussian or another predefined family, the diffusion process does not impose a simple parametric shape on the reward distribution. This makes multimodal outcomes easier to represent.
- One architecture for different supervision: The same design handles multi-attribute regression as well as pairwise preference data.
- Distribution-aware inference: Multiple samples for one input form an empirical reward distribution. Those samples can be summarized by a mean-like scalar, a variance, or quantiles, depending on the decision being made.
What the paper reports
Across five benchmarks, DRM matches or outperforms baselines when the data and backbone are controlled. Despite its modest training scale, it remains competitive with substantially larger discriminative, distributional, and generative reward models. The paper also reports that conventional reward heads tend to collapse multimodal preference structure into a point, while DRM can recover multiple modes.
The distribution is useful beyond analysis. The authors test uncertainty-aware rejection and lower-confidence-bound aggregation. These strategies allow a system to consider not only how high a predicted reward is, but also how uncertain or risky that estimate may be. In downstream RLHF experiments, using DRM as the training-time reward leads to improved policy performance, providing evidence that the extra structure can affect optimization rather than merely produce richer plots.
Why it matters
DRM points toward a broader view of reward modeling. A reward model can be a conditional generative model of judgments, rather than a deterministic scoring function. Such a representation may be useful for alignment settings where annotators disagree, multiple quality dimensions coexist, or a confident but fragile score could lead the policy in the wrong direction.
The available material does not provide detailed scores for each benchmark, sampling costs, or a full comparison of aggregation methods. Therefore, the results should not be read as proof that diffusion rewards universally replace conventional heads. Deployment will still require a trade-off between additional sampling and the value of uncertainty information, as well as checks that the learned distribution reflects the intended population of preferences. Even with those caveats, DRM offers a concrete route for moving RLHF from scalar rewards toward structured and uncertainty-aware feedback.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...