Back to articles
Reinforcement Learning

Diffusion Reward Models Move RLHF Beyond a Single Preference Score

3 min read

Introduction

Reward models are central to reinforcement learning from human feedback. They translate judgments about helpfulness, safety, or quality into signals that can optimize a language model. Most existing designs, however, compress each prompt–response pair into one scalar reward, or assume that the reward follows a predefined parametric distribution. This makes training convenient, but it can hide genuine disagreement: the same answer may be judged positively for different reasons, and several evaluations may be reasonable at the same time.

A paper from the Tsinghua NLP Group proposes Diffusion Reward Model, or DRM, to model that uncertainty directly. Instead of estimating only a point value, DRM formulates reward modeling as conditional density estimation over p(r|x,y), where the input prompt and response condition a distribution of possible rewards.

How DRM works

  • Diffusion-based reward generation: A frozen large language model encodes the prompt and response. A lightweight Diffusion Transformer then starts from Gaussian noise and denoises it into a reward vector.
  • No fixed output family: Unlike a head tied to a Gaussian or another predefined family, the diffusion process does not impose a simple parametric shape on the reward distribution. This makes multimodal outcomes easier to represent.
  • One architecture for different supervision: The same design handles multi-attribute regression as well as pairwise preference data.
  • Distribution-aware inference: Multiple samples for one input form an empirical reward distribution. Those samples can be summarized by a mean-like scalar, a variance, or quantiles, depending on the decision being made.

What the paper reports

Across five benchmarks, DRM matches or outperforms baselines when the data and backbone are controlled. Despite its modest training scale, it remains competitive with substantially larger discriminative, distributional, and generative reward models. The paper also reports that conventional reward heads tend to collapse multimodal preference structure into a point, while DRM can recover multiple modes.

The distribution is useful beyond analysis. The authors test uncertainty-aware rejection and lower-confidence-bound aggregation. These strategies allow a system to consider not only how high a predicted reward is, but also how uncertain or risky that estimate may be. In downstream RLHF experiments, using DRM as the training-time reward leads to improved policy performance, providing evidence that the extra structure can affect optimization rather than merely produce richer plots.

Why it matters

DRM points toward a broader view of reward modeling. A reward model can be a conditional generative model of judgments, rather than a deterministic scoring function. Such a representation may be useful for alignment settings where annotators disagree, multiple quality dimensions coexist, or a confident but fragile score could lead the policy in the wrong direction.

The available material does not provide detailed scores for each benchmark, sampling costs, or a full comparison of aggregation methods. Therefore, the results should not be read as proof that diffusion rewards universally replace conventional heads. Deployment will still require a trade-off between additional sampling and the value of uncertainty information, as well as checks that the learned distribution reflects the intended population of preferences. Even with those caveats, DRM offers a concrete route for moving RLHF from scalar rewards toward structured and uncertainty-aware feedback.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
VLA-Precision: Making Real-World Online RL More Stable for VLA Robots
Reinforcement Learning
cctest.ai

VLA-Precision: Making Real-World Online RL More Stable for VLA Robots

VLA-Precision addresses policy drift and system overhead in real-world online reinforcement learning for vision-language-action models. Its ACoB algorithm combines staged value calibration with reference-regularized updates, while ACoB-Stream targets the throughput bottleneck of large VLAs.

Read more