Skip to content
Read the original: OpenBMB· Published 42/100AI score42/100

Diffusion Reward Models learn full human preference distributions, not single scores

Original titleA reward of “3” can mean two completely different things.

AISummary

OpenBMB introduces Diffusion Reward Models (DRM), which learn the full reward distribution of human preferences instead of collapsing them into one scalar score.

The approach preserves disagreement and uncertainty, enabling distribution-aware Best-of-N ranking and a new test-time scaling axis by sampling more reward outputs.

DRM also improves downstream policy performance over scalar reward baselines when used as the reward in RLHF, according to the post.

Read the original x.com

Source: OpenBMB · x.comPublished · added here