Uncertainty-Normalized Margins Enhance Direct Preference Optimization
Researchers introduce a new method (UNM-DPO) for direct preference optimization in language models by incorporating prompt-dependent uncertainty and preference-strength normalization. ULNM-DPO-WR, a variant with length normalization, outperforms prior DPO methods on several evaluation benchmarks.
The paper proposes a novel family of training objectives to address limitations in traditional Direct Preference Optimization (DPO), which previously modeled preferences using a fixed-noise Bradley-Terry approach. The new approach, uncertainty-normalized margin DPO (UNM-DPO), allows the optimization to account for prompt-dependent uncertainty and preference strength.

What Changed?
- DPO has been extended to include strength-dependent margins and a learned prompt-dependent noise scale.
- Two new objectives were introduced: advantage-only (AO) and whole-residual (WR), differing in how margin and scale are applied during optimization.
- A practical method for learning the prompt scale was presented.
- ULNM-DPO-WR further normalizes reward by response length, showing improved performance.
Evaluation Results
On HelpSteer2 and HelpSteer3 (using Skywork reward model), ULNM-DPO-WR achieved tie-adjusted win rates of 68.00% and 65.31% versus matched DPO, with higher mean rewards and shorter responses. On AlpacaEval (evaluated with GPT-4.1), it reached a length-controlled win rate of 21.62%, outperforming DPO (16.39%) and SimPO (15.30%).
What Developers Should Do
- Consider adopting UNM-DPO or its length-normalized variant ULNM-DPO-WR when training models with human preference data, especially if task-specific uncertainty or variable response lengths are factors.
