Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains

Luca Viano; Ruida Zhou; Yifan Sun; Mahdi Namazifar; Volkan Cevher; Shoham Sabach; and Mohammad Ghavamzadeh

arXiv:2602.00603·cs.LG·February 3, 2026

Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains

Luca Viano, Ruida Zhou, Yifan Sun, Mahdi Namazifar, Volkan Cevher, Shoham Sabach, and Mohammad Ghavamzadeh

PDF

Open Access

TL;DR

This paper introduces algorithms that leverage rating gap information to improve preference optimization in language models, achieving faster learning and robustness to inaccuracies, with strong empirical results across various benchmarks.

Contribution

The paper proposes new algorithms that utilize rating gap data for preference optimization, providing theoretical and empirical advantages over existing DPO methods.

Findings

01

Faster statistical convergence with accurate rating gaps.

02

Robustness of algorithms to rating gap inaccuracies.

03

Superior performance across multiple LLM benchmarks.

Abstract

The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with very limited feedback in the form of pairwise preferences and fine-tune models to align with these preferences without explicitly learning a reward model. While the form of feedback used by these algorithms makes the data collection process easy and relatively more accurate, its ambiguity in terms of the quality of responses could have negative implications. For example, it is not clear if a decrease (increase) in the likelihood of preferred (dispreferred) responses during the execution of these algorithms could be interpreted as a positive or negative phenomenon. In this paper, we study how to design algorithms that can leverage additional information in the form of rating gap, which informs the learner how…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsConstraint Satisfaction and Optimization · Machine Learning and Data Classification · Advanced Multi-Objective Optimization Algorithms