Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph

Ning Liu; Chuanneng Sun; Kristina Klinkner; Shervin Malmasi

arXiv:2605.08037·cs.LG·May 11, 2026

Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph

Ning Liu, Chuanneng Sun, Kristina Klinkner, Shervin Malmasi

PDF

TL;DR

This paper introduces GraphDPO, a novel preference optimization method that models complex preference structures as graphs, improving alignment of language models over traditional pairwise approaches.

Contribution

GraphDPO generalizes DPO by operating over preference graphs, capturing transitivity and rich preference relations, leading to more stable and effective model alignment.

Findings

01

GraphDPO outperforms pairwise DPO on reasoning and program synthesis tasks.

02

It efficiently encodes preference transitivity using graph structures.

03

The method maintains linear complexity despite leveraging full graph information.

Abstract

Direct Preference Optimization (DPO) aligns language models using pairwise preference comparisons, offering a simple and effective alternative to Reinforcement Learning (RL) from human feedback. However, in many practical settings, training data consists of multiple rollouts per prompt, inducing rich preference structure that pairwise DPO fails to exploit. Collapsing such data into independent pairs discards transitivity, introduces redundant or conflicting supervision, and can lead to unstable optimization. We propose Graph Direct Preference Optimization (GraphDPO), a principled generalization of DPO that operates over directed acyclic preference graphs induced by rollout rankings. GraphDPO encodes dominance relations as edges and optimizes a graph-structured Plackett--Luce-inspired objective that aggregates supervision over graph neighborhoods, enforcing transitivity while recovering…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.