Beyond Pairs: DPO Secretly Optimizes a Preference Graph, Not Just Pairwise Comparisons
arXiv·medium signal
Reveals that Direct Preference Optimization implicitly optimizes over a full preference graph rather than isolated pairs — reframing how alignment practitioners should think about preference data. The graph structure means DPO extracts more signal from existing datasets than previously understood, and the authors show how to exploit this for better data efficiency. Practical implications for anyone fine-tuning LLMs with RLHF/DPO.