AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation

arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC

Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational efficiency, and implicit modeling of human preferences. Interestingly, iterative extensions of DPO have achieved stronger performance on academic benchmarks, raising two key questions: (i) Why do iterative methods generally outperform offline ones? (ii) Can their advantages be incorporated into offline alignment? To answer the first question, our controlled experiments reveal that the explicit preference model, additionally introduced in the iterativ

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:52:10.320Z. This is not the publication date.