Direct preference alignment methods are favored for aligning LLMs with human preferences due to their computational and memory efficiency. To overcome likelihood displacement when preference pairs have small likelihood margins, the authors introduce Comparison-based Preference Optimization (ComPO), a zeroth-order alignment technique that relies on comparison oracles. ComPO extracts directional information from preference pairs to improve alignment without relying on likelihood estimates.
Read original
huggingface/daily-papers