A DPO variant that kernelizes the preference loss and swaps KL for other divergences is claimed to improve alignment, but the math and evaluation do not support the state-of-the-art claim.
It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
A DPO variant that kernelizes the preference loss and swaps KL for other divergences is claimed to improve alignment, but the math and evaluation do not support the state-of-the-art claim.