Reweighting offline DPO pairs by the likelihood ratio pi_theta/pi_ref reduces reward over-optimization and keeps the model closer to the reference policy.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling
Reweighting offline DPO pairs by the likelihood ratio pi_theta/pi_ref reduces reward over-optimization and keeps the model closer to the reference policy.