Mixing on-policy and off-policy preference data in equal proportions (SIMPLEMIX) improves DPO alignment over either source alone, with on-policy data best for reasoning tasks and off-policy data best for open-ended tasks.
cc/paper_files/paper/2020/file/ 1f89885d556929e98d3ef9b86448f951-Paper
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
Mixing on-policy and off-policy preference data in equal proportions (SIMPLEMIX) improves DPO alignment over either source alone, with on-policy data best for reasoning tasks and off-policy data best for open-ended tasks.