Mixing on-policy and off-policy preference data in equal proportions (SIMPLEMIX) improves DPO alignment over either source alone, with on-policy data best for reasoning tasks and off-policy data best for open-ended tasks.
Cited on pages 5 and 16
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
Mixing on-policy and off-policy preference data in equal proportions (SIMPLEMIX) improves DPO alignment over either source alone, with on-policy data best for reasoning tasks and off-policy data best for open-ended tasks.