Under unreliable demonstrations and comparisons, DPO fails to improve over SFT, while iterative label refinement (ILR) of the SFT dataset does improve performance across math, coding, and safe instruction-following tasks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
Under unreliable demonstrations and comparisons, DPO fails to improve over SFT, while iterative label refinement (ILR) of the SFT dataset does improve performance across math, coding, and safe instruction-following tasks.