Under unreliable demonstrations and comparisons, DPO fails to improve over SFT, while iterative label refinement (ILR) of the SFT dataset does improve performance across math, coding, and safe instruction-following tasks.
Input: How can I compute the area of a circle with radius 5? Response A: The area of it is 25π
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
Under unreliable demonstrations and comparisons, DPO fails to improve over SFT, while iterative label refinement (ILR) of the SFT dataset does improve performance across math, coding, and safe instruction-following tasks.