Replacing DPO's relative reward margin with min(r_w, -α r_l) improves math benchmark accuracy by 5-10 points in this paper, though the exact gain figure is misreported and α is tuned on test data.
Bradley and Milton E
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
BPO: Revisiting Preference Modeling in Direct Preference Optimization
Replacing DPO's relative reward margin with min(r_w, -α r_l) improves math benchmark accuracy by 5-10 points in this paper, though the exact gain figure is misreported and α is tuned on test data.