Selective-DPO applies DPO only to the top 40% of tokens ranked by policy-reference log-probability difference, reporting improved benchmark scores but with tuned hyperparameters on test sets and omitted negative results.
the method of paired comparisons
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization
Selective-DPO applies DPO only to the top 40% of tokens ranked by policy-reference log-probability difference, reporting improved benchmark scores but with tuned hyperparameters on test sets and omitted negative results.