Training reward models and DPO policies on response-conditioned preference pairs, where a length constraint is added to or withheld from the same prompt-response pair, reduces length bias and improves length instruction following.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling
Training reward models and DPO policies on response-conditioned preference pairs, where a length constraint is added to or withheld from the same prompt-response pair, reduces length bias and improves length instruction following.