KPO extends the Plackett-Luce preference model used in DPO to top-K partial rankings, with query-adaptive K and curriculum learning, and reports improved LLM ranking accuracy.
Title resolution pending
1 Pith paper cite this work, alongside 6 external citations. Polarity classification is still indexing.
1
Pith paper citing it
6
external citations · OpenAlex
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
K-order Ranking Preference Optimization for Large Language Models
KPO extends the Plackett-Luce preference model used in DPO to top-K partial rankings, with query-adaptive K and curriculum learning, and reports improved LLM ranking accuracy.