Preference Optimization, a DPO-style training loss that ranks sampled solutions by their objective value, speeds up and improves RL-based neural solvers for combinatorial problems.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Preference Optimization for Combinatorial Optimization Problems
Preference Optimization, a DPO-style training loss that ranks sampled solutions by their objective value, speeds up and improves RL-based neural solvers for combinatorial problems.