MPPO trains LLMs without a reference model by treating the geometric mean of response-token probabilities as the reward and jointly suppressing multiple negative responses; the Pair-MNM variant reports the best MT-Bench score in the paper.
gradient descent
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
MPPO trains LLMs without a reference model by treating the geometric mean of response-token probabilities as the reward and jointly suppressing multiple negative responses; the Pair-MNM variant reports the best MT-Bench score in the paper.