GFRIEND generates chain-of-thought preference judgments, scores them by perplexity, and uses weighted multi-level preference optimization so a reward model trained on 3,000 samples rivals models trained on much larger datasets.
This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO
GFRIEND generates chain-of-thought preference judgments, scores them by perplexity, and uses weighted multi-level preference optimization so a reward model trained on 3,000 samples rivals models trained on much larger datasets.