GFRIEND generates chain-of-thought preference judgments, scores them by perplexity, and uses weighted multi-level preference optimization so a reward model trained on 3,000 samples rivals models trained on much larger datasets.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO
GFRIEND generates chain-of-thought preference judgments, scores them by perplexity, and uses weighted multi-level preference optimization so a reward model trained on 3,000 samples rivals models trained on much larger datasets.