Ranking-based reward construction, using win counts against sampled responses or reference anchors, lets generative reward models improve LLM reinforcement learning more than probability-based rewards.
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
Ranking-based reward construction, using win counts against sampled responses or reference anchors, lets generative reward models improve LLM reinforcement learning more than probability-based rewards.