ReasonGRM uses a likelihood-based metric R* to select correct, high-confidence reasoning paths for supervised fine-tuning and then fine-tunes on hard cases with GRPO, reaching an average score of 83.3 across RewardBench, RM-Bench, and RMB.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models
ReasonGRM uses a likelihood-based metric R* to select correct, high-confidence reasoning paths for supervised fine-tuning and then fine-tunes on hard cases with GRPO, reaching an average score of 83.3 across RewardBench, RM-Bench, and RMB.