Reward models that think before judging, trained via reinforcement learning without human-written reasoning traces, outperform standard reward models and improve with more test-time compute.
Title resolution pending
1 Pith paper cite this work, alongside 7 external citations. Polarity classification is still indexing.
1
Pith paper citing it
7
external citations · OpenAlex
citation-role summary
other 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
Reward Reasoning Model
Reward models that think before judging, trained via reinforcement learning without human-written reasoning traces, outperform standard reward models and improve with more test-time compute.