RewardBench 2 is a new benchmark that supplies challenging fresh human prompts for reward model evaluation, yielding lower average scores but higher correlation with downstream best-of-N sampling and RLHF training performance.
Evaluating robustness of reward models for mathematical reasoning
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
dataset 1
citation-polarity summary
years
2025 2verdicts
UNVERDICTED 2roles
dataset 1polarities
background 1representative citing papers
An expository book that systematically presents RLHF methods, from reward modeling to direct alignment algorithms, aimed at readers with quantitative backgrounds.
citing papers explorer
-
RewardBench 2: Advancing Reward Model Evaluation
RewardBench 2 is a new benchmark that supplies challenging fresh human prompts for reward model evaluation, yielding lower average scores but higher correlation with downstream best-of-N sampling and RLHF training performance.
-
Reinforcement Learning from Human Feedback
An expository book that systematically presents RLHF methods, from reward modeling to direct alignment algorithms, aimed at readers with quantitative backgrounds.