Pith. sign in

Evaluating robustness of reward models for mathematical reasoning

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

citation-role summary

dataset 1

citation-polarity summary

fields

cs.CL 1 cs.LG 1

years

2025 2

verdicts

UNVERDICTED 2

roles

dataset 1

polarities

background 1

representative citing papers

RewardBench 2: Advancing Reward Model Evaluation

cs.CL · 2025-06-02 · unverdicted · novelty 6.0

RewardBench 2 is a new benchmark that supplies challenging fresh human prompts for reward model evaluation, yielding lower average scores but higher correlation with downstream best-of-N sampling and RLHF training performance.

Reinforcement Learning from Human Feedback

cs.LG · 2025-04-16 · unverdicted · novelty 0.0

An expository book that systematically presents RLHF methods, from reward modeling to direct alignment algorithms, aimed at readers with quantitative backgrounds.

citing papers explorer

Showing 2 of 2 citing papers.

  • RewardBench 2: Advancing Reward Model Evaluation cs.CL · 2025-06-02 · unverdicted · none · ref 18

    RewardBench 2 is a new benchmark that supplies challenging fresh human prompts for reward model evaluation, yielding lower average scores but higher correlation with downstream best-of-N sampling and RLHF training performance.

  • Reinforcement Learning from Human Feedback cs.LG · 2025-04-16 · unverdicted · none · ref 95

    An expository book that systematically presents RLHF methods, from reward modeling to direct alignment algorithms, aimed at readers with quantitative backgrounds.