Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Reward-model benchmarks that measure correctness do not predict how well the models guide LLM reasoning.

desk verdict Useful survey with a good point about RM evaluation; treat the headline experiment as suggestive, not conclusive. read the letter →

arxiv 2510.01925 v3 pith:U5QKEWGY submitted 2025-10-02 cs.CL

classification cs.CL
keywords rewardmodelsLLMreasoningprocessoutcomegenerativetest-timescalingmodelevaluationreinforcementlearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that the most common metrics used to evaluate reward models—how often they rank pairs correctly or spot an incorrect final answer—do not predict how useful they are in the two main jobs they are built for: choosing good solutions at test time and supplying training signals in reinforcement learning. The authors support this with experiments on six open-source process reward models, where a model with middling step-correctness scores ranked first in several downstream test-time tasks. They also find that generative reward models generally outperform discriminative ones, that process (step-level) models beat outcome (whole-solution) models for selecting answers but not for online RL, and that most reward models generalize poorly out of distribution. If the claim holds, practitioners should evaluate reward models by measuring their downstream effect—best-of-N accuracy, search-guiding quality, and the policy produced by RL—not by benchmark precision alone.

What carries the argument

Reward models are learned verifiers that map a question and a reasoning trace to a scalar score, and the survey's organizing taxonomy is the double distinction between discriminative vs generative RMs (scalar-only vs critique-producing) and outcome vs process RMs (whole-solution vs step-level). The analysis that carries the argument is a comparative experiment design: same-base-model comparisons of generative versus discriminative verifiers, PRM-versus-ORM comparisons on test-time selection and online RL, and a correlation analysis plotting step-correctness accuracy against best-of-N, beam-search, and MCTS outcomes for a set of process reward models under two different policy models.

What would settle it

Re-run the paper's correlation experiment with a much larger set of process reward models (several dozen), two or more judge prompts, and three or more policy models. If step-correctness accuracy and downstream best-of-N/MCTS/beam scores consistently track each other (Spearman above roughly 0.8), the claim that correctness metrics are insufficient would collapse; if the spread persists, it is confirmed.

Watch

Extended reading notes

Core claim

The paper's central discovery is that an RM's score on correctness-style benchmarks is a weak guide to its real-world value. In their experiment, step-level correctness accuracy (the metric used by popular process-reward benchmarks) shows only a modest positive correlation with best-of-N and search-guiding scores, and the relative ranking of reward models changes depending on which policy model generates the candidate solutions. A process reward model that ranks low on correctness can rank first on multiple downstream test-time tasks. From this the paper concludes that evaluation practices need to move toward directly measuring task-level performance, while retaining process-level metrics to

Load-bearing premise

The co-evolution claim rests on a measurement of discrimination that uses one judge prompt checking only final-answer correctness on 100 random questions per dataset; this assumes the prompt and sample represent a model's general judging ability and are not distorted by the judge's own training data or prompt sensitivity.

Editorial extensions

If this is right

  • RM evaluation should include downstream task metrics such as best-of-N accuracy and search-guiding score, not just pairwise or correctness accuracy.
  • Generative reward models, despite higher cost, are the safer choice when out-of-distribution generalization matters.
  • Process reward models are worth the extra step-level supervision for test-time selection, but should not be assumed to improve online RL over outcome rewards.
  • Improving the reasoning ability of the base model that serves as a generative RM should improve its judging accuracy, making reasoning training and reward-model training mutually reinforcing.
  • Most current RMs, especially discriminative ones, need task-specific retraining or domain adaptation when deployed outside their training distribution.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension: for online RL, a reward model's usefulness may hinge less on its ranking accuracy than on properties like reward variance and signal-to-noise ratio; the survey's cited evidence points this way but the authors do not make it their headline.
  • If correctness-style benchmarks continue to misalign with task performance, leaderboard rankings of RMs should probably be re-computed on downstream tasks across several policy models, and users should treat benchmark leader positions as weak evidence.
  • The co-evolution result suggests a concrete test: train the same base model alternately on generation and verification objectives and measure whether each stage raises the other; if it does, deliberate alternating training schedules could outperform separate reward-model and policy-model pipelines.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper is an analytical survey of reward models (RMs) for LLM reasoning. It develops taxonomies (discriminative vs. generative RMs, ORMs vs. PRMs, pointwise vs. pairwise), reviews evaluation benchmarks, and surveys three application areas: test-time guidance, synthetic data curation/self-improvement, and online RL. The paper also offers four analytical findings (Q1–Q4) on RM selection, OOD generalization, the co-evolution of generative and discriminative ability, and the adequacy of current RM evaluation metrics. The empirical contribution is small: Table V measures the correlation between generation and discrimination for ten LLMs, and Figure 5/Table IX compares ProcessBench correctness with downstream BoN/MCTS/Beam performance for six PRMs. The central empirical claim is that correctness-focused metrics, especially ProcessBench, may not predict real downstream performance, so practitioners should evaluate RMs with BoN-style metrics.

Significance. If the Q3 and Q4 findings hold, the survey provides actionable guidance for RM selection and evaluation, and it consolidates a rapidly growing literature in a useful way. The survey's taxonomies and coverage of recent methods/benchmarks are generally accurate and well organized. The paper's strengths include broad literature coverage, explicit comparisons of external results, and carefully hedged qualitative claims. However, the two new empirical analyses are small and lack statistical rigor, and the online-RL half of Q4 rests entirely on citations. The paper would be strengthened by making the new experiments reproducible and robust, or by explicitly scoping the 'we find' claims to the evidence provided. As a survey, the central claims remain defensible, but the novel empirical support needs attention.

major comments (3)
  1. [Section VI-D / Figure 5 / Table IX] The headline rank inversion — Skywork-PRM-7B 'ranks first in 4 out of 6 downstream tasks' despite a moderate ProcessBench score — is the key direct support for Q4, but Table IX reports no variance, no repeated seeds, and no significance tests. With six PRMs and margins as small as 1.0 pt (Qwen Beam: 82.2 vs 81.2; Mistral BoN: 49.6 vs 48.0), the inversion may be sampling noise. Please provide bootstrap confidence intervals, per-question error bars, or repeated-seed results, and state the exact number of MATH500 items used. If such robustness cannot be supplied, reframe this as a case study and lean on the cited external correlations.
  2. [Section VI-C / Table V] Q3's 'strong correlation' between generative and discriminative ability is asserted without a quantified correlation coefficient, confidence interval, or test. Discrimination is measured with a single fixed LLM-as-a-judge prompt that checks only final-answer correctness (Appendix C) on 100 random questions per dataset (Appendix B). This makes the trend vulnerable to prompt sensitivity and to contamination by the judge's own training data — an issue the paper itself raises for GPT-4o but does not resolve. Please report Spearman/pairwise correlations with uncertainty, vary judge prompts, and cross-check a subset with verifiable labels (e.g., exact-match). Also clarify the mixture of officially reported generation scores and 32-trial averages in the 'Avg.' column.
  3. [Section VI-D / Abstract / Q4] The claim that existing RM evaluation metrics are insufficient for online RL is supported only by external citations [235]–[237]; the authors' new experiments address test-time guidance (BoN/MCTS/Beam) exclusively. Since the abstract and introduction present Q4 as based partly on 'our empirical findings,' the paper should explicitly scope the claim: for test-time guidance the authors' own experiment is suggestive; for online RL the argument is a literature-based synthesis. This scoping is necessary to avoid overclaiming the paper's direct evidence.
minor comments (5)
  1. [Section II-C, Eq. (1)-(2)] Several formulas are incomplete or mis-rendered: `Rpoint_theta(p, tau) = r` and `Rpair_theta(P, tau1, tau2) = tau*` are stub equations, and Eq. (1) has an unmatched parenthesis in the expectation. Please fix the notation.
  2. [Appendix B] The paper does not state whether code, data, random seeds, or judge-prompt variants will be released. Given that Figure 5 and Table V are new empirical contributions, a reproducibility statement (even 'available upon request') is needed.
  3. [Tables III/IV] Values marked with * are described as read from published figures. This should be noted in the captions themselves, and the precision of such digitized values should be treated cautiously when drawing conclusions.
  4. [Appendix / Figure 6] Figure 6, the Spearman correlation heatmap among test-time strategies and ProcessBench, appears in the appendix but is not referenced in the main text. It would strengthen Section VI-D and should be cited there.
  5. [Section IV-A] Minor typo: 'REST-MCTS*' should be 'ReST-MCTS*' to match the referenced work. There are also occasional spacing/brace issues in the Appendix C prompts.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's analytical claims rest on external citations and fresh experiments; any self-citations are list-level and non-load-bearing.

full rationale

The paper's derivation chain is not circular. The central empirical claim (Section VI-D) is that correctness metrics such as ProcessBench are insufficient to predict downstream test-time performance; this is supported by a new experiment (Figure 5, Table IX) that compares ProcessBench-MATH500 accuracy against BoN@8, MCTS, and beam-search accuracies on MATH500 for six PRMs. No fitted parameter is later renamed as a prediction: the linear-regression trend lines in Figure 5 are descriptive summaries, and the paper uses residual cases (e.g., Skywork-PRM-7B) to argue that correctness scores are insufficient, which is an ordinary empirical comparison rather than a constructional equivalence. The co-evolution claim in Section VI-C similarly rests on separately measured generation scores and LLM-as-a-judge discrimination accuracy (Table V, Appendix B); the correlation is empirical, not definitional. The only author-overlapping self-citations, notably SelfCheck [165] in the sampling/selection list and possibly RM-Bench [80] in Table I, appear in method enumerations and are not load-bearing: the main arguments about RM selection, generalization, evaluation, and online RL are supported by external references (e.g., [79], [81], [89], [34], [235], [236], [237]) and by the paper's own new experiments. Concerns such as the small number of PRMs in Table IX, the single MATH500 dataset, the absence of variance estimates, and the reliance on one LLM-as-a-judge prompt in Table V are threats to statistical robustness or measurement validity, but they are not circularity and do not indicate that the conclusions reduce to their inputs.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

No mathematical free parameters or invented entities. The empirical claims rest on several chosen experimental settings and representativeness assumptions, which are the main fragility points.

free parameters (4)
  • beam_size = 4
    Section VI-D/Appendix B: used for beam search; no sensitivity analysis provided; could affect PRM downstream rankings.
  • mcts_simulations = 4
    Section VI-D/Appendix B: MCTS used 4 simulation paths; no sensitivity analysis provided.
  • generation_temperature = 0.7
    Appendix B: set for all evaluations; may affect BoN and search results.
  • num_questions_per_dataset = 100
    Appendix B: 100 random questions per dataset for discrimination evaluation; small sample, no confidence intervals.
assumptions (3)
  • domain assumption The cited literature accurately reports the results attributed to it.
    A survey's comparative claims (e.g., GRMs > DRMs, PRMs > ORMs at test time) rest on the reliability of the cited studies.
  • domain assumption The six PRMs and two policy models in Figure 5 are representative of the broader PRM landscape.
    Section VI-D: conclusions about ProcessBench vs downstream performance are based on this small set.
  • domain assumption The LLM-as-a-judge prompt in Appendix C measures discriminative ability rather than prompt-following or answer-format matching.
    Section VI-C: Table V correlations would change if the judge prompt rewarded step-level reasoning or penalized style artifacts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey." pith.science (2026). https://pith.science/paper/U5QKEWGY

@misc{pith2026251001925,
  author       = {Pith},
  title        = {Pith review of: Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U5QKEWGY}},
  note         = {Machine review of arXiv:2510.01925}
}
read the original abstract

Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune LLMs during reinforcement learning (RL) and help select the best answer from multiple candidates during inference. In this paper, we provide a systematic introduction to RMs, along with a comprehensive survey of their applications in LLM reasoning. We first review fundamental concepts of RMs, including their architectures, training methodologies, and evaluation techniques. Then, we explore their key applications: (1) guiding generation and selecting optimal outputs during LLM inference, (2) facilitating data synthesis and iterative self-improvement for LLMs, and (3) providing training signals in RL-based finetuning. Finally, we discuss critical open questions regarding the selection, generalization, evaluation, and enhancement of RMs, based on existing research and our own empirical findings. Our analysis aims to provide actionable insights for the effective deployment and advancement of RMs for LLM reasoning.

Figures

Figures reproduced from arXiv: 2510.01925 by the authors.

Figure 1
Figure 1. Illustration of three main applications of reward models in LLM reasoning. Green/red blocks denote higher/lower-quality [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of current research on process reward models [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Applications of RMs in LLM reasoning frequently (i.e., a majority vote over final answers) without an explicit verifier. In contrast, the generator-verifier paradigm equips selection with reward scores from PRMs or ORMs to explicitly verify the correctness of each solution. Whereas self-consistency may fail when the policy model has a higher probability of generating incorrect answers, selection methods with reward … view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Comparisons of Llama and Qwen response styles in an example math question [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: The relationship between correctness scores (ProcessBench), BoN scores, and search-guiding performance (MCTS [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: Heatmap illustrating Spearman correlation coefficients among the evaluated test-time search strategies and ProcessBench [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: The discrimination accuracy for responses generated from different models [PITH_FULL_IMAGE:figures/full_fig_p023_7.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    This survey introduces the Generate-Filter-Control-Replay (GFCR) taxonomy to structure rollout pipelines for RL-based post-training of reasoning LLMs.

  2. Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    PASS middleware independently standardizes process/outcome/format streams, derives value-homogeneous chunks, and converts cumulative returns to average value density, yielding consistent pass@1 gains over GRPO baselin...

Reference graph

Works this paper leans on

237 extracted references · cited by 2 Pith papers

  1. [235]

    The accuracy paradox in rlhf: When better reward models don’t yield better language models,

    Y . Chen, D. Zhu, Y . Sun, X. Chen, W. Zhang, and X. Shen, “The accuracy paradox in rlhf: When better reward models don’t yield better language models,” 2024

  2. [237]

    What makes a reward model a good teacher? an optimization perspective,

    N. Razin, Z. Wang, H. Strauss, S. Wei, J. D. Lee, and S. Arora, “What makes a reward model a good teacher? an optimization perspective,” 2025

  3. [1]

    Sparks of artificial general intelligence: Early experiments with gpt-4,

    S. Bubeck, V . Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y . T. Lee, Y . Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y . Zhang, “Sparks of artificial general intelligence: Early experiments with gpt-4,” 2023

  4. [2]

    A survey on medical large language models: Technology, application, trustworthiness, and future directions,

    L. Liu, X. Yang, J. Lei, Y . Shen, J. Wang, P. Wei, Z. Chu, Z. Qin, and K. Ren, “A survey on medical large language models: Technology, application, trustworthiness, and future directions,” 2024

  5. [3]

    Vision-language models for vision tasks: A survey,

    J. Zhang, J. Huang, S. Jin, and S. Lu, “Vision-language models for vision tasks: A survey,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5625–5644, 2024

  6. [4]

    Bridging the linguistic divide: A survey on leveraging large language models for machine translation,

    B. Gain, D. Bandyopadhyay, and A. Ekbal, “Bridging the linguistic divide: A survey on leveraging large language models for machine translation,” 2025

  7. [5]

    A survey of large language model agents for question answering,

    M. Yue, “A survey of large language model agents for question answering,” 2025

  8. [6]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” 2023. 16

Show all 237 references
  1. [7]

    Tree of thoughts: Deliberate problem solving with large language models,

    S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y . Cao, and K. R. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” inThirty-seventh Conference on Neural Information Processing Systems, 2023

  2. [8]

    Solving quantitative reasoning problems with language models,

    A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V . Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y . Wu, B. Neyshabur, G. Gur-Ari, and V . Misra, “Solving quantitative reasoning problems with language models,” 2022

  3. [9]

    Mint: Boosting generalization in mathematical reasoning via multi-view fine- tuning,

    Z. Liang, D. Yu, X. Pan, W. Yao, Q. Zeng, X. Zhang, and D. Yu, “Mint: Boosting generalization in mathematical reasoning via multi-view fine- tuning,” 2023

  4. [10]

    Robust visual question answering: Datasets, methods, and future challenges,

    J. Ma, P. Wang, D. Kong, Z. Wang, J. Liu, H. Pei, and J. Zhao, “Robust visual question answering: Datasets, methods, and future challenges,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5575–5594, 2024

  5. [11]

    Learning from mistakes makes llm better reasoner,

    S. An, Z. Ma, Z. Lin, N. Zheng, J.-G. Lou, and W. Chen, “Learning from mistakes makes llm better reasoner,” 2024

  6. [12]

    Openai o1 system card,

    OpenAIet al., “Openai o1 system card,” 2024

  7. [13]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

    DeepSeek-AIet al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” 2025

  8. [14]

    Tulu 3: Pushing frontiers in open language model post-training,

    N. Lambert, J. Morrison, V . Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V . Miranda, A. Liu, N. Dziri, S. Lyu, Y . Gu, S. Malik, V . Graf, J. D. Hwang, J. Yang, R. L. Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y . Wang, P. Dasigi, and H. Hajishirzi, “Tulu 3: ...

  9. [15]

    Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback,

    H. Lee, S. Phatale, H. Mansoor, T. Mesnard, J. Ferret, K. Lu, C. Bishop, E. Hall, V . Carbune, A. Rastogi, and S. Prakash, “Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback,” 2024

  10. [16]

    Improve mathematical reasoning in language models by automated process supervision,

    L. Luo, Y . Liu, R. Liu, S. Phatale, M. Guo, H. Lara, Y . Li, L. Shu, Y . Zhu, L. Meng, J. Sun, and A. Rastogi, “Improve mathematical reasoning in language models by automated process supervision,” 2024

  11. [17]

    Advancing process verification for large language models via tree-based preference learning,

    M. He, Y . Shen, W. Zhang, Z. Tan, and W. Lu, “Advancing process verification for large language models via tree-based preference learning,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Ed...

  12. [18]

    Token-supervised value models for enhancing mathematical problem- solving capabilities of large language models,

    J. H. Lee, J. Y . Yang, B. Heo, D. Han, K. Kim, E. Yang, and K. M. Yoo, “Token-supervised value models for enhancing mathematical problem- solving capabilities of large language models,” 2025

  13. [19]

    Coarse-to-fine process reward modeling for mathematical reasoning,

    Y . Hu, G. Chen, J. Zhao, S. Ouyang, and Y . Liu, “Coarse-to-fine process reward modeling for mathematical reasoning,” 2025

  14. [20]

    Visualprm: An effective process reward model for multimodal reasoning,

    W. Wang, Z. Gao, L. Chen, Z. Chen, J. Zhu, X. Zhao, Y . Liu, Y . Cao, S. Ye, X. Zhu, L. Lu, H. Duan, Y . Qiao, J. Dai, and W. Wang, “Visualprm: An effective process reward model for multimodal reasoning,” 2025

  15. [21]

    Towards hierarchical multi-step reward models for enhanced reasoning in large language models,

    T. Wang, Z. Jiang, Z. He, W. Yang, Y . Zheng, Z. Li, Z. He, S. Tong, and H. Gong, “Towards hierarchical multi-step reward models for enhanced reasoning in large language models,” 2025

  16. [22]

    Adaptivestep: Automatically dividing reasoning step through model confidence,

    Y . Liu, J. Lu, Z. Chen, C. Qu, J. K. Liu, C. Liu, Z. Cai, Y . Xia, L. Zhao, J. Bian, C. Zhang, W. Shen, and Z. Lin, “Adaptivestep: Automatically dividing reasoning step through model confidence,” 2025

  17. [23]

    Vilbench: A suite for vision-language process reward modeling,

    H. Tu, W. Feng, H. Chen, H. Liu, X. Tang, and C. Xie, “Vilbench: A suite for vision-language process reward modeling,” 2025

  18. [24]

    Retrieval-augmented process reward model for generalizable mathematical reasoning,

    J. Zhu, C. Zheng, J. Lin, K. Du, Y . Wen, Y . Yu, J. Wang, and W. Zhang, “Retrieval-augmented process reward model for generalizable mathematical reasoning,” 2025

  19. [25]

    Making large language models better reasoners with step-aware verifier,

    Y . Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen, “Making large language models better reasoners with step-aware verifier,” 2023

  20. [26]

    OVM, outcome-supervised value models for planning in mathematical reasoning,

    F. Yu, A. Gao, and B. Wang, “OVM, outcome-supervised value models for planning in mathematical reasoning,” inFindings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico: Association for Computational Linguis...

  21. [27]

    Let’s verify step by step,

    H. Lightman, V . Kosaraju, Y . Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” 2023

  22. [28]

    Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations,

    P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y . Li, D. Chen, Y . Wu, and Z. Sui, “Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L....

  23. [29]

    Multi-step problem solving through a verifier: An empirical analysis on model- induced process supervision,

    Z. Wang, Y . Li, Y . Wu, L. Luo, L. Hou, H. Yu, and J. Shang, “Multi-step problem solving through a verifier: An empirical analysis on model- induced process supervision,” 2024

  24. [30]

    Glore: When, where, and how to improve llm reasoning via global and local refinements,

    A. Havrilla, S. Raparthy, C. Nalmpantis, J. Dwivedi-Yu, M. Zhuravin- skyi, E. Hambro, and R. Raileanu, “Glore: When, where, and how to improve llm reasoning via global and local refinements,” 2024

  25. [31]

    Autopsv: Automated process-supervised verifier,

    J. Lu, Z. Dou, H. Wang, Z. Cao, J. Dai, Y . Wan, Y . Feng, and Z. Guo, “Autopsv: Automated process-supervised verifier,” inAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curra...

  26. [32]

    Rewarding progress: Scaling automated process verifiers for llm reasoning,

    A. Setlur, C. Nagpal, A. Fisch, X. Geng, J. Eisenstein, R. Agarwal, A. Agarwal, J. Berant, and A. Kumar, “Rewarding progress: Scaling automated process verifiers for llm reasoning,” 2024

  27. [33]

    Entropy-regularized process reward model,

    H. Zhang, P. Wang, S. Diao, Y . Lin, R. Pan, H. Dong, D. Zhang, P. Molchanov, and T. Zhang, “Entropy-regularized process reward model,” 2024

  28. [34]

    The lessons of developing process reward models in mathematical reasoning,

    Z. Zhang, C. Zheng, Y . Wu, B. Zhang, R. Lin, B. Yu, D. Liu, J. Zhou, and J. Lin, “The lessons of developing process reward models in mathematical reasoning,” 2025

  29. [35]

    Athena: Enhancing multimodal reasoning with data-efficient process reward models,

    S. Wang, Z. Liu, J. Wei, X. Yin, D. Li, and E. Barsoum, “Athena: Enhancing multimodal reasoning with data-efficient process reward models,” 2025

  30. [36]

    Reasonflux- prm: Trajectory-aware prms for long chain-of-thought reasoning in llms,

    J. Zou, L. Yang, J. Gu, J. Qiu, K. Shen, J. He, and M. Wang, “Reasonflux- prm: Trajectory-aware prms for long chain-of-thought reasoning in llms,” 2025

  31. [37]

    Better process supervision with bi-directional rewarding signals,

    W. Chen, W. He, Z. Xi, H. Guo, B. Hong, J. Zhang, R. Zheng, N. Li, T. Gui, Y . Li, Q. Zhang, and X. Huang, “Better process supervision with bi-directional rewarding signals,” 2025

  32. [38]

    Duashepherd: Integrating stepwise correctness and potential rewards for mathematical reasoning,

    Y . Wu, J. Song, H. Zhang, T. Zhang, and C. Niu, “Duashepherd: Integrating stepwise correctness and potential rewards for mathematical reasoning,” 2025

  33. [39]

    Llm critics help catch bugs in mathematics: Towards a better mathematical verifier with natural language feedback,

    B. Gao, Z. Cai, R. Xu, P. Wang, C. Zheng, R. Lin, K. Lu, D. Liu, C. Zhou, W. Xiao, J. Hu, T. Liu, and B. Chang, “Llm critics help catch bugs in mathematics: Towards a better mathematical verifier with natural language feedback,” 2024

  34. [40]

    Verifierq: Enhancing llm test time compute with q-learning-based verifiers,

    J. Qi, H. Tang, and Z. Zhu, “Verifierq: Enhancing llm test time compute with q-learning-based verifiers,” 2024

  35. [41]

    Process reward model with q-value rankings,

    W. Li and Y . Li, “Process reward model with q-value rankings,” 2025

  36. [42]

    Free process rewards without process labels,

    L. Yuan, W. Li, H. Chen, G. Cui, N. Ding, K. Zhang, B. Zhou, Z. Liu, and H. Peng, “Free process rewards without process labels,” 2024

  37. [43]

    Tdrm: Smooth reward models with temporal difference for llm rl and inference,

    D. Zhang, M. Cai, J. Li, Z. Hu, Y . Yue, Y . Dong, and J. Tang, “Tdrm: Smooth reward models with temporal difference for llm rl and inference,” 2025

  38. [44]

    Cold: Counterfactually-guided length debiasing for process reward models,

    C. Zheng, J. Zhu, J. Lin, X. Dai, Y . Yu, W. Zhang, and M. Yang, “Cold: Counterfactually-guided length debiasing for process reward models,” 2025

  39. [45]

    Judging LLM-as-a-judge with MT-bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging LLM-as-a-judge with MT-bench and chatbot arena,” inThirty-seventh Conference on Neural Information Processing Systems Datasets and ...

  40. [46]

    R-prm: Reasoning-driven process reward modeling,

    S. She, J. Liu, Y . Liu, J. Chen, X. Huang, and S. Huang, “R-prm: Reasoning-driven process reward modeling,” 2025

  41. [47]

    Genprm: Scaling test-time compute of process reward models via generative reasoning,

    J. Zhao, R. Liu, K. Zhang, Z. Zhou, J. Gao, D. Li, J. Lyu, Z. Qian, B. Qi, X. Li, and B. Zhou, “Genprm: Scaling test-time compute of process reward models via generative reasoning,” 2025

  42. [48]

    Scaling evaluation-time compute with reasoning models as process evaluators,

    S. Kim, I. Wu, J. Lee, X. Yue, S. Lee, M. Moon, K. Gashteovski, C. Lawrence, J. Hockenmaier, G. Neubig, and S. Welleck, “Scaling evaluation-time compute with reasoning models as process evaluators,” 2025

  43. [49]

    Spc: Evolving self-play critic via adversarial games for llm reasoning,

    J. Chen, B. Zhang, R. Ma, P. Wang, X. Liang, Z. Tu, X. Li, and K.-Y . K. Wong, “Spc: Evolving self-play critic via adversarial games for llm reasoning,” 2025

  44. [50]

    Process reward models that think,

    M. Khalifa, R. Agarwal, L. Logeswaran, J. Kim, H. Peng, M. Lee, H. Lee, and L. Wang, “Process reward models that think,” 2025

  45. [51]

    Stepwiser: Stepwise generative judges for wiser reasoning,

    W. Xiong, W. Zhao, W. Yuan, O. Golovneva, T. Zhang, J. Weston, and S. Sukhbaatar, “Stepwiser: Stepwise generative judges for wiser reasoning,” 2025

  46. [52]

    Solving math word problems with process- and outcome-based feedback,

    J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins, “Solving math word problems with process- and outcome-based feedback,” 2022

  47. [53]

    Training verifiers to solve math word problems,

    K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, “Training verifiers to solve math word problems,” 2021

  48. [54]

    Direct preference optimization: Your language model is secretly a reward model,

    R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” 2024

  49. [55]

    Inference-time scaling for generalist reward modeling,

    Z. Liu, P. Wang, R. Xu, S. Ma, C. Ruan, P. Li, Y . Liu, and Y . Wu, “Inference-time scaling for generalist reward modeling,” 2025. 17

  50. [56]

    Rm-r1: Reward modeling as reasoning,

    X. Chen, G. Li, Z. Wang, B. Jin, C. Qian, Y . Wang, H. Wang, Y . Zhang, D. Zhang, T. Zhang, H. Tong, and H. Ji, “Rm-r1: Reward modeling as reasoning,” 2025

  51. [57]

    Ticking all the boxes: Generated checklists improve llm evaluation and generation,

    J. Cook, T. Rocktäschel, J. Foerster, D. Aumiller, and A. Wang, “Ticking all the boxes: Generated checklists improve llm evaluation and generation,” 2024

  52. [58]

    Generative verifiers: Reward modeling as next-token prediction,

    L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal, “Generative verifiers: Reward modeling as next-token prediction,” 2025

  53. [59]

    Critique- out-loud reward models,

    Z. Ankner, M. Paul, B. Cui, J. D. Chang, and P. Ammanabrolu, “Critique- out-loud reward models,” 2024

  54. [60]

    Learning to reason for factuality,

    X. Chen, I. Kulikov, V .-P. Berges, B. O ˘guz, R. Shao, G. Ghosh, J. Weston, and W. tau Yih, “Learning to reason for factuality,” 2025

  55. [61]

    Internlm2 technical report,

    Z. Caiet al., “Internlm2 technical report,” 2024

  56. [62]

    Advancing llm reasoning generalists with preference trees,

    L. Yuan, G. Cui, H. Wang, N. Ding, X. Wang, J. Deng, B. Shan, H. Chen, R. Xie, Y . Lin, Z. Liu, B. Zhou, H. Peng, Z. Liu, and M. Sun, “Advancing llm reasoning generalists with preference trees,” 2024

  57. [63]

    Interpretable preferences via multi-objective reward modeling and mixture-of-experts,

    H. Wang, W. Xiong, T. Xie, H. Zhao, and T. Zhang, “Interpretable preferences via multi-objective reward modeling and mixture-of-experts,” 2024

  58. [64]

    Llm-blender: Ensembling large language models with pairwise ranking and generative fusion,

    D. Jiang, X. Ren, and B. Y . Lin, “Llm-blender: Ensembling large language models with pairwise ranking and generative fusion,” 2023

  59. [65]

    Helpsteer2-preference: Complementing ratings with preferences,

    Z. Wang, A. Bukharin, O. Delalleau, D. Egert, G. Shen, J. Zeng, O. Kuchaiev, and Y . Dong, “Helpsteer2-preference: Complementing ratings with preferences,” 2025

  60. [66]

    Kto: Model alignment as prospect theoretic optimization,

    K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela, “Kto: Model alignment as prospect theoretic optimization,” 2024

  61. [67]

    Bootstrapping language models with dpo implicit rewards,

    C. Chen, Z. Liu, C. Du, T. Pang, Q. Liu, A. Sinha, P. Varakantham, and M. Lin, “Bootstrapping language models with dpo implicit rewards,” 2025

  62. [68]

    Generative judge for evaluating alignment,

    J. Li, S. Sun, W. Yuan, R.-Z. Fan, hai zhao, and P. Liu, “Generative judge for evaluating alignment,” inThe Twelfth International Conference on Learning Representations, 2024

  63. [69]

    Prometheus 2: An open source language model specialized in evaluating other language models,

    S. Kim, J. Suk, S. Longpre, B. Y . Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo, “Prometheus 2: An open source language model specialized in evaluating other language models,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Proc...

  64. [70]

    Foundational autoraters: Taming large language models for better automatic evaluation,

    T. Vu, K. Krishna, S. Alzubi, C. Tar, M. Faruqui, and Y .-H. Sung, “Foundational autoraters: Taming large language models for better automatic evaluation,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and ...

  65. [71]

    Compassjudger-1: All-in-one judge model helps model evaluation and evolution,

    M. Cao, A. Lam, H. Duan, H. Liu, S. Zhang, and K. Chen, “Compassjudger-1: All-in-one judge model helps model evaluation and evolution,” 2024

  66. [72]

    Learning LLM-as-a-judge for preference alignment,

    Z. Ye, X. Li, Q. Li, Q. Ai, Y . Zhou, W. Shen, D. Yan, and Y . LIU, “Learning LLM-as-a-judge for preference alignment,” inThe Thirteenth International Conference on Learning Representations, 2025

  67. [73]

    Atla selene mini: A general purpose evaluation model,

    A. Alexandru, A. Calvi, H. Broomfield, J. Golden, K. Dai, M. Leys, M. Burger, M. Bartolo, R. Engeler, S. Pisupati, T. Drane, and Y . S. Park, “Atla selene mini: A general purpose evaluation model,” 2025

  68. [74]

    One token to fool llm-as-a-judge,

    Y . Zhao, H. Liu, D. Yu, S. Y . Kung, H. Mi, and D. Yu, “One token to fool llm-as-a-judge,” 2025

  69. [75]

    Judgelrm: Large reasoning models as a judge,

    N. Chen, Z. Hu, Q. Zou, J. Wu, Q. Wang, B. Hooi, and B. He, “Judgelrm: Large reasoning models as a judge,” 2025

  70. [76]

    Unified multimodal chain-of-thought reward model through reinforcement fine- tuning,

    Y . Wang, Z. Li, Y . Zang, C. Wang, Q. Lu, C. Jin, and J. Wang, “Unified multimodal chain-of-thought reward model through reinforcement fine- tuning,” 2025

  71. [78]

    Pairjudge rm: Perform best-of-n sampling with knockout tournament,

    Y . Liu, Z. Yao, R. Min, Y . Cao, L. Hou, and J. Li, “Pairjudge rm: Perform best-of-n sampling with knockout tournament,” 2025

  72. [79]

    Rewardbench: Evaluating reward models for language modeling,

    N. Lambert, V . Pyatkin, J. Morrison, L. Miranda, B. Y . Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y . Choi, N. A. Smith, and H. Hajishirzi, “Rewardbench: Evaluating reward models for language modeling,” 2024

  73. [80]

    Rm-bench: Benchmarking reward models of language models with subtlety and style,

    Y . Liu, Z. Yao, R. Min, Y . Cao, L. Hou, and J. Li, “Rm-bench: Benchmarking reward models of language models with subtlety and style,” 2024

  74. [81]

    Rmb: Comprehensively benchmarking reward models in llm alignment,

    E. Zhou, G. Zheng, B. Wang, Z. Xi, S. Dou, R. Bao, W. Shen, L. Xiong, J. Fan, Y . Mou, R. Zheng, T. Gui, Q. Zhang, and X. Huang, “Rmb: Comprehensively benchmarking reward models in llm alignment,” 2025

  75. [82]

    How to evaluate reward models for rlhf,

    E. Frick, T. Li, C. Chen, W.-L. Chiang, A. N. Angelopoulos, J. Jiao, B. Zhu, J. E. Gonzalez, and I. Stoica, “How to evaluate reward models for rlhf,” 2024

  76. [83]

    Rag- rewardbench: Benchmarking reward models in retrieval augmented generation for preference alignment,

    Z. Jin, H. Yuan, T. Men, P. Cao, Y . Chen, K. Liu, and J. Zhao, “Rag- rewardbench: Benchmarking reward models in retrieval augmented generation for preference alignment,” 2024

  77. [84]

    Acemath: Advancing frontier math reasoning with post-training and reward modeling,

    Z. Liu, Y . Chen, M. Shoeybi, B. Catanzaro, and W. Ping, “Acemath: Advancing frontier math reasoning with post-training and reward modeling,” 2025

  78. [85]

    M- rewardbench: Evaluating reward models in multilingual settings,

    S. Gureja, L. J. V . Miranda, S. B. Islam, R. Maheshwary, D. Sharma, G. Winata, N. Lambert, S. Ruder, S. Hooker, and M. Fadaee, “M- rewardbench: Evaluating reward models in multilingual settings,” 2024

  79. [86]

    Rewardbench 2: Advancing reward model evaluation,

    S. Malik, V . Pyatkin, S. Land, J. Morrison, N. A. Smith, H. Hajishirzi, and N. Lambert, “Rewardbench 2: Advancing reward model evaluation,” 2025

  80. [87]

    Rewardanything: Generalizable principle- following reward models,

    Z. Yu, J. Zeng, W. Gu, Y . Wang, J. Wang, F. Meng, J. Zhou, Y . Zhang, S. Zhang, and W. Ye, “Rewardanything: Generalizable principle- following reward models,” 2025

  81. [88]

    Posterior-grpo: Rewarding reasoning processes in code generation,

    L. Fan, Y . Zhang, M. Chen, and Z. Liu, “Posterior-grpo: Rewarding reasoning processes in code generation,” 2025

  82. [89]

    Processbench: Identifying process errors in mathematical reasoning,

    C. Zheng, Z. Zhang, B. Zhang, R. Lin, K. Lu, B. Yu, D. Liu, J. Zhou, and J. Lin, “Processbench: Identifying process errors in mathematical reasoning,” 2024

  83. [90]

    Aurora:automated training framework of universal process reward models via ensemble prompting and reverse verification,

    X. Tan, T. Yao, C. Qu, B. Li, M. Yang, D. Lu, H. Wang, X. Qiu, W. Chu, Y . Xu, and Y . Qi, “Aurora:automated training framework of universal process reward models via ensemble prompting and reverse verification,” 2025

  84. [91]

    Prmbench: A fine- grained and challenging benchmark for process-level reward models,

    M. Song, Z. Su, X. Qu, J. Zhou, and Y . Cheng, “Prmbench: A fine- grained and challenging benchmark for process-level reward models,” 2025

  85. [92]

    Evaluating judges as evaluators: The jetts benchmark of llm-as-judges as test-time scaling evaluators,

    Y . Zhou, A. Xu, P. Wang, C. Xiong, and S. Joty, “Evaluating judges as evaluators: The jetts benchmark of llm-as-judges as test-time scaling evaluators,” 2025

  86. [93]

    Mr-gsm8k: A meta- reasoning benchmark for large language model evaluation,

    Z. Zeng, P. Chen, S. Liu, H. Jiang, and J. Jia, “Mr-gsm8k: A meta- reasoning benchmark for large language model evaluation,” 2024

  87. [94]

    Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms,

    Z. Zeng, Y . Liu, Y . Wan, J. Li, P. Chen, J. Dai, Y . Yao, R. Xu, Z. Qi, W. Zhao, L. Shen, J. Lu, H. Tan, Y . Chen, H. Zhang, Z. Shi, B. Wang, Z. Guo, and J. Jia, “Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms,” 2024

  88. [95]

    Vlrewardbench: A challenging benchmark for vision-language generative reward models,

    L. Li, Y . Wei, Z. Xie, X. Yang, Y . Song, P. Wang, C. An, T. Liu, S. Li, B. Y . Lin, L. Kong, and Q. Liu, “Vlrewardbench: A challenging benchmark for vision-language generative reward models,” 2024

  89. [96]

    Mj-bench: Is your multimodal reward model really a good judge for text-to-image generation?

    Z. Chen, Y . Du, Z. Wen, Y . Zhou, C. Cui, Z. Weng, H. Tu, C. Wang, Z. Tong, Q. Huang, C. Chen, Q. Ye, Z. Zhu, Y . Zhang, J. Zhou, Z. Zhao, R. Rafailov, C. Finn, and H. Yao, “Mj-bench: Is your multimodal reward model really a good judge for text-to-image generation?” 2024

  90. [97]

    Multimodal rewardbench: Holistic evaluation of reward models for vision language models,

    M. Yasunaga, L. Zettlemoyer, and M. Ghazvininejad, “Multimodal rewardbench: Holistic evaluation of reward models for vision language models,” 2025

  91. [98]

    Vlrmbench: A comprehensive and challenging benchmark for vision-language reward models,

    J. Ruan, W. Yuan, X. Gao, Y . Guo, D. Zhang, Z. Xu, Y . Hu, T. Liu, and Y . Fu, “Vlrmbench: A comprehensive and challenging benchmark for vision-language reward models,” 2025

  92. [99]

    Large language monkeys: Scaling inference compute with repeated sampling,

    B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V . Le, C. Ré, and A. Mirhoseini, “Large language monkeys: Scaling inference compute with repeated sampling,” 2024

  93. [100]

    Scaling llm test-time compute optimally can be more effective than scaling model parameters,

    C. Snell, J. Lee, K. Xu, and A. Kumar, “Scaling llm test-time compute optimally can be more effective than scaling model parameters,” 2024

  94. [101]

    Sample, don’t search: Rethinking test-time alignment for language models,

    G. Faria and N. A. Smith, “Sample, don’t search: Rethinking test-time alignment for language models,” 2025

  95. [102]

    Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment,

    A. Huang, A. Block, Q. Liu, N. Jiang, A. Krishnamurthy, and D. J. Foster, “Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment,” 2025

  96. [103]

    When to solve, when to verify: Compute-optimal problem solving and generative verification for llm reasoning,

    N. Singhi, H. Bansal, A. Hosseini, A. Grover, K.-W. Chang, M. Rohrbach, and A. Rohrbach, “When to solve, when to verify: Compute-optimal problem solving and generative verification for llm reasoning,” 2025

  97. [104]

    Reasoning with language model is planning with world model,

    S. Hao, Y . Gu, H. Ma, J. Hong, Z. Wang, D. Wang, and Z. Hu, “Reasoning with language model is planning with world model,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association f...

  98. [105]

    Grace: Discriminator-guided chain-of-thought reasoning,

    M. Khalifa, L. Logeswaran, M. Lee, H. Lee, and L. Wang, “Grace: Discriminator-guided chain-of-thought reasoning,” 2023

  99. [106]

    Let’s reward step by step: Step-level reward model as the navigators for reasoning,

    Q. Ma, H. Zhou, T. Liu, J. Yuan, P. Liu, Y . You, and H. Yang, “Let’s reward step by step: Step-level reward model as the navigators for reasoning,” 2023. 18

  100. [107]

    Deductive beam search: Decoding deducible rationale for chain-of-thought reasoning,

    T. Zhu, K. Zhang, J. Xie, and Y . Su, “Deductive beam search: Decoding deducible rationale for chain-of-thought reasoning,” 2024

  101. [108]

    Mindstar: Enhancing math reasoning in pre-trained llms at inference time,

    J. Kang, X. Z. Li, X. Chen, A. Kazemi, Q. Sun, B. Chen, D. Li, X. He, Q. He, F. Wen, J. Hao, and J. Yao, “Mindstar: Enhancing math reasoning in pre-trained llms at inference time,” 2024

  102. [109]

    Q*: Improving multi-step reasoning for llms with deliberative planning,

    C. Wang, Y . Deng, Z. Lyu, L. Zeng, J. He, S. Yan, and B. An, “Q*: Improving multi-step reasoning for llms with deliberative planning,” 2024

  103. [110]

    Ensembling large language models with process reward-guided tree search for better complex reasoning,

    S. Park, X. Liu, Y . Gong, and E. Choi, “Ensembling large language models with process reward-guided tree search for better complex reasoning,” 2024

  104. [111]

    Llm2: Let large language models harness system 2 reasoning,

    C. Yang, C. Shi, S. Li, B. Shui, Y . Yang, and W. Lam, “Llm2: Let large language models harness system 2 reasoning,” 2025

  105. [112]

    Agentrm: Enhancing agent generalization with reward modeling,

    Y . Xia, J. Fan, W. Chen, S. Yan, X. Cong, Z. Zhang, Y . Lu, Y . Lin, Z. Liu, and M. Sun, “Agentrm: Enhancing agent generalization with reward modeling,” 2025

  106. [113]

    Mt-rewardtree: A comprehensive framework for advancing llm-based machine translation via reward modeling,

    Z. Feng, J. Ren, J. Su, J. Zheng, Z. Tang, H. Wang, and Z. Liu, “Mt-rewardtree: A comprehensive framework for advancing llm-based machine translation via reward modeling,” 2025

  107. [114]

    Value-guided search for efficient chain-of-thought reasoning,

    K. Wang, J. P. Zhou, J. Chang, Z. Gao, N. Kallus, K. Brantley, and W. Sun, “Value-guided search for efficient chain-of-thought reasoning,” 2025

  108. [115]

    REFINER: Reasoning feedback on intermediate representations,

    D. Paul, M. Ismayilzada, M. Peyrard, B. Borges, A. Bosselut, R. West, and B. Faltings, “REFINER: Reasoning feedback on intermediate representations,” inProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long P...

  109. [116]

    LM2: A simple society of language models solves complex reasoning,

    G. Juneja, S. Dutta, and T. Chakraborty, “LM2: A simple society of language models solves complex reasoning,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Associa...

  110. [117]

    Enhancing mathematical reasoning in llms by stepwise correction,

    Z. Wu, Q. Zeng, Z. Zhang, Z. Tan, C. Shen, and M. Jiang, “Enhancing mathematical reasoning in llms by stepwise correction,” 2024

  111. [118]

    Enhancing llm reasoning via critique models with test-time and training-time supervision,

    Z. Xi, D. Yang, J. Huang, J. Tang, G. Li, Y . Ding, W. He, B. Hong, S. Do, W. Zhan, X. Wang, R. Zheng, T. Ji, X. Shi, Y . Zhai, R. Weng, J. Wang, X. Cai, T. Gui, Z. Wu, Q. Zhang, X. Qiu, X. Huang, and Y .-G. Jiang, “Enhancing llm reasoning via critique models with test-time an...

  112. [119]

    Reinforcing thinking through reasoning-enhanced reward models,

    D. Yang, L. Zeng, K. Chen, and Y . Zhang, “Reinforcing thinking through reasoning-enhanced reward models,” 2024

  113. [120]

    Reward-guided speculative decoding for efficient llm reasoning,

    B. Liao, Y . Xu, H. Dong, J. Li, C. Monz, S. Savarese, D. Sahoo, and C. Xiong, “Reward-guided speculative decoding for efficient llm reasoning,” 2025

  114. [121]

    Rest- mcts*: Llm self-training via process reward guided tree search,

    D. Zhang, S. Zhoubian, Z. Hu, Y . Yue, Y . Dong, and J. Tang, “Rest- mcts*: Llm self-training via process reward guided tree search,” 2024

  115. [122]

    Imitate, explore, and self-improve: A reproduction report on slow- thinking reasoning systems,

    Y . Min, Z. Chen, J. Jiang, J. Chen, J. Deng, Y . Hu, Y . Tang, J. Wang, X. Cheng, H. Song, W. X. Zhao, Z. Liu, Z. Wang, and J.-R. Wen, “Imitate, explore, and self-improve: A reproduction report on slow- thinking reasoning systems,” 2024

  116. [123]

    Enhancing reasoning through process supervision with monte carlo tree search,

    S. Li, S. Dong, K. Luan, X. Di, and C. Ding, “Enhancing reasoning through process supervision with monte carlo tree search,” 2025

  117. [124]

    rstar-math: Small llms can master math reasoning with self-evolved deep thinking,

    X. Guan, L. L. Zhang, Y . Liu, N. Shang, Y . Sun, Y . Zhu, F. Yang, and M. Yang, “rstar-math: Small llms can master math reasoning with self-evolved deep thinking,” 2025

  118. [125]

    Self-rewarding language models,

    W. Yuan, R. Y . Pang, K. Cho, X. Li, S. Sukhbaatar, J. Xu, and J. Weston, “Self-rewarding language models,” 2025

  119. [126]

    Learning planning-based reasoning by trajectories collection and process reward synthesizing,

    F. Jiao, C. Qin, Z. Liu, N. F. Chen, and S. Joty, “Learning planning-based reasoning by trajectories collection and process reward synthesizing,” 2024

  120. [127]

    Monte carlo tree search boosts reasoning via iterative preference learning,

    Y . Xie, A. Goyal, W. Zheng, M.-Y . Kan, T. P. Lillicrap, K. Kawaguchi, and M. Shieh, “Monte carlo tree search boosts reasoning via iterative preference learning,” 2024

  121. [128]

    Alphamath almost zero: Process supervision without process,

    G. Chen, M. Liao, C. Li, and K. Fan, “Alphamath almost zero: Process supervision without process,” inAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, Inc., 20...

  122. [129]

    Step-dpo: Step-wise preference optimization for long-chain reasoning of llms,

    X. Lai, Z. Tian, Y . Chen, S. Yang, X. Peng, and J. Jia, “Step-dpo: Step-wise preference optimization for long-chain reasoning of llms,” 2024

  123. [130]

    Enhancing llm reasoning with reward-guided tree search,

    J. Jiang, Z. Chen, Y . Min, J. Chen, X. Cheng, J. Wang, Y . Tang, H. Sun, J. Deng, W. X. Zhao, Z. Liu, D. Yan, J. Xie, Z. Wang, and J.-R. Wen, “Enhancing llm reasoning with reward-guided tree search,” 2024

  124. [131]

    Longreward: Improving long-context large language models with ai feedback,

    J. Zhang, Z. Hou, X. Lv, S. Cao, Z. Hou, Y . Niu, L. Hou, Y . Dong, L. Feng, and J. Li, “Longreward: Improving long-context large language models with ai feedback,” 2024

  125. [132]

    Step-kto: Optimizing mathematical reasoning through stepwise binary feedback,

    Y .-T. Lin, D. Jin, T. Xu, T. Wu, S. Sukhbaatar, C. Zhu, Y . He, Y .-N. Chen, J. Weston, Y . Tian, A. Rahnama, S. Wang, H. Ma, and H. Fang, “Step-kto: Optimizing mathematical reasoning through stepwise binary feedback,” 2025

  126. [133]

    Full-step-dpo: Self-supervised preference optimization with step-wise rewards for mathematical reasoning,

    H. Xu, X. Mao, F.-L. Li, X. Wu, W. Chen, W. Zhang, and A. T. Luu, “Full-step-dpo: Self-supervised preference optimization with step-wise rewards for mathematical reasoning,” 2025

  127. [134]

    Enhancing llm reasoning with iterative dpo: A comprehensive empirical investigation,

    S. Tu, J. Lin, X. Tian, Q. Zhang, L. Li, Y . Fu, N. Xu, W. He, X. Lan, D. Jiang, and D. Zhao, “Enhancing llm reasoning with iterative dpo: A comprehensive empirical investigation,” 2025

  128. [135]

    Process-based self-rewarding language models,

    S. Zhang, X. Liu, X. Zhang, J. Liu, Z. Luo, S. Huang, and Y . Gong, “Process-based self-rewarding language models,” 2025

  129. [136]

    Cot-self-instruct: Building high- quality synthetic prompts for reasoning and non-reasoning tasks,

    P. Yu, J. Lanchantin, T. Wang, W. Yuan, O. Golovneva, I. Kulikov, S. Sukhbaatar, J. Weston, and J. Xu, “Cot-self-instruct: Building high- quality synthetic prompts for reasoning and non-reasoning tasks,” 2025

  130. [137]

    Reinforced self-training (rest) for language modeling,

    C. Gulcehre, T. L. Paine, S. Srinivasan, K. Konyushkova, L. Weerts, A. Sharma, A. Siddhant, A. Ahern, M. Wang, C. Gu, W. Macherey, A. Doucet, O. Firat, and N. de Freitas, “Reinforced self-training (rest) for language modeling,” 2023

  131. [138]

    Preference-guided reflective sampling for aligning language models,

    H. Ye and H. T. Ng, “Preference-guided reflective sampling for aligning language models,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Association for Computation...

  132. [139]

    Offline reinforcement learning for llm multi-step reasoning,

    H. Wang, S. Hao, H. Dong, S. Zhang, Y . Bao, Z. Yang, and Y . Wu, “Offline reinforcement learning for llm multi-step reasoning,” 2024

  133. [140]

    Improving multi-step reasoning abilities of large language models with direct advantage policy optimization,

    J. Liu, C. Wang, C. Y . Liu, L. Zeng, R. Yan, Y . Sun, Y . Liu, and Y . Zhou, “Improving multi-step reasoning abilities of large language models with direct advantage policy optimization,” 2024

  134. [141]

    Synthetic data generation & multi-step rl for reasoning & tool use,

    A. Goldie, A. Mirhoseini, H. Zhou, I. Cai, and C. D. Manning, “Synthetic data generation & multi-step rl for reasoning & tool use,” 2025

  135. [142]

    Proximal policy optimization algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017

  136. [143]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions...

  137. [144]

    Deepseekmath: Pushing the limits of mathematical reasoning in open language models,

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . K. Li, Y . Wu, and D. Guo, “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” 2024

  138. [145]

    Dapo: An open-source llm reinforcement learning system at scale,

    Q. Yu, Z. Zhang, R. Zhu, Y . Yuan, X. Zuo, Y . Yue, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y . Tong, C. Zhang, M. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, W. Dai, Y . Song, X. Wei, H. Zhou, J. Liu, W.-Y . Ma, Y .-Q. Zhang, L....

  139. [146]

    Crossing the reward bridge: Expanding rl with verifiable rewards across diverse domains,

    Y . Su, D. Yu, L. Song, J. Li, H. Mi, Z. Tu, M. Zhang, and D. Yu, “Crossing the reward bridge: Expanding rl with verifiable rewards across diverse domains,” 2025

  140. [147]

    Understanding r1-zero-like training: A critical perspective,

    Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin, “Understanding r1-zero-like training: A critical perspective,” 2025

  141. [148]

    Gpg: A simple and strong reinforcement learning baseline for model reasoning,

    X. Chu, H. Huang, X. Zhang, F. Wei, and Y . Wang, “Gpg: A simple and strong reinforcement learning baseline for model reasoning,” 2025

  142. [149]

    Ttrl: Test-time reinforcement learning,

    Y . Zuo, K. Zhang, S. Qu, L. Sheng, X. Zhu, B. Qi, Y . Sun, G. Cui, N. Ding, and B. Zhou, “Ttrl: Test-time reinforcement learning,” 2025

  143. [150]

    Evolving language models without labels: Majority drives selection, novelty promotes variation,

    Y . Zhou, Z. Liang, H. Liu, W. Yu, K. Panaganti, L. Song, D. Yu, X. Zhang, H. Mi, and D. Yu, “Evolving language models without labels: Majority drives selection, novelty promotes variation,” 2025

  144. [151]

    Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks,

    Y . Yue, Y . Yuan, Q. Yu, X. Zuo, R. Zhu, W. Xu, J. Chen, C. Wang, T. Fan, Z. Du, X. Wei, X. Yu, G. Liu, J. Liu, L. Liu, H. Lin, Z. Lin, B. Ma, C. Zhang, M. Zhang, W. Zhang, H. Zhu, R. Zhang, X. Liu, M. Wang, Y . Wu, and L. Yan, “Vapo: Efficient and reliable reinforcement lear...

  145. [152]

    Putting the value back in rl: Better test-time scaling by unifying llm reasoners with verifiers,

    K. Sareen, M. M. Moss, A. Sordoni, R. Agarwal, and A. Hosseini, “Putting the value back in rl: Better test-time scaling by unifying llm reasoners with verifiers,” 2025

  146. [153]

    Segment policy optimization: Effective segment-level credit assignment in rl for large language models,

    Y . Guo, L. Xu, J. Liu, D. Ye, and S. Qiu, “Segment policy optimization: Effective segment-level credit assignment in rl for large language models,” 2025

  147. [154]

    Rlver: 19 Reinforcement learning with verifiable emotion rewards for empathetic agents,

    P. Wang, R. Ma, B. Zhang, X. Chen, Z. He, K. Luo, Q. Lv, Q. Jiang, Z. Xie, S. Wang, Y . Li, F. Ye, J. Li, Y . Yang, Z. Tu, and X. Li, “Rlver: 19 Reinforcement learning with verifiable emotion rewards for empathetic agents,” 2025

  148. [155]

    On designing effective rl reward at training time for llm reasoning,

    J. Gao, S. Xu, W. Ye, W. Liu, C. He, W. Fu, Z. Mei, G. Wang, and Y . Wu, “On designing effective rl reward at training time for llm reasoning,” 2024

  149. [156]

    Steptool: Enhancing multi-step tool usage in llms through step-grained reinforcement learning,

    Y . Yu, Z. Wang, W. Ma, S. Wang, C. Wu, Z. Guo, and M. Zhang, “Steptool: Enhancing multi-step tool usage in llms through step-grained reinforcement learning,” 2025

  150. [157]

    BackMATH: Towards backward reasoning for solving math problems step by step,

    S. Zhang and D. Xiong, “BackMATH: Towards backward reasoning for solving math problems step by step,” inProceedings of the 31st International Conference on Computational Linguistics: Industry Track, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schockae...

  151. [158]

    Process reinforcement through implicit rewards,

    G. Cui, L. Yuan, Z. Wang, H. Wang, W. Li, B. He, Y . Fan, T. Yu, Q. Xu, W. Chen, J. Yuan, H. Chen, K. Zhang, X. Lv, S. Wang, Y . Yao, X. Han, H. Peng, Y . Cheng, Z. Liu, M. Sun, B. Zhou, and N. Ding, “Process reinforcement through implicit rewards,” 2025

  152. [159]

    Stop summation: Min-form credit assignment is all process reward model needs for reasoning,

    J. Cheng, R. Qiao, L. Li, C. Guo, J. Wang, G. Xiong, Y . Lv, and F.-Y . Wang, “Stop summation: Min-form credit assignment is all process reward model needs for reasoning,” 2025

  153. [160]

    Exploring the limit of outcome reward for learning mathematical reasoning,

    C. Lyu, S. Gao, Y . Gu, W. Zhang, J. Gao, K. Liu, Z. Wang, S. Li, Q. Zhao, H. Huang, W. Cao, J. Liu, H. Liu, J. Liu, S. Zhang, D. Lin, and K. Chen, “Exploring the limit of outcome reward for learning mathematical reasoning,” 2025

  154. [161]

    o1-coder: an o1 replication for coding,

    Y . Zhang, S. Wu, Y . Yang, J. Shu, J. Xiao, C. Kong, and J. Sang, “o1-coder: an o1 replication for coding,” 2024

  155. [162]

    Reward- sql: Boosting text-to-sql via stepwise reasoning and process-supervised rewards,

    Y . Zhang, M. Fan, J. Fan, M. Yi, Y . Luo, J. Tan, and G. Li, “Reward- sql: Boosting text-to-sql via stepwise reasoning and process-supervised rewards,” 2025

  156. [163]

    Beyond correctness: Harmonizing process and outcome rewards through rl training,

    C. Ye, Z. Yu, Z. Zhang, H. Chen, N. Sadagopan, J. Huang, T. Zhang, and A. Beniwal, “Beyond correctness: Harmonizing process and outcome rewards through rl training,” 2025

  157. [164]

    Self-consistency improves chain of thought reasoning in language models,

    X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdh- ery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” 2023

  158. [165]

    Selfcheck: Using llms to zero- shot check their own step-by-step reasoning,

    N. Miao, Y . W. Teh, and T. Rainforth, “Selfcheck: Using llms to zero- shot check their own step-by-step reasoning,” 2023

  159. [166]

    ARGS: Alignment as reward- guided search,

    M. Khanov, J. Burapacheep, and Y . Li, “ARGS: Alignment as reward- guided search,” inThe Twelfth International Conference on Learning Representations, 2024

  160. [167]

    Language agent tree search unifies reasoning acting and planning in language models,

    A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y .-X. Wang, “Language agent tree search unifies reasoning acting and planning in language models,” 2024

  161. [168]

    Alphazero-like tree-search can guide large language model decoding and training,

    X. Feng, Z. Wan, M. Wen, S. M. McAleer, Y . Wen, W. Zhang, and J. Wang, “Alphazero-like tree-search can guide large language model decoding and training,” 2024

  162. [169]

    Teaching large language models to self-debug,

    X. Chen, M. Lin, N. Schärli, and D. Zhou, “Teaching large language models to self-debug,” 2023

  163. [170]

    Critic: Large language models can self-correct with tool-interactive critiquing,

    Z. Gou, Z. Shao, Y . Gong, Y . Shen, Y . Yang, N. Duan, and W. Chen, “Critic: Large language models can self-correct with tool-interactive critiquing,” 2024

  164. [171]

    Large language models cannot self-correct reasoning yet,

    J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou, “Large language models cannot self-correct reasoning yet,” 2024

  165. [172]

    When can LLMs actually correct their own mistakes? a critical survey of self- correction of LLMs,

    R. Kamoi, Y . Zhang, N. Zhang, J. Han, and R. Zhang, “When can LLMs actually correct their own mistakes? a critical survey of self- correction of LLMs,”Transactions of the Association for Computational Linguistics, vol. 12, pp. 1417–1440, 2024

  166. [173]

    Recursive introspection: Teaching language model agents how to self-improve,

    Y . Qu, T. Zhang, N. Garg, and A. Kumar, “Recursive introspection: Teaching language model agents how to self-improve,” 2024

  167. [174]

    Training language models to self-correct via reinforcement learning,

    A. Kumar, V . Zhuang, R. Agarwal, Y . Su, J. D. Co-Reyes, A. Singh, K. Baumli, S. Iqbal, C. Bishop, R. Roelofs, L. M. Zhang, K. McKinney, D. Shrivastava, C. Paduraru, G. Tucker, D. Precup, F. Behbahani, and A. Faust, “Training language models to self-correct via reinforcement ...

  168. [175]

    S 2r: Teaching llms to self-verify and self-correct via reinforcement learning,

    R. Ma, P. Wang, C. Liu, X. Liu, J. Chen, B. Zhang, X. Zhou, N. Du, and J. Li, “S 2r: Teaching llms to self-verify and self-correct via reinforcement learning,” 2025

  169. [176]

    Teaching llms for step-level automatic math correction via reinforcement learning,

    J. Li, J. Zhou, Y . Yang, B. Zhan, Q. Pan, Y . Ding, Q. Chen, J. Bo, X. Lin, and L. He, “Teaching llms for step-level automatic math correction via reinforcement learning,” 2025

  170. [177]

    Self- rewarding correction for mathematical reasoning,

    W. Xiong, H. Zhang, C. Ye, L. Chen, N. Jiang, and T. Zhang, “Self- rewarding correction for mathematical reasoning,” 2025

  171. [178]

    Pag: Multi-turn reinforced llm self-correction with policy as generative verifier,

    Y . Jiang, Y . Xiong, Y . Yuan, C. Xin, W. Xu, Y . Yue, Q. Zhao, and L. Yan, “Pag: Multi-turn reinforced llm self-correction with policy as generative verifier,” 2025

  172. [179]

    Self-refine: Iterative refinement with self-feedback,

    A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y . Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Self-refine: Iterative refinement with self-feedback,” 2023

  173. [180]

    Reflexion: Language agents with verbal reinforcement learning,

    N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” 2023

  174. [181]

    Raft: Reward ranked finetuning for generative foundation model alignment,

    H. Dong, W. Xiong, D. Goyal, Y . Zhang, W. Chow, R. Pan, S. Diao, J. Zhang, K. Shum, and T. Zhang, “Raft: Reward ranked finetuning for generative foundation model alignment,” 2023

  175. [182]

    Rrhf: Rank responses to align language models with human feedback without tears,

    Z. Yuan, H. Yuan, C. Tan, W. Wang, S. Huang, and F. Huang, “Rrhf: Rank responses to align language models with human feedback without tears,” 2023

  176. [183]

    Star: Bootstrapping reasoning with reasoning,

    E. Zelikman, Y . Wu, J. Mu, and N. D. Goodman, “Star: Bootstrapping reasoning with reasoning,” 2022

  177. [184]

    Scaling relationship on learning mathematical reasoning with large language models,

    Z. Yuan, H. Yuan, C. Li, G. Dong, K. Lu, C. Tan, C. Zhou, and J. Zhou, “Scaling relationship on learning mathematical reasoning with large language models,” 2023

  178. [185]

    V-star: Training verifiers for self-taught reasoners,

    A. Hosseini, X. Yuan, N. Malkin, A. Courville, A. Sordoni, and R. Agarwal, “V-star: Training verifiers for self-taught reasoners,” 2024

  179. [186]

    Self-evolved reward learning for llms,

    C. Huang, Z. Fan, L. Wang, F. Yang, P. Zhao, Z. Lin, Q. Lin, D. Zhang, S. Rajmohan, and Q. Zhang, “Self-evolved reward learning for llms,” 2025

  180. [187]

    Personalizing reinforcement learning from human feedback with variational preference learning,

    S. Poddar, Y . Wan, H. Ivison, A. Gupta, and N. Jaques, “Personalizing reinforcement learning from human feedback with variational preference learning,” 2024

  181. [188]

    Iterative reasoning preference optimization,

    R. Y . Pang, W. Yuan, K. Cho, H. He, S. Sukhbaatar, and J. Weston, “Iterative reasoning preference optimization,” 2024

  182. [189]

    Magistral,

    Mistral-AIet al., “Magistral,” 2025

  183. [190]

    Policy gradient methods for reinforcement learning with function approximation,

    R. S. Sutton, D. McAllester, S. Singh, and Y . Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems, S. Solla, T. Leen, and K. Müller, Eds., vol. 12. MIT Press, 1999

  184. [191]

    Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms,

    A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker, “Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms,” 2024

  185. [192]

    Reinforce++: An efficient rlhf algorithm with robustness to both prompt and reward models,

    J. Hu, J. K. Liu, and W. Shen, “Reinforce++: An efficient rlhf algorithm with robustness to both prompt and reward models,” 2025

  186. [193]

    Defining and characterizing reward hacking,

    J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger, “Defining and characterizing reward hacking,” 2025

  187. [194]

    Concrete problems in ai safety,

    D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané, “Concrete problems in ai safety,” 2016

  188. [195]

    Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective,

    T. Everitt, M. Hutter, R. Kumar, and V . Krakovna, “Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective,” 2021

  189. [196]

    The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation and artificial life research communities,

    J. Lehman, J. Clune, D. Misevic, C. Adami, L. Altenberg, J. Beaulieu, P. J. Bentley, S. Bernard, G. Beslon, D. M. Bryson, P. Chrabaszcz, N. Cheney, A. Cully, S. Doncieux, F. C. Dyer, K. O. Ellefsen, R. Feldt, S. Fischer, S. Forrest, A. Frénoy, C. Gagné, L. L. Goff, L. M. Grabo...

  190. [197]

    Language models learn to mislead humans via rlhf,

    J. Wen, R. Zhong, A. Khan, E. Perez, J. Steinhardt, M. Huang, S. R. Bowman, H. He, and S. Feng, “Language models learn to mislead humans via rlhf,” 2024

  191. [198]

    Sycophancy to subterfuge: Investigating reward-tampering in large language models,

    C. Denison, M. MacDiarmid, F. Barez, D. Duvenaud, S. Kravec, S. Marks, N. Schiefer, R. Soklaski, A. Tamkin, J. Kaplan, B. Shlegeris, S. R. Bowman, E. Perez, and E. Hubinger, “Sycophancy to subterfuge: Investigating reward-tampering in large language models,” 2024

  192. [199]

    Towards understanding sycophancy in language models,

    M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, N. Cheng, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, S. Kravec, T. Maxwell, S. McCandlish, K. Ndousse, O. Rausch, N. Schiefer, D. Yan, M. Zhang, and E. Perez, “Towards understanding sycophancy in language ...

  193. [200]

    Chaos with keywords: Exposing large language models sycophantic hallucination to misleading keywords and evaluating defense strategies,

    A. RRV , N. Tyagi, M. N. Uddin, N. Varshney, and C. Baral, “Chaos with keywords: Exposing large language models sycophantic hallucination to misleading keywords and evaluating defense strategies,” 2024

  194. [201]

    Sycophancy in large language models: Causes and mitigations,

    L. Malmqvist, “Sycophancy in large language models: Causes and mitigations,” 2024

  195. [202]

    A long way to go: Investigating length correlations in rlhf,

    P. Singhal, T. Goyal, J. Xu, and G. Durrett, “A long way to go: Investigating length correlations in rlhf,” 2024. 20

  196. [203]

    Monitoring reasoning models for misbehavior and the risks of promoting obfuscation,

    B. Baker, J. Huizinga, L. Gao, Z. Dou, M. Y . Guan, A. Madry, W. Zaremba, J. Pachocki, and D. Farhi, “Monitoring reasoning models for misbehavior and the risks of promoting obfuscation,” 2025

  197. [204]

    Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking,

    J. Eisenstein, C. Nagpal, A. Agarwal, A. Beirami, A. D’Amour, D. Dvijotham, A. Fisch, K. Heller, S. Pfohl, D. Ramachandran, P. Shaw, and J. Berant, “Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking,” 2024

  198. [205]

    Spontaneous reward hacking in iterative self-refinement,

    J. Pan, H. He, S. R. Bowman, and S. Feng, “Spontaneous reward hacking in iterative self-refinement,” 2024

  199. [206]

    Feedback loops with language models drive in-context reward hacking,

    A. Pan, E. Jones, M. Jagadeesan, and J. Steinhardt, “Feedback loops with language models drive in-context reward hacking,” 2024

  200. [207]

    Reward model ensembles help mitigate overoptimization,

    T. Coste, U. Anwar, R. Kirk, and D. Krueger, “Reward model ensembles help mitigate overoptimization,” 2024

  201. [208]

    Warm: On the benefits of weight averaged reward models,

    A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret, “Warm: On the benefits of weight averaged reward models,” 2024

  202. [209]

    Agentic reward modeling: Integrating human preferences with verifiable correctness signals for reliable reward systems,

    H. Peng, Y . Qi, X. Wang, Z. Yao, B. Xu, L. Hou, and J. Li, “Agentic reward modeling: Integrating human preferences with verifiable correctness signals for reliable reward systems,” 2025

  203. [210]

    Beyond reward hacking: Causal rewards for large language model alignment,

    C. Wang, Z. Zhao, Y . Jiang, Z. Chen, C. Zhu, Y . Chen, J. Liu, L. Zhang, X. Fan, H. Ma, and S. Wang, “Beyond reward hacking: Causal rewards for large language model alignment,” 2025

  204. [211]

    Reward shaping to mitigate reward hacking in rlhf,

    J. Fu, X. Zhao, C. Yao, H. Wang, Q. Han, and Y . Xiao, “Reward shaping to mitigate reward hacking in rlhf,” 2025

  205. [212]

    Odin: Disentangled reward mitigates hacking in rlhf,

    L. Chen, C. Zhu, D. Soselia, J. Chen, T. Zhou, T. Goldstein, H. Huang, M. Shoeybi, and B. Catanzaro, “Odin: Disentangled reward mitigates hacking in rlhf,” 2024

  206. [213]

    Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback,

    W. Shen, R. Zheng, W. Zhan, J. Zhao, S. Dou, T. Gui, Q. Zhang, and X. Huang, “Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback,” 2023

  207. [214]

    Robust reward modeling via causal rubrics,

    P. Srivastava, H. Singh, R. Madhavan, G. Patil, S. Addepalli, A. Suggala, R. Aravamudhan, S. Sharma, A. Laha, A. Raghuveer, K. Shanmugam, and D. Precup, “Robust reward modeling via causal rubrics,” 2025

  208. [215]

    Rank analysis of incomplete block designs: I. the method of paired comparisons,

    R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. the method of paired comparisons,”Biometrika, vol. 39, no. 3/4, pp. 324–345, 1952

  209. [216]

    On the robustness of reward models for language model alignment,

    J. Hong, N. Lee, E. Kim, G. Son, W. Chung, A. Gupta, S. Tang, and J. Thorne, “On the robustness of reward models for language model alignment,” 2025

  210. [217]

    Infopo: On mutual information maximization for large language model alignment,

    T. Xiao, Z. Ge, S. Sanghavi, T. Wang, J. Katz-Samuels, M. Versage, Q. Cui, and T. Chilimbi, “Infopo: On mutual information maximization for large language model alignment,” 2025

  211. [218]

    Secrets of rlhf in large language models part ii: Reward modeling,

    B. Wang, R. Zheng, L. Chen, Y . Liu, S. Dou, C. Huang, W. Shen, S. Jin, E. Zhou, C. Shi, S. Gao, N. Xu, Y . Zhou, X. Fan, Z. Xi, J. Zhao, X. Wang, T. Ji, H. Yan, L. Shen, Z. Chen, T. Gui, Q. Zhang, X. Qiu, X. Huang, Z. Wu, and Y .-G. Jiang, “Secrets of rlhf in large language m...

  212. [219]

    On the limited generalization capability of the implicit reward model induced by direct preference optimization,

    Y . Lin, S. Seto, M. ter Hoeve, K. Metcalf, B.-J. Theobald, X. Wang, Y . Zhang, C. Huang, and T. Zhang, “On the limited generalization capability of the implicit reward model induced by direct preference optimization,” 2024

  213. [220]

    Regularizing hidden states enables learning generalizable reward model for llms,

    R. Yang, R. Ding, Y . Lin, H. Zhang, and T. Zhang, “Regularizing hidden states enables learning generalizable reward model for llms,” 2024

  214. [221]

    Generative reward models,

    D. Mahan, D. V . Phung, R. Rafailov, C. Blagden, N. Lile, L. Castricato, J.-P. Fränken, C. Finn, and A. Albalak, “Generative reward models,” 2024

  215. [222]

    Training compute-optimal large language models,

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifr...

  216. [223]

    Is prm necessary? problem-solving rl implicitly induces prm capability in llms,

    Z. Feng, Q. Chen, N. Lu, Y . Li, S. Cheng, S. Peng, D. Tang, S. Liu, and Z. Zhang, “Is prm necessary? problem-solving rl implicitly induces prm capability in llms,” 2025

  217. [224]

    Do we need to verify step by step? rethinking process supervision from a theoretical perspective,

    Z. Jia, A. Rakhlin, and T. Xie, “Do we need to verify step by step? rethinking process supervision from a theoretical perspective,” 2025

  218. [225]

    A baseline analysis of reward models’ ability to accurately analyze foundation models under distribution shift,

    W. LeVine, B. Pikus, A. Chen, and S. Hendryx, “A baseline analysis of reward models’ ability to accurately analyze foundation models under distribution shift,” 2024

  219. [226]

    Easy- to-hard generalization: Scalable alignment beyond human supervision,

    Z. Sun, L. Yu, Y . Shen, W. Liu, Y . Yang, S. Welleck, and C. Gan, “Easy- to-hard generalization: Scalable alignment beyond human supervision,” 2024

  220. [227]

    Thinkbench: Dynamic out-of-distribution evaluation for robust llm reasoning,

    S. Huang, L. Yang, Y . Song, S. Chen, L. Cui, Z. Wan, Q. Zeng, Y . Wen, K. Shao, W. Zhang, J. Wang, and Y . Zhang, “Thinkbench: Dynamic out-of-distribution evaluation for robust llm reasoning,” 2025

  221. [228]

    Uncertainty- aware reward model: Teaching reward models to know what is unknown,

    X. Lou, D. Yan, W. Shen, Y . Yan, J. Xie, and J. Zhang, “Uncertainty- aware reward model: Teaching reward models to know what is unknown,” 2025

  222. [229]

    Helpsteer2: Open-source dataset for training top-performing reward models,

    Z. Wang, Y . Dong, O. Delalleau, J. Zeng, G. Shen, D. Egert, J. J. Zhang, M. N. Sreedhar, and O. Kuchaiev, “Helpsteer2: Open-source dataset for training top-performing reward models,” 2024

  223. [230]

    Versaprm: Multi-domain process reward model via synthetic reasoning data,

    T. Zeng, S. Zhang, S. Wu, C. Classen, D. Chae, E. Ewer, M. Lee, H. Kim, W. Kang, J. Kunde, Y . Fan, J. Kim, H. I. Koo, K. Ramchandran, D. Papailiopoulos, and K. Lee, “Versaprm: Multi-domain process reward model via synthetic reasoning data,” 2025

  224. [231]

    Worldpm: Scaling human preference modeling,

    B. Wang, R. Lin, K. Lu, L. Yu, Z. Zhang, F. Huang, C. Zheng, K. Dang, Y . Fan, X. Ren, A. Yang, B. Hui, D. Liu, T. Gui, Q. Zhang, X. Huang, Y .-G. Jiang, B. Yu, J. Zhou, and J. Lin, “Worldpm: Scaling human preference modeling,” 2025

  225. [232]

    Qwen2.5: A party of foundation models,

    Q. Team, “Qwen2.5: A party of foundation models,” September 2024

  226. [233]

    An implementation of generative prm,

    W. Xiong, H. Zhang, N. Jiang, and T. Zhang, “An implementation of generative prm,” https://github.com/RLHFlow/RLHF-Reward-Modeling, 2024

  227. [234]

    Skywork-o1 open series,

    J. He, T. Wei, R. Yan, J. Liu, C. Wang, Y . Gan, S. Tu, C. Y . Liu, L. Zeng, X. Wang, B. Wang, Y . Li, F. Zhang, J. Xu, B. An, Y . Liu, and Y . Zhou, “Skywork-o1 open series,” https://huggingface.co/Skywork, November 2024

  228. [236]

    Rethinking reward model evaluation: Are we barking up the wrong tree?

    X. Wen, J. Lou, Y . Lu, H. Lin, X. Yu, X. Lu, B. He, X. Han, D. Zhang, and L. Sun, “Rethinking reward model evaluation: Are we barking up the wrong tree?” 2025

  229. [238]

    Openr: An open source framework for advanced reasoning with large language models,

    J. Wang, M. Fang, Z. Wan, M. Wen, J. Zhu, A. Liu, Z. Gong, Y . Song, L. Chen, L. M. Ni, L. Yang, Y . Wen, and W. Zhang, “Openr: An open source framework for advanced reasoning with large language models,” 2024. 21 TABLE VI: As a representative of generative RMs, GenPRM-7B [47]...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.