Pith. sign in

REVIEW 4 major objections 6 minor 35 references

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims that process-supervised rewards should depend on both the accuracy and the length of a reasoning chain, and that the relationship is nonlinear, reporting that a method built on this principle outperforms mainstream…

desk verdict A sensible reward-shaping idea undermined by an ablation table that, as printed, contradicts the paper's main claim. read the letter →

arxiv 2411.11681 v3 pith:GKT5YPL2 submitted 2024-11-18 cs.AI cs.LG

classification cs.AIcs.LG
keywords processsupervisionrewardshapingchain-of-thoughtreasoningalignmentpolicyoptimizationWeibulldistributionmathematicallargelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that process supervision of chain-of-thought reasoning should reward not just the correctness of each step but also the length of the reasoning chain, and that the two combine nonlinearly into an overall reward score. It proposes a full workflow, PSPO*, that trains a step-level reward model and then optimizes a policy against a reward in which per-step probabilities are combined by geometric mean and then shaped by an adjusted Weibull distribution that peaks at a preferred number of steps. The paper asserts that it is the first to claim nonlinearity of the reward score in reasoning alignment, and it reports that the resulting method, PSPO-WRS, outperforms current baselines on six mathematical reasoning datasets while also improving out-of-distribution generalization. A sympathetic reading takes the central claim to be that length-aware, nonlinear process rewards make reasoning policies produce more accurate and more complete chains.

What carries the argument

The load-bearing object is the combined reward: a geometric-mean accumulation function $F = (\prod_{j=1}^t R^s_j)^{1/t}$, where $R^s_j$ is the reward model's probability that step $j$ is positive, multiplied by an adjusted Weibull reward-shaping function $R_s = C \frac{k}{\lambda}(\frac{t}{\lambda})^{k-1} e^{-(t/\lambda)^k}$. The geometric mean removes the automatic penalty that plain multiplication imposes on long chains, and the Weibull factor injects the prior that chains of roughly $\lambda = 8$ steps deserve the largest shaping reward, with $k=1.5$ controlling how sharply the bonus falls off. The policy objective is $J(\pi) = E_\pi[R_s F - \beta D_{KL}(\pi \| \pi_{ref})]$. This machinery carries the paper's evidence: the ablation that removes $R_s$ degenerates to the standard product-accumulation method of Lightman et al. and loses performance, while the full method shifts the policy toward longer, higher-scoring chains.

What would settle it

Take the same PSPO-WRS pipeline and run it on problems whose correct solutions require very long chains, for example 15-25 steps, while also sweeping $\lambda$ over a grid from 4 to 16 with all else fixed: the nonlinearity claim predicts the method should keep beating the linear baseline across a range of $\lambda$, whereas the length-prior artifact predicts performance peaks only near $\lambda=8$ and collapses on long-chain problems.

Watch

Extended reading notes

Core claim

Process supervision, as usually implemented, scores a reasoning chain by multiplying the per-step correctness probabilities, which silently punishes longer chains even when every step is right. The paper's central discovery claim is that this multiplicative scheme is wrong in two ways: the overall reward should depend on the number of steps as well as their accuracy, and the dependence is nonlinear rather than linear. Its concrete proposal replaces the plain product with the geometric mean $F = (\prod_{j=1}^t P(y_j=1 \mid x, y_{<j}))^{1/t}$, which removes the length bias, and then multiplies $F$ by a Weibull-based shaping factor $R_s = C \frac{k}{\lambda}(\frac{t}{\lambda})^{k-1} e^{-(t/\lambda)^k}$ with $C=10.735$, $k=1.5$, $\lambda=8.0$, encoding a prior that chains of moderate length are best. In experiments on the six MATH-derived datasets, PSPO-WRS beats the Abel-7B baseline by between 6.56 and 30.94 percentage points and often matches or exceeds much larger models such as GPT-3.5 and Qwen2-72B; ablation shows the nonlinear module also raises the average number of generated steps from 2.624 to 3.051. The paper concludes that nonlinear, length-aware rewards are what make process supervision work.

Load-bearing premise

The load-bearing premise is that the hand-chosen Weibull curve, which gives the largest shaping reward to chains of about eight steps, correctly captures the true relationship between step count and reward quality; a single bad curve could erase the reported gains.

Editorial extensions

If this is right

  • Process supervision pipelines should stop scoring chains with a raw product of step probabilities, because that construction systematically discourages longer reasoning.
  • Step-count-normalized rewards plus a length prior can be added to any PRM-trained policy, not just the specific base model used here.
  • Models trained this way generate more complete reasoning chains on average, which is the behavior the paper ties to higher accuracy.
  • Because the gain persists on out-of-distribution datasets, the improvement is attributed to stronger reasoning rather than to memorizing the evaluation sets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper tests exactly one shaping curve, so its experiments do not yet separate 'nonlinearity helps' from 'this particular Weibull shape helps'; our inference is that a comparison against other nonlinear shapes, such as a log-length bonus or a bounded length cap, is needed to pin down the active ingredient.
  • The chosen peak at eight steps is likely tuned to the difficulty of the six datasets; our inference is that portable versions of the method would need a procedure for setting $\lambda$ per task or per problem, otherwise the same method could hurt on puzzles whose correct chains are much shorter or much longer.
  • The geometric mean already removes the main length penalty by itself; our inference is that an ablation keeping the geometric mean but dropping the Weibull factor would reveal how much of the gain is due to normalization versus the nonlinear prior.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript proposes PSPO*, a paradigm for process-supervised policy optimization, and PSPO-WRS, a concrete instance that combines geometric-mean step accuracy with an adjusted Weibull reward-shaping function. The central claims are that process-supervision effectiveness depends on both accuracy and reasoning-chain length, that the reward score is nonlinear, and that PSPO-WRS consistently outperforms current models on six mathematical reasoning datasets. The paper includes an out-of-distribution generalization check and a code link.

Significance. If the empirical claims hold, the paper offers a useful systematization of process supervision and draws attention to length bias in multiplicative step rewards. The out-of-distribution experiments in Table 3 are a positive check against overfitting, and the promise of released code supports reproducibility. However, the sole ablation isolating the nonlinear-reward contribution is internally inconsistent as printed, and the fixed Weibull parameters are neither tuned nor subjected to sensitivity analysis, so the central mechanistic claim is currently not established.

major comments (4)
  1. [Table 4 / Ablation Analysis] As printed, Table 4's PSPO-WRS column is misaligned: the entries 64.91, 67.60, 71.57, 52.29, and 54.70 are the Table 2 values for NewsNLI, RedditNLI, RTE-Quant, StressTest, and QQA, respectively, not the values for the rows in which they appear. Reading the table at face value, the w/o-nonlinear baseline outperforms PSPO-WRS on AwpNLI (69.12 vs 64.91), RTE-Quant (59.66 vs 52.29), and QQA (54.94 vs 54.70), directly contradicting the text that nonlinear rewards improve performance across all evaluated datasets. StressTest also shows 53.57 in Table 4 versus 52.29 in Table 2. Since Table 4 is the only experiment that isolates the nonlinear-reward contribution, the main mechanistic claim is not reproducible from the manuscript as written; the authors must correct the table and re-state the conclusions accordingly.
  2. [Experimental Setups / Eq. (9)] The Weibull reward-shaping parameters C=10.735, k=1.5, and lambda=8.0 are fixed without any tuning procedure, grid search, or sensitivity analysis. Because Eq. (9) directly determines how the reward depends on step count, the observed gains could be artifacts of these specific values. Please report an ablation over these parameters or justify the choice from properties of the datasets.
  3. [Eqs. (9)-(10) / Ablation Analysis] The nonlinearity is introduced by construction: Eq. (10) length-normalizes the step-product reward and Eq. (9) applies a Weibull-shaped length prior. The comparison against a baseline without length shaping therefore tests the presence of a length-dependent reward, not the specific Weibull form or the broader claim that the underlying reward is nonlinear. I recommend comparing against alternative length reweighting schemes (e.g., linear normalization, logarithmic penalty, or a quadratic length model) and reporting whether the Weibull form gives materially better results.
  4. [Experimental Results / Metrics and Parameters] All experimental comparisons appear to be single runs with no error bars, confidence intervals, or multiple training seeds. PPO training is known to be high-variance, and some reported differences in Table 4 are small (e.g., QQA 54.94 vs 54.70). Please report means and variances over at least three seeds, or explicitly justify why seed variation is negligible.
minor comments (6)
  1. [Throughout] The dataset name is written inconsistently as "RTE-Quant" and "RTE Quant" across the text and tables; please standardize.
  2. [Introduction] The phrase "we concrete the PSPO* paradigm" should be "we instantiate" or "we concretize the PSPO* paradigm".
  3. [Eqs. (2)-(3)] The notation is overloaded: Eq. (2) writes p(zi) = sigma(zi), but sigma is later defined as softmax in Eq. (3), which is not a component-wise sigmoid; please define the activation consistently.
  4. [Figure 4] The caption contains the typo "Prameter settings" and should read "Parameter settings".
  5. [Table 4] The row "Num. of steps" is not defined in the caption; please state which policy's average step count is reported and over which evaluation set.
  6. [Introduction / Contribution 2] The claim "We are the first to assert that the reward score in the reasoning alignment is nonlinear" is stronger than what the related-work discussion supports; product-of-probabilities rewards are already nonlinear in step count, so prior work implicitly involves nonlinearity even if not framed that way.

Circularity Check

1 steps flagged · score 6.0 of 10

The claimed 'discovery' that process-supervision reward scores are nonlinear is built into the reward definition (Eqs. 9-11), and the ablation meant to validate it is internally inconsistent with Table 2.

  1. self definitional [Section 'PSPO-WRS: Process-supervised Policy Optimization with Nonlinear Reward Shaping', Eqs. (9)-(11); Ablation Analysis, Table 4]
    "Additionally, we are the first to assert that the reward score in the reasoning alignment is nonlinear. ... Based on this prior knowledge, we employ the Adjusted Weibull distribution to shape the rewards ... Rs = C ∗ k λ ( t λ )k−1e−(t/λ)k , (9) ... F = [ tY j=1 P (yj = 1|xj, yj pre)]1/t. (10) ... Following the removal of the nonlinearity module, there is a noticeable decline in the performance of PSPO-WRS. ... These findings underscore the critical importance of nonlinear rewards in the process supervision."

    The paper's headline claim that the process-supervision reward score is nonlinear is an input, not a derived result. Eq. (9) defines the shaping term Rs as a Weibull function of the step count t, and Eq. (10) defines F as a geometric mean of per-step probabilities; both are nonlinear functions, so the optimized objective RsF in Eq. (11) is nonlinear by construction. The ablation analysis then treats a comparison between this constructed reward and a version without the Weibull factor as empirical support for the 'hypothesis' that reward scores are nonlinear. No independent, parameter-free measurement of a ground-truth nonlinear relationship is presented, and the 'w/o nonlinear' baseline still contains the nonlinear geometric-mean accumulation F.

full rationale

The PSPO* paradigm itself and the PSPO-WRS instantiation are largely self-contained engineering contributions: reward-model training, PPO-style policy optimization, and head-to-head comparisons against external baselines do not depend on any load-bearing self-citation chain or on a fitted parameter being relabeled as a prediction. The circularity is concentrated in the conceptual claim of nonlinearity. Equations (9)-(11) make the final reward score a product of a Weibull function of step count and a geometric mean of step accuracies, so the statement that the reward score is nonlinear is true by definition; the later statement that experiments 'validate' the nonlinearity hypothesis is therefore a check of the authors' own design choice, not an independent discovery. The out-of-distribution results (Table 3) support the overall method but do not isolate the nonlinear contribution, so they cannot repair this definitional circularity. Additionally, Table 4 as printed contradicts Table 2 (e.g., the PSPO-WRS entries 64.91, 67.60, 71.57, 52.29, and 54.70 are Table 2's values for different datasets, and StressTest is shown as 53.57 instead of 52.29), so the only ablation that claims to validate the nonlinear module is not reproducible as written. These are correctness concerns as well, but they reinforce the conclusion that the central mechanistic claim currently stands on the definition rather than on solid evidence. A score of 6 reflects partial circularity: the key theoretical claim reduces to construction, while the empirical result that a length-aware Weibull-shaped reward can beat baselines retains independent content.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central method rests on three hand-chosen Weibull parameters and two assumed functional forms; no entity is invented. The reward model training uses standard classification, and the policy optimization uses standard PPO/KL, so the novel content is concentrated in the reward-shaping choices.

free parameters (3)
  • C = 10.735
    Constant multiplier in Weibull reward shaping, equation (9), chosen by hand to adjust overall reward magnitude; no tuning procedure given.
  • k = 1.5
    Weibull shape parameter controlling the peak and skew of the step-length reward prior; fixed without sensitivity analysis.
  • lambda = 8.0
    Weibull scale parameter setting the characteristic step count; fixed without sensitivity analysis.
assumptions (3)
  • ad hoc to paper An adjusted Weibull distribution is an appropriate prior for how reasoning reward should depend on step count.
    Introduced in equation (9) from the generic statement that too few steps hurt accuracy and too many hurt efficiency; the specific distribution family and parameter values are not derived.
  • ad hoc to paper The geometric-mean accumulation function is the right way to combine step scores while removing length bias.
    Equation (10) normalizes the product by 1/t without justification beyond 'to eliminate the linear trend'; other normalizations are possible.
  • domain assumption Human step annotations are reliable and consistent across six datasets.
    The reward model is trained entirely on these labels; no inter-annotator agreement or label noise analysis is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment." pith.science (2026). https://pith.science/paper/GKT5YPL2

@misc{pith2026241111681,
  author       = {Pith},
  title        = {Pith review of: PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GKT5YPL2}},
  note         = {Machine review of arXiv:2411.11681}
}
read the original abstract

Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack of effective process supervision methods, even advanced large language models are prone to logical errors and redundant reasoning. We claim that the effectiveness of process supervision significantly depends on both the accuracy and the length of reasoning chains. Moreover, we identify that these factors exhibit a nonlinear relationship with the overall reward score of the reasoning process. Inspired by these insights, we propose a novel process supervision paradigm, PSPO*, which systematically outlines the workflow from reward model training to policy optimization, and highlights the importance of nonlinear rewards in process supervision. Based on PSPO*, we develop the PSPO-WRS, which considers the number of reasoning steps in determining reward scores and utilizes an adjusted Weibull distribution for nonlinear reward shaping. Experimental results on six mathematical reasoning datasets demonstrate that PSPO-WRS consistently outperforms current mainstream models.

Figures

Figures reproduced from arXiv: 2411.11681 by the authors.

Figure 1
Figure 1. An example from QQA dataset. The reasoning er [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The data annotation approach for PRM. Unlike [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The overall method of PSPO*. In the workflow of process supervision, we encompass a nonlinear accumulation [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: The results compared with ultra LLMs. It is note [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: The relationship between the length of reasoning [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: The relationship between the length of reasoning [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 10 canonical work pages

  1. [1]

    G.; Guo, Z

    Azar, M. G.; Guo, Z. D.; Piot, B.; Munos, R.; Rowland, M.; Valko, M.; and Calandriello, D. 2024. A General Theoretical Paradigm to Understand Learning from Human Preferences. In Dasgupta, S.; Mandt, S.; and Li, Y., eds., International Conference on Artificial Intelligence and Statistics, 2-4 May 2024, Palau de Congressos, Valencia, Spain, volume 238 of Pr...

  2. [2]

    Chen, C.-C.; Takamura, H.; Kobayashi, I.; and Miyao, Y. 2023. Improving Numeracy by Input Reframing and Quantitative Pre-Finetuning Task. In Vlachos, A.; and Augenstein, I., eds., Findings of the Association for Computational Linguistics: EACL 2023, 69--77. Dubrovnik, Croatia: Association for Computational Linguistics

  3. [3]

    Chern, E.; Zou, H.; Li, X.; Hu, J.; Feng, K.; Li, J.; and Liu, P. 2023. Generative AI for Math: Abel. https://github.com/GAIR-NLP/abel

  4. [4]

    Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve Math Word Problems. CoRR, abs/2110.14168

  5. [5]

    Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapo...

  6. [6]

    Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; Goyal, A.; Hartshorn, A.; Yang, A.; Mitra, A.; Sravankumar, A.; Korenev, A.; Hinsvark, A.; Rao, A.; Zhang, A.; Rodriguez, A.; Gregerson, A.; Spataru, A.; Roziere, B.; Biron, B.; Tang, B.; Chern, B.; Caucheteux, C.; Nayak, C.; Bi, C.; Marra...

  7. [7]

    Gao, L.; Schulman, J.; and Hilton, J. 2023. Scaling Laws for Reward Model Overoptimization. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , volume 202 of Proceedings of Machine Learning Research, 10835--10866. PMLR

  8. [8]

    J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W

    Hu, E. J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net

Show all 35 references
  1. [9]

    Hu, Y.; Wang, W.; Jia, H.; Wang, Y.; Chen, Y.; Hao, J.; Wu, F.; and Fan, C. 2020. Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing System...

  2. [10]

    Jin, M.; Yu, Q.; Shu, D.; Zhao, H.; Hua, W.; Meng, Y.; Zhang, Y.; and Du, M. 2024. The Impact of Reasoning Step Length on Large Language Models. CoRR, abs/2401.04925

  3. [11]

    S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y

    Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2023. Large Language Models are Zero-Shot Reasoners. arXiv:2205.11916

  4. [12]

    Lai, X.; Tian, Z.; Chen, Y.; Yang, S.; Peng, X.; and Jia, J. 2024. Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs. CoRR, abs/2406.18629

  5. [13]

    Levine, S.; Popovic, Z.; and Koltun, V. 2011. Nonlinear Inverse Reinforcement Learning with Gaussian Processes. In Shawe - Taylor, J.; Zemel, R. S.; Bartlett, P. L.; Pereira, F. C. N.; and Weinberger, K. Q., eds., Advances in Neural Information Processing Systems 24: 25th Annu...

  6. [14]

    Liang, X.; Li, J.; Yang, Y.; and Gao, Y. 2024. Bit \_ numeval at S em E val-2024 Task 7: Enhance Numerical Sensitivity and Reasoning Completeness for Quantitative Understanding. In Ojha, A. K.; Do g ru \"o z, A. S.; Tayyar Madabushi, H.; Da San Martino, G.; Rosenthal, S.; and ...

  7. [15]

    Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023. Let's Verify Step by Step. CoRR, abs/2305.20050

  8. [16]

    Luo, H.; Sun, Q.; Xu, C.; Zhao, P.; Lou, J.; Tao, C.; Geng, X.; Lin, Q.; Chen, S.; and Zhang, D. 2023. WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct. CoRR, abs/2308.09583

  9. [17]

    Ma, Q.; Zhou, H.; Liu, T.; Yuan, J.; Liu, P.; You, Y.; and Yang, H. 2023. Let's reward step by step: Step-Level reward model as the Navigators for Reasoning. CoRR, abs/2310.10080

  10. [18]

    L.; Bari, M

    Muennighoff, N.; Wang, T.; Sutawika, L.; Roberts, A.; Biderman, S.; Scao, T. L.; Bari, M. S.; Shen, S.; Yong, Z. X.; Schoelkopf, H.; Tang, X.; Radev, D.; Aji, A. F.; Almubarak, K.; Albanie, S.; Alyafeai, Z.; Webson, A.; Raff, E.; and Raffel, C. 2023. Crosslingual Generalizatio...

  11. [19]

    L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P

    Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P. F.; Leike, J.; and Lowe, R. 2022. Training language mo...

  12. [20]

    D.; Ermon, S.; and Finn, C

    Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Informat...

  13. [21]

    Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal Policy Optimization Algorithms. CoRR, abs/1707.06347

  14. [22]

    Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Canton - Ferrer, C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goy...

  15. [24]

    F.; Siegel, N

    Uesato, J.; Kushman, N.; Kumar, R.; Song, H. F.; Siegel, N. Y.; Wang, L.; Creswell, A.; Irving, G.; and Higgins, I. 2022 b . Solving math word problems with process- and outcome-based feedback. CoRR, abs/2211.14275

  16. [25]

    X.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z

    Wang, P.; Li, L.; Shao, Z.; Xu, R. X.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z. 2023 a . Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations. CoRR, abs/2312.08935

  17. [26]

    V.; Chi, E

    Wang, X.; Wei, J.; Schuurmans, D.; Le, Q. V.; Chi, E. H.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023 b . Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, ...

  18. [27]

    Wang, Y.; Zhong, W.; Li, L.; Mi, F.; Zeng, X.; Huang, W.; Shang, L.; Jiang, X.; and Liu, Q. 2023 c . Aligning Large Language Models with Human: A Survey. CoRR, abs/2307.12966

  19. [28]

    Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2023. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903

  20. [29]

    H.; Le, Q

    Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; and Oh, A., eds., Advances in Neu...

  21. [30]

    Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Tang, J.; Wang, J.; Yang, J.; Tu, J.; Zhang, J.; Ma, J.; Yang, J.; Xu, J.; Zhou, J.; Bai, J.; He, J.; Lin, J.; Dang, K.; Lu, K.; Chen, K.; Yang, K.; Li, M....

  22. [31]

    Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Information ...

  23. [32]

    Zeng, A.; Xu, B.; Wang, B.; Zhang, C.; Yin, D.; Rojas, D.; Feng, G.; Zhao, H.; Lai, H.; Yu, H.; Wang, H.; Sun, J.; Zhang, J.; Cheng, J.; Gui, J.; Tang, J.; Zhang, J.; Li, J.; Zhao, L.; Wu, L.; Zhong, L.; Liu, M.; Huang, M.; Zhang, P.; Zheng, Q.; Lu, R.; Duan, S.; Zhang, S.; Ca...

  24. [33]

    Zhang, D.; Zhoubian, S.; Yue, Y.; Dong, Y.; and Tang, J. 2024. ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search. CoRR, abs/2406.03816

  25. [34]

    Zhong, H.; Feng, G.; Xiong, W.; Zhao, L.; He, D.; Bian, J.; and Wang, L. 2024. DPO Meets PPO: Reinforced Token Optimization for RLHF . CoRR, abs/2404.18922

  26. [35]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  27. [36]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.