REVIEW 4 major objections 6 minor 35 references
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that process-supervised rewards should depend on both the accuracy and the length of a reasoning chain, and that the relationship is nonlinear, reporting that a method built on this principle outperforms mainstream…
desk verdict A sensible reward-shaping idea undermined by an ablation table that, as printed, contradicts the paper's main claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the combined reward: a geometric-mean accumulation function $F = (\prod_{j=1}^t R^s_j)^{1/t}$, where $R^s_j$ is the reward model's probability that step $j$ is positive, multiplied by an adjusted Weibull reward-shaping function $R_s = C \frac{k}{\lambda}(\frac{t}{\lambda})^{k-1} e^{-(t/\lambda)^k}$. The geometric mean removes the automatic penalty that plain multiplication imposes on long chains, and the Weibull factor injects the prior that chains of roughly $\lambda = 8$ steps deserve the largest shaping reward, with $k=1.5$ controlling how sharply the bonus falls off. The policy objective is $J(\pi) = E_\pi[R_s F - \beta D_{KL}(\pi \| \pi_{ref})]$. This machinery carries the paper's evidence: the ablation that removes $R_s$ degenerates to the standard product-accumulation method of Lightman et al. and loses performance, while the full method shifts the policy toward longer, higher-scoring chains.
What would settle it
Take the same PSPO-WRS pipeline and run it on problems whose correct solutions require very long chains, for example 15-25 steps, while also sweeping $\lambda$ over a grid from 4 to 16 with all else fixed: the nonlinearity claim predicts the method should keep beating the linear baseline across a range of $\lambda$, whereas the length-prior artifact predicts performance peaks only near $\lambda=8$ and collapses on long-chain problems.
Extended reading notes
Core claim
Process supervision, as usually implemented, scores a reasoning chain by multiplying the per-step correctness probabilities, which silently punishes longer chains even when every step is right. The paper's central discovery claim is that this multiplicative scheme is wrong in two ways: the overall reward should depend on the number of steps as well as their accuracy, and the dependence is nonlinear rather than linear. Its concrete proposal replaces the plain product with the geometric mean $F = (\prod_{j=1}^t P(y_j=1 \mid x, y_{<j}))^{1/t}$, which removes the length bias, and then multiplies $F$ by a Weibull-based shaping factor $R_s = C \frac{k}{\lambda}(\frac{t}{\lambda})^{k-1} e^{-(t/\lambda)^k}$ with $C=10.735$, $k=1.5$, $\lambda=8.0$, encoding a prior that chains of moderate length are best. In experiments on the six MATH-derived datasets, PSPO-WRS beats the Abel-7B baseline by between 6.56 and 30.94 percentage points and often matches or exceeds much larger models such as GPT-3.5 and Qwen2-72B; ablation shows the nonlinear module also raises the average number of generated steps from 2.624 to 3.051. The paper concludes that nonlinear, length-aware rewards are what make process supervision work.
Load-bearing premise
The load-bearing premise is that the hand-chosen Weibull curve, which gives the largest shaping reward to chains of about eight steps, correctly captures the true relationship between step count and reward quality; a single bad curve could erase the reported gains.
Editorial extensions
If this is right
- Process supervision pipelines should stop scoring chains with a raw product of step probabilities, because that construction systematically discourages longer reasoning.
- Step-count-normalized rewards plus a length prior can be added to any PRM-trained policy, not just the specific base model used here.
- Models trained this way generate more complete reasoning chains on average, which is the behavior the paper ties to higher accuracy.
- Because the gain persists on out-of-distribution datasets, the improvement is attributed to stronger reasoning rather than to memorizing the evaluation sets.
Reading between the lines
- The paper tests exactly one shaping curve, so its experiments do not yet separate 'nonlinearity helps' from 'this particular Weibull shape helps'; our inference is that a comparison against other nonlinear shapes, such as a log-length bonus or a bounded length cap, is needed to pin down the active ingredient.
- The chosen peak at eight steps is likely tuned to the difficulty of the six datasets; our inference is that portable versions of the method would need a procedure for setting $\lambda$ per task or per problem, otherwise the same method could hurt on puzzles whose correct chains are much shorter or much longer.
- The geometric mean already removes the main length penalty by itself; our inference is that an ablation keeping the geometric mean but dropping the Weibull factor would reveal how much of the gain is due to normalization versus the nonlinear prior.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes PSPO*, a paradigm for process-supervised policy optimization, and PSPO-WRS, a concrete instance that combines geometric-mean step accuracy with an adjusted Weibull reward-shaping function. The central claims are that process-supervision effectiveness depends on both accuracy and reasoning-chain length, that the reward score is nonlinear, and that PSPO-WRS consistently outperforms current models on six mathematical reasoning datasets. The paper includes an out-of-distribution generalization check and a code link.
Significance. If the empirical claims hold, the paper offers a useful systematization of process supervision and draws attention to length bias in multiplicative step rewards. The out-of-distribution experiments in Table 3 are a positive check against overfitting, and the promise of released code supports reproducibility. However, the sole ablation isolating the nonlinear-reward contribution is internally inconsistent as printed, and the fixed Weibull parameters are neither tuned nor subjected to sensitivity analysis, so the central mechanistic claim is currently not established.
major comments (4)
- [Table 4 / Ablation Analysis] As printed, Table 4's PSPO-WRS column is misaligned: the entries 64.91, 67.60, 71.57, 52.29, and 54.70 are the Table 2 values for NewsNLI, RedditNLI, RTE-Quant, StressTest, and QQA, respectively, not the values for the rows in which they appear. Reading the table at face value, the w/o-nonlinear baseline outperforms PSPO-WRS on AwpNLI (69.12 vs 64.91), RTE-Quant (59.66 vs 52.29), and QQA (54.94 vs 54.70), directly contradicting the text that nonlinear rewards improve performance across all evaluated datasets. StressTest also shows 53.57 in Table 4 versus 52.29 in Table 2. Since Table 4 is the only experiment that isolates the nonlinear-reward contribution, the main mechanistic claim is not reproducible from the manuscript as written; the authors must correct the table and re-state the conclusions accordingly.
- [Experimental Setups / Eq. (9)] The Weibull reward-shaping parameters C=10.735, k=1.5, and lambda=8.0 are fixed without any tuning procedure, grid search, or sensitivity analysis. Because Eq. (9) directly determines how the reward depends on step count, the observed gains could be artifacts of these specific values. Please report an ablation over these parameters or justify the choice from properties of the datasets.
- [Eqs. (9)-(10) / Ablation Analysis] The nonlinearity is introduced by construction: Eq. (10) length-normalizes the step-product reward and Eq. (9) applies a Weibull-shaped length prior. The comparison against a baseline without length shaping therefore tests the presence of a length-dependent reward, not the specific Weibull form or the broader claim that the underlying reward is nonlinear. I recommend comparing against alternative length reweighting schemes (e.g., linear normalization, logarithmic penalty, or a quadratic length model) and reporting whether the Weibull form gives materially better results.
- [Experimental Results / Metrics and Parameters] All experimental comparisons appear to be single runs with no error bars, confidence intervals, or multiple training seeds. PPO training is known to be high-variance, and some reported differences in Table 4 are small (e.g., QQA 54.94 vs 54.70). Please report means and variances over at least three seeds, or explicitly justify why seed variation is negligible.
minor comments (6)
- [Throughout] The dataset name is written inconsistently as "RTE-Quant" and "RTE Quant" across the text and tables; please standardize.
- [Introduction] The phrase "we concrete the PSPO* paradigm" should be "we instantiate" or "we concretize the PSPO* paradigm".
- [Eqs. (2)-(3)] The notation is overloaded: Eq. (2) writes p(zi) = sigma(zi), but sigma is later defined as softmax in Eq. (3), which is not a component-wise sigmoid; please define the activation consistently.
- [Figure 4] The caption contains the typo "Prameter settings" and should read "Parameter settings".
- [Table 4] The row "Num. of steps" is not defined in the caption; please state which policy's average step count is reported and over which evaluation set.
- [Introduction / Contribution 2] The claim "We are the first to assert that the reward score in the reasoning alignment is nonlinear" is stronger than what the related-work discussion supports; product-of-probabilities rewards are already nonlinear in step count, so prior work implicitly involves nonlinearity even if not framed that way.
Circularity Check
The claimed 'discovery' that process-supervision reward scores are nonlinear is built into the reward definition (Eqs. 9-11), and the ablation meant to validate it is internally inconsistent with Table 2.
-
self definitional
[Section 'PSPO-WRS: Process-supervised Policy Optimization with Nonlinear Reward Shaping', Eqs. (9)-(11); Ablation Analysis, Table 4]
"Additionally, we are the first to assert that the reward score in the reasoning alignment is nonlinear. ... Based on this prior knowledge, we employ the Adjusted Weibull distribution to shape the rewards ... Rs = C ∗ k λ ( t λ )k−1e−(t/λ)k , (9) ... F = [ tY j=1 P (yj = 1|xj, yj pre)]1/t. (10) ... Following the removal of the nonlinearity module, there is a noticeable decline in the performance of PSPO-WRS. ... These findings underscore the critical importance of nonlinear rewards in the process supervision."
The paper's headline claim that the process-supervision reward score is nonlinear is an input, not a derived result. Eq. (9) defines the shaping term Rs as a Weibull function of the step count t, and Eq. (10) defines F as a geometric mean of per-step probabilities; both are nonlinear functions, so the optimized objective RsF in Eq. (11) is nonlinear by construction. The ablation analysis then treats a comparison between this constructed reward and a version without the Weibull factor as empirical support for the 'hypothesis' that reward scores are nonlinear. No independent, parameter-free measurement of a ground-truth nonlinear relationship is presented, and the 'w/o nonlinear' baseline still contains the nonlinear geometric-mean accumulation F.
full rationale
The PSPO* paradigm itself and the PSPO-WRS instantiation are largely self-contained engineering contributions: reward-model training, PPO-style policy optimization, and head-to-head comparisons against external baselines do not depend on any load-bearing self-citation chain or on a fitted parameter being relabeled as a prediction. The circularity is concentrated in the conceptual claim of nonlinearity. Equations (9)-(11) make the final reward score a product of a Weibull function of step count and a geometric mean of step accuracies, so the statement that the reward score is nonlinear is true by definition; the later statement that experiments 'validate' the nonlinearity hypothesis is therefore a check of the authors' own design choice, not an independent discovery. The out-of-distribution results (Table 3) support the overall method but do not isolate the nonlinear contribution, so they cannot repair this definitional circularity. Additionally, Table 4 as printed contradicts Table 2 (e.g., the PSPO-WRS entries 64.91, 67.60, 71.57, 52.29, and 54.70 are Table 2's values for different datasets, and StressTest is shown as 53.57 instead of 52.29), so the only ablation that claims to validate the nonlinear module is not reproducible as written. These are correctness concerns as well, but they reinforce the conclusion that the central mechanistic claim currently stands on the definition rather than on solid evidence. A score of 6 reflects partial circularity: the key theoretical claim reduces to construction, while the empirical result that a length-aware Weibull-shaped reward can beat baselines retains independent content.
Assumptions & free parameters
free parameters (3)
- C =
10.735
- k =
1.5
- lambda =
8.0
assumptions (3)
- ad hoc to paper An adjusted Weibull distribution is an appropriate prior for how reasoning reward should depend on step count.
- ad hoc to paper The geometric-mean accumulation function is the right way to combine step scores while removing length bias.
- domain assumption Human step annotations are reliable and consistent across six datasets.
Cite this review
Pith. "Pith review of PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment." pith.science (2026). https://pith.science/paper/GKT5YPL2
@misc{pith2026241111681,
author = {Pith},
title = {Pith review of: PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/GKT5YPL2}},
note = {Machine review of arXiv:2411.11681}
}
read the original abstract
Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack of effective process supervision methods, even advanced large language models are prone to logical errors and redundant reasoning. We claim that the effectiveness of process supervision significantly depends on both the accuracy and the length of reasoning chains. Moreover, we identify that these factors exhibit a nonlinear relationship with the overall reward score of the reasoning process. Inspired by these insights, we propose a novel process supervision paradigm, PSPO*, which systematically outlines the workflow from reward model training to policy optimization, and highlights the importance of nonlinear rewards in process supervision. Based on PSPO*, we develop the PSPO-WRS, which considers the number of reasoning steps in determining reward scores and utilizes an adjusted Weibull distribution for nonlinear reward shaping. Experimental results on six mathematical reasoning datasets demonstrate that PSPO-WRS consistently outperforms current mainstream models.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Azar, M. G.; Guo, Z. D.; Piot, B.; Munos, R.; Rowland, M.; Valko, M.; and Calandriello, D. 2024. A General Theoretical Paradigm to Understand Learning from Human Preferences. In Dasgupta, S.; Mandt, S.; and Li, Y., eds., International Conference on Artificial Intelligence and Statistics, 2-4 May 2024, Palau de Congressos, Valencia, Spain, volume 238 of Pr...
work page 2024
-
[2]
Chen, C.-C.; Takamura, H.; Kobayashi, I.; and Miyao, Y. 2023. Improving Numeracy by Input Reframing and Quantitative Pre-Finetuning Task. In Vlachos, A.; and Augenstein, I., eds., Findings of the Association for Computational Linguistics: EACL 2023, 69--77. Dubrovnik, Croatia: Association for Computational Linguistics
work page 2023
-
[3]
Chern, E.; Zou, H.; Li, X.; Hu, J.; Feng, K.; Li, J.; and Liu, P. 2023. Generative AI for Math: Abel. https://github.com/GAIR-NLP/abel
2023
-
[4]
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve Math Word Problems. CoRR, abs/2110.14168
arXiv 2021
-
[5]
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapo...
2019
-
[6]
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; Goyal, A.; Hartshorn, A.; Yang, A.; Mitra, A.; Sravankumar, A.; Korenev, A.; Hinsvark, A.; Rao, A.; Zhang, A.; Rodriguez, A.; Gregerson, A.; Spataru, A.; Roziere, B.; Biron, B.; Tang, B.; Chern, B.; Caucheteux, C.; Nayak, C.; Bi, C.; Marra...
arXiv 2024
-
[7]
Gao, L.; Schulman, J.; and Hilton, J. 2023. Scaling Laws for Reward Model Overoptimization. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , volume 202 of Proceedings of Machine Learning Research, 10835--10866. PMLR
work page 2023
-
[8]
J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W
Hu, E. J.; Shen, Y.; Wallis, P.; Allen - Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
2022
Show all 35 references
-
[9]
Hu, Y.; Wang, W.; Jia, H.; Wang, Y.; Chen, Y.; Hao, J.; Wu, F.; and Fan, C. 2020. Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing System...
2020
-
[10]
Jin, M.; Yu, Q.; Shu, D.; Zhao, H.; Hua, W.; Meng, Y.; Zhang, Y.; and Du, M. 2024. The Impact of Reasoning Step Length on Large Language Models. CoRR, abs/2401.04925
2024 arXiv
-
[11]
S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2023. Large Language Models are Zero-Shot Reasoners. arXiv:2205.11916
2023 arXiv
-
[12]
Lai, X.; Tian, Z.; Chen, Y.; Yang, S.; Peng, X.; and Jia, J. 2024. Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs. CoRR, abs/2406.18629
2024 arXiv
-
[13]
Levine, S.; Popovic, Z.; and Koltun, V. 2011. Nonlinear Inverse Reinforcement Learning with Gaussian Processes. In Shawe - Taylor, J.; Zemel, R. S.; Bartlett, P. L.; Pereira, F. C. N.; and Weinberger, K. Q., eds., Advances in Neural Information Processing Systems 24: 25th Annu...
2011
-
[14]
Liang, X.; Li, J.; Yang, Y.; and Gao, Y. 2024. Bit \_ numeval at S em E val-2024 Task 7: Enhance Numerical Sensitivity and Reasoning Completeness for Quantitative Understanding. In Ojha, A. K.; Do g ru \"o z, A. S.; Tayyar Madabushi, H.; Da San Martino, G.; Rosenthal, S.; and ...
2024
-
[15]
Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023. Let's Verify Step by Step. CoRR, abs/2305.20050
2023 arXiv
-
[16]
Luo, H.; Sun, Q.; Xu, C.; Zhao, P.; Lou, J.; Tao, C.; Geng, X.; Lin, Q.; Chen, S.; and Zhang, D. 2023. WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct. CoRR, abs/2308.09583
2023 arXiv
-
[17]
Ma, Q.; Zhou, H.; Liu, T.; Yuan, J.; Liu, P.; You, Y.; and Yang, H. 2023. Let's reward step by step: Step-Level reward model as the Navigators for Reasoning. CoRR, abs/2310.10080
2023 arXiv
-
[18]
L.; Bari, M
Muennighoff, N.; Wang, T.; Sutawika, L.; Roberts, A.; Biderman, S.; Scao, T. L.; Bari, M. S.; Shen, S.; Yong, Z. X.; Schoelkopf, H.; Tang, X.; Radev, D.; Aji, A. F.; Almubarak, K.; Albanie, S.; Alyafeai, Z.; Webson, A.; Raff, E.; and Raffel, C. 2023. Crosslingual Generalizatio...
2023
-
[19]
L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P. F.; Leike, J.; and Lowe, R. 2022. Training language mo...
2022
-
[20]
D.; Ermon, S.; and Finn, C
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Informat...
2023
-
[21]
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal Policy Optimization Algorithms. CoRR, abs/1707.06347
2017 arXiv
-
[22]
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Canton - Ferrer, C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goy...
2023 arXiv
-
[24]
F.; Siegel, N
Uesato, J.; Kushman, N.; Kumar, R.; Song, H. F.; Siegel, N. Y.; Wang, L.; Creswell, A.; Irving, G.; and Higgins, I. 2022 b . Solving math word problems with process- and outcome-based feedback. CoRR, abs/2211.14275
2022 arXiv
-
[25]
X.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z
Wang, P.; Li, L.; Shao, Z.; Xu, R. X.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z. 2023 a . Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations. CoRR, abs/2312.08935
2023 arXiv
-
[26]
V.; Chi, E
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q. V.; Chi, E. H.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023 b . Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, ...
2023
-
[27]
Wang, Y.; Zhong, W.; Li, L.; Mi, F.; Zeng, X.; Huang, W.; Shang, L.; Jiang, X.; and Liu, Q. 2023 c . Aligning Large Language Models with Human: A Survey. CoRR, abs/2307.12966
2023 arXiv
-
[28]
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2023. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903
2023 arXiv
-
[29]
H.; Le, Q
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Koyejo, S.; Mohamed, S.; Agarwal, A.; Belgrave, D.; Cho, K.; and Oh, A., eds., Advances in Neu...
2022
-
[30]
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Tang, J.; Wang, J.; Yang, J.; Tu, J.; Zhang, J.; Ma, J.; Yang, J.; Xu, J.; Zhou, J.; Bai, J.; He, J.; Lin, J.; Dang, K.; Lu, K.; Chen, K.; Yang, K.; Li, M....
2024 arXiv
-
[31]
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Information ...
2023
-
[32]
Zeng, A.; Xu, B.; Wang, B.; Zhang, C.; Yin, D.; Rojas, D.; Feng, G.; Zhao, H.; Lai, H.; Yu, H.; Wang, H.; Sun, J.; Zhang, J.; Cheng, J.; Gui, J.; Tang, J.; Zhang, J.; Li, J.; Zhao, L.; Wu, L.; Zhong, L.; Liu, M.; Huang, M.; Zhang, P.; Zheng, Q.; Lu, R.; Duan, S.; Zhang, S.; Ca...
2024 arXiv
-
[33]
Zhang, D.; Zhoubian, S.; Yue, Y.; Dong, Y.; and Tang, J. 2024. ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search. CoRR, abs/2406.03816
2024 arXiv
-
[34]
Zhong, H.; Feng, G.; Xiong, W.; Zhao, L.; He, D.; Bian, J.; and Wang, L. 2024. DPO Meets PPO: Reinforced Token Optimization for RLHF . CoRR, abs/2404.18922
2024 arXiv
-
[35]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[36]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.