Pith. sign in

REVIEW 3 major objections 5 minor 43 references

A model can learn to keep or revise its own answers using only self-made verdicts and confidence, and that signal can allocate test-time compute better than fixed budgets.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 06:46 UTC pith:SG6WVWVP

load-bearing objection Solid adaptive test-time systems paper: joint verdict–confidence stopping works on the reported frontier, with real but openly measured stop-error limits. the 3 major comments →

arxiv 2607.28457 v1 pith:SG6WVWVP submitted 2026-07-30 cs.AI cs.CL

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

classification cs.AI cs.CL
keywords Self-VerificationAdaptive Test-Time ComputeConfidence CalibrationMulti-turn Reinforcement LearningGRPOAnswer RetentionMathematical Reasoning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Uniform extra thinking wastes tokens on easy problems and can overwrite answers that were already right. External verifiers can guide refinement, but they may be costly or unavailable at deployment. This paper argues that a single policy can both refine solutions and produce a structured self-check—Correct/Incorrect/Unsure plus a confidence score—that decides whether to stop or continue, without ever seeing ground-truth labels in the refinement loop. Training uses fixed-length multi-turn trajectories and a trajectory-level reward that rewards correct solutions, calibrated confidence, error recognition, and stop-ready correct states; early stopping is turned on only at inference. On seven math benchmarks with a 2B model, the method reaches higher macro accuracy than strong multi-turn baselines and a fixed-budget oracle-feedback reference while using about three turns on average instead of ten. The practical point is that internal self-verification, if trained as a control signal, can retain the right intermediate answer and spend compute where it still helps.

Core claim

Learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute: the policy emits a solution plus a discrete verdict and confidence, stops only when the verdict is Correct and confidence clears a threshold, and otherwise refines from its own self-check—without oracle feedback in prompts or at inference—yielding higher aggregate accuracy at far fewer turns than fixed-budget multi-turn baselines.

What carries the argument

Self-Verifying Refinement (SVR) with Joint Verdict–Confidence Reinforcement Learning: fixed-horizon multi-turn GRPO trajectories whose averaged return mixes solve progress, Brier-style calibration, overconfidence penalties, error detection, and stop-readiness, with gradients on the final completion only; at inference a Correct-and-confidence≥γ gate decides whether to keep the answer or continue.

Load-bearing premise

Training only on fixed-length trajectories, without ever rewarding fewer turns, still produces intermediate self-checks reliable enough that one global confidence threshold can safely stop early on many problems.

What would settle it

On the same seven math benchmarks and backbone, run the trained policy with the stated stop rule and check whether adaptive stopping still beats every fixed shared turn budget and the fixed-budget oracle-score baseline on macro accuracy at materially lower average turns; a collapse of that accuracy–compute gap, or PSE that erases the gains, would falsify the claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Test-time compute can be allocated turn-by-turn from the model’s own verdict–confidence state rather than from a pre-chosen budget or an external verifier.
  • A single trained policy can support different accuracy–cost trade-offs by changing only the deployment confidence threshold, without retraining.
  • History-conditioned self-refinement with learned stopping can match multi-sample majority voting accuracy at roughly half the token cost.
  • Uniform extra refinement is not just wasteful: without an answer-retention signal it can destroy correct intermediate solutions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If self-verification is the controller, calibration and asymmetric overconfidence become first-class systems problems, not post-hoc reporting niceties.
  • The same stop interface could be tried outside contest math wherever a binary correctness check exists only in training (code tests, tool outcomes) but must not be exposed at deployment.
  • High premature-stop error on hard sets suggests pairing the gate with a cheap secondary check or deferral action before treating early stop as final.
  • Training that never prices realized tokens may still under-teach frugality; adding a mild inference-cost term could tighten the compute side without changing the oracle-free prompt boundary.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Self-Verifying Refinement (SVR), an oracle-free multi-turn RL framework in which a single policy emits a solution plus a discrete verdict and confidence, and uses that self-verification both to condition the next refinement prompt and, at inference only, to decide whether to stop. Training uses fixed-horizon GRPO trajectories with a trajectory-averaged return combining solve, calibration-aware verification (Brier, overconfidence, detection, stop-readiness), and format terms; ground-truth correctness enters only the reward, never refinement prompts or the inference controller. On seven math benchmarks with Qwen3.5-2B, adaptive SVR reports All-7 accuracy 0.563 at 2.99 mean turns, outperforming single-turn GRPO, several multi-turn baselines, and a fixed-budget oracle score-feedback reference while using far fewer tokens than fixed ten-turn inference, and matching Maj@10 GRPO at roughly half the tokens. Supporting analyses include same-policy fixed-budget curves, a global threshold sweep, reward/interface ablations, and three-seed stability.

Significance. If the results hold, the work is a concrete contribution to adaptive test-time compute: it shows that a policy-internal verdict–confidence interface can drive instance-dependent answer retention without an external verifier or separate controller at deployment. Strengths include a clear information boundary (oracle-free refinement), a complete-system baseline suite (Table 1), same-policy fixed-vs-adaptive isolation (Fig. 3 / Table 10), token-matched majority-vote comparison (Table 2), structured ablations (Tables 3–4), and multi-seed reporting (Appendix D.4). These make the accuracy–compute claim more falsifiable than typical multi-turn RL writeups. The practical significance is moderated by a single small backbone, high premature-stop error on harder sets, and a training–inference schedule mismatch that the paper does not fully close analytically.

major comments (3)
  1. [§3.3–3.4, Eqs. (7)–(10); §4.3] §3.3–3.4 and Eqs. (7)–(10): the central control claim depends on intermediate (v_t, c_t) being reliable enough for early stopping under γ=0.85, yet training forces T_tr=3, never runs the stop rule, and applies the GRPO loss only to final-completion tokens. Most deployment mass is early (All-7 mean 2.99 turns, ESR 86.3%; GSM8K 1.12 turns)—i.e., turns that never receive direct policy gradients. Trajectory averaging and context construction give indirect supervision, but the manuscript should quantify turn-level calibration/verdict quality (Brier, overconf., stop eligibility) at t=1,2 vs t=3 under the trained policy, or provide an ablation with multi-turn credit assignment / stop-on-policy training, so the reader can judge whether early-stop reliability is learned or largely inherited from easy domains and later-turn strength.
  2. [§4.1–4.2, Table 1] §4.2 / Table 1 and §4.1: the claim that SVR “exceeds … a fixed-budget oracle-guided score-feedback reference” is a complete-system comparison. The oracle baseline receives privileged y_{t−1} in prompts but has no policy-generated stopping signal and is evaluated at a fixed ten-turn budget, while SVR stops adaptively. This confounds feedback source with answer-retention policy and compute schedule. Either reframe the claim strictly as system-level (as the text sometimes does) and avoid implying superior self-verification quality versus oracle feedback, or add a matched controller (e.g., oracle score with the same stop rule, or SVR forced to K=10) so the contribution of internal verification versus adaptive retention is isolated.
  3. [§4.3, Fig. 4, Table 9] §4.3, Fig. 4, Table 9: Premature Stop Error remains 29.9% All-7 and 37.0% Math-5 (53.3% on MinervaMath) at the reported operating point, so a large fraction of examples commit early and wrong even while aggregate accuracy rises. Fixed-budget prefixes of the same policy top out at 0.450 All-7 versus adaptive 0.563, which supports retention value, but does not by itself show that the control signal is well calibrated on hard instances. The paper should break PSE by whether the trajectory ever contained a correct answer (harmful early stop vs never-solved), report conditional error among early stops (PSE/ESR), and discuss failure modes more prominently in the main text—not only as a future-work sentence—because this directly bounds the “effective internal control signal” claim.
minor comments (5)
  1. [Figure 1] Figure 1 caption and axis: “Avg. total tokens (×10^3)” is clear, but the main text sometimes switches between turns and tokens without restating that multi-turn prompt growth is included; a one-sentence reminder near Fig. 1 would help.
  2. [Table 1, §4.2] AIME26 (n=30) and AMC23 (n=40) drive some of the Math-5 narrative; the caution already in §4.2 should also appear near Table 1 when ranking methods on those columns.
  3. [§4.1; Appendix B.3] Missing-confidence priors (0.8/0.2/0.5 for C/I/U) in Appendix B.3 interact with γ=0.85 (default C without confidence cannot stop). Mention this briefly in the main inference protocol so readers do not treat γ as the only gate parameter.
  4. [§3.1, §3.4] Notation: T_tr vs T_max is introduced cleanly in §3.1 but occasionally “horizon” is used ambiguously in §3.4; keep the two symbols consistent in prose.
  5. [§2] Related work cites C3RL/CAS and CoRefine appropriately; a short explicit contrast on closed-loop refinement history vs independent sampling would sharpen positioning without lengthening much.

Circularity Check

0 steps flagged

No significant circularity: empirical RL method with explicit oracle-free boundary and external held-out evaluation.

full rationale

SVR is an empirical multi-turn RL paper, not a first-principles derivation. Ground-truth correctness enters only reward construction (Eqs. 7–9); refinement prompts and the inference controller never see labels (§3.1–3.2). Adaptive accuracy, turns, tokens, ESR, and PSE are measured on held-out benchmarks against external task evaluators (§4, App. C). The threshold γ and reward coefficients are declared hyperparameters, not fitted quantities re-presented as predictions. Fixed-horizon training with final-completion GRPO and adaptive inference are a design choice with acknowledged limitations (e.g., PSE), not a definitional identity that forces the reported All-7 accuracy. Related-work citations (GRPO, Reflexion, etc.) supply background methods, not load-bearing uniqueness theorems by the same authors. No step reduces the central accuracy–compute claim to its inputs by construction.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

Load-bearing content is methodological and empirical, not axiomatic physics. The claim rests on standard RL/LM assumptions, a specific reward factorization and optimization restriction (final-turn gradients only), hand-chosen control hyperparameters, and the modeling choice that a policy-generated verdict–confidence pair is a sufficient Markov control state for stop/refine.

free parameters (6)
  • deployment confidence threshold γ = 0.85
    Single global stop threshold used for all main results; diagnostic sweep picks 0.85 as best aggregate point but it is not learned.
  • self-verification reward coefficients (λ_cal, λ_over, λ_detect, λ_ready) = 0.5, 0.8, 0.2, 0.3
    Hand-set weights that shape calibration, overconfidence penalty, error detection, and stop-readiness; ablations show the central claim depends on this mix.
  • solve/format reward coefficients (λ_abs, λ_Δ, α, λ_keep, λ_reg, λ_fail, λ_trunc, λ_fmt) = 1.0, 0.3, 0.5, 0.3, 0.5, 0.3, 0.5, 0.4
    Hand-tuned trajectory shaping terms that define what ‘good refinement’ means during GRPO.
  • training horizon T_tr and group size G = T_tr=3, G=8
    Fixed forced continuation depth and number of trajectories per input; determine state coverage and optimization sample count.
  • inference budget T_max = 10
    Hard cap on adaptive refinement; interacts with γ to set the accuracy–compute operating point.
  • missing-confidence priors (0.8/0.2/0.5 for C/I/U) = 0.8, 0.2, 0.5
    Parser fallback confidences that affect whether incomplete self-checks can stop; chosen below γ for Correct.
axioms (5)
  • domain assumption Group Relative Policy Optimization with clipped objectives is a valid multi-turn trainer when advantages are trajectory-level and gradients are applied only to final-completion tokens.
    §3.4 adopts GRPO and the final-turn-only loss without proving credit assignment optimality for intermediate self-checks.
  • domain assumption Binary task evaluators (Countdown expression check; GSM8K/MATH numeric/symbolic equivalence) are reliable ground truth for rewards and accuracy.
    §C.2; all y_t and reported accuracies inherit evaluator errors or parse failures as incorrect.
  • ad hoc to paper A first-order Markov refinement context (problem + truncated previous draft + parsed self-check) is sufficient history for continued improvement and stopping.
    Eq. (6) and §3.2 deliberately discard full trajectory history.
  • ad hoc to paper Forced fixed-horizon continuation during training yields better controller learning than on-policy early stopping under an immature verifier.
    §3.4 motivation; not independently validated against train-time adaptive schedules.
  • domain assumption Standard transformer LM + RLVR setup (Qwen3.5-2B, ms-swift, vLLM rollouts) is an adequate testbed for general claims about self-verifying compute control.
    All experiments use one 2B backbone and math-only domains (§4.1).
invented entities (3)
  • Self-Verifying Refinement (SVR) controller no independent evidence
    purpose: Unify solution generation and stop/refine decisions in one policy via structured self-checks.
    Named system-level construct; evaluated only inside this paper’s training and benchmarks.
  • Joint Verdict–Confidence Reinforcement Learning objective no independent evidence
    purpose: Trajectory return combining solve, Brier-style calibration, overconfidence, detection, stop-readiness, and format terms.
    Paper-specific reward aggregation (Eqs. 7–9); existence justified by ablations, not external theory.
  • Stop-readiness reward term r_ready no independent evidence
    purpose: Encourage correct answers to carry high confidence Correct verdicts so the inference gate can fire.
    Introduced in Eq. (9); ablation w/o R_ready sharply hurts All-7 accuracy and raises turns.

pith-pipeline@v1.2.0-daily-grok45 · 41278 in / 4267 out tokens · 86126 ms · 2026-07-31T06:46:05.549587+00:00 · methodology

0 comments
read the original abstract

Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues refinement using its own self-verification. Ground-truth correctness is used only to construct training rewards and is never exposed to the policy through refinement prompts or required at inference. SVR is trained with GRPO on fixed-horizon trajectories using rewards that promote solution correctness, calibration-aware self-verification, and stop-ready correct states; adaptive stopping is activated only at inference. On seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves a macro-average accuracy of 0.563 with only 2.99 inference turns on average. In the evaluated complete-system comparison, it exceeds standard GRPO, strong multi-turn baselines, and a fixed-budget oracle-guided score-feedback reference while requiring substantially fewer turns than fixed ten-turn inference. These results demonstrate that learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute allocation.

Figures

Figures reproduced from arXiv: 2607.28457 by Guangrun Wang, Hongyu Chen, Liang Lin.

Figure 1
Figure 1. Figure 1: Accuracy and cumulative inference tokens on All-7 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of fixed-horizon training and adaptive inference in SVR. (a) Per-turn solve, self-verification, and format [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: Sensitivity of adaptive SVR to the confidence thresh [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 3
Figure 3. Figure 3: Fixed-budget and adaptive inference using the same [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

43 extracted references · 14 linked inside Pith

  1. [1]

    Mohammad Ali Alomrani, Yingxue Zhang, Derek Li, Qianyi Sun, Soumyasundar Pal, Zhanguang Zhang, Yaochen Hu, Rohan Deepak Ajwani, Antonios Valkanas, Raika Karimi, Peng Cheng, Yunzhou Wang, Pengyi Liao, Hanrui Huang, Bin Wang, Jianye Hao, and Mark Coates. 2025. Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs.CoRRabs/250...

  2. [2]

    Glenn W. Brier. 1950. Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review78, 1 (1950), 1–3. doi:10.1175/1520-0493(1950)078<0001: VOFEIT>2.0.CO;2

  3. [3]

    Jiefeng Chen, Jie Ren, Xinyun Chen, Chengrun Yang, Ruoxi Sun, Jinsung Yoon, and Sercan Ö. Arik. 2025. SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling.Trans. Mach. Learn. Res.2025 (2025)

  4. [4]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training Verifiers to Solve Math Word Problems

  5. [5]

    Mehul Damani, Idan Shenfeld, Andi Peng, Andreea Bobu, and Jacob Andreas. 2025. Learning How Hard to Think: Input-Adaptive Allocation of LM Computation. In ICLR. OpenReview.net

  6. [6]

    DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.CoRRabs/2501.12948 (2025)

  7. [7]

    Jasper Dekoninck, Nikola Jovanovic, Tim Gehrunger, Kári Rögnvaldsson, Ivo Petrov, Chenhao Sun, and Martin T. Vechev. 2026. Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.CoRR abs/2605.00674 (2026)

  8. [8]

    Chanakya Ekbote, Vijay Lingam, Behrooz Omidvar-Tehrani, Jun Huan, Sujay Sanghavi, Anoop Deoras, and Stefano Soatto. 2025. MURPHY: Multi-Turn GRPO for Self Correcting Code Generation.CoRRabs/2511.07833 (2025)

  9. [9]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. InICML (Proceedings of Machine Learning Research, Vol. 70). PMLR, 1321–1330

  10. [10]

    Ali Hatamizadeh, Shrimai Prabhumoye, Igor Gitman, Ximing Lu, Seungju Han, Wei Ping, Yejin Choi, and Jan Kautz. 2026. iGRPO: Self-Feedback-Driven LLM Reasoning.CoRRabs/2602.09000 (2026)

  11. [11]

    Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024. OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems. InACL (1). Association for Computational Linguistics, 3828–3850

  12. [12]

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring Mathematical Problem Solving With the MATH Dataset. InNeurIPS Datasets and Benchmarks

  13. [13]

    Jingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang, Xiangyu Zhang, and Heung- Yeung Shum. 2025. Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model. InNeurIPS

  14. [14]

    Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How Can We KnowWhenLanguage Models Know? On the Calibration of Language Models for Question Answering.Trans. Assoc. Comput. Linguistics9 (2021), 962–977

  15. [15]

    Chen Jin, Ryutaro Tanno, Tom Diethe, and Philip Teare. 2026. CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute.CoRR abs/2602.08948 (2026)

  16. [16]

    Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran- Johnson, Scott Johnston, Sheer El Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec...

  17. [17]

    Ryo Kamoi, Yusen Zhang, Nan Zhang, Sarkar Snigdha Sarathi Das, and Rui Zhang

  18. [18]

    Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M

    Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal M. P. Behbahani, and Aleksandra Faust. 2025. Training Language Models to Self-Correct via Reinforcement Learning. I...

  19. [19]

    Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

    Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving Quantitative Reasoning Problems with Language Models. InNeurIPS

  20. [20]

    Lianrui Li, Dakuan Lu, Jiawei Shao, Chi Zhang, and Xuelong Li. 2025. ScRPO: From Errors to Insights.CoRRabs/2511.06065 (2025)

  21. [21]

    Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. Let’s Verify Step by Step. InICLR. OpenReview.net

  22. [22]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.CoRRabs/2402.03300 (2024)

  23. [23]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: language agents with verbal reinforcement learning. InNeurIPS

  24. [24]

    Utsav Singh, Sidhaarth Sredharan Murali, Souradip Chakraborty, Danush Khanna, Mubarak Shah, and Amrit Singh Bedi. 2026. Multi-Level Multi-Turn RL Outper- forms GRPO: Reasoning with Textual Feedback. InInternational Conference on Learning Representations

  25. [25]

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling LLM Test- Time Compute Optimally can be More Effective than Scaling Model Parameters. CoRRabs/2408.03314 (2024)

  26. [26]

    Paul Stangel, David Bani-Harouni, Chantal Pellegrini, Ege Özsoy, Kamilia Zaripova, Matthias Keicher, and Nassir Navab. 2025. Rewarding Doubt: A Re- inforcement Learning Approach to Calibrated Confidence Expression of Large Language Models. arXiv:2503.02623 [cs.CL]

  27. [27]

    Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Na Zou, Hanjie Chen, and Xia Hu. 2025. Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.Trans. Mach. Learn. Res.2025 (2025)

  28. [28]

    Liaoyaqi Wang, Chunsheng Zuo, William Jurayj, Benjamin Van Durme, and Anqi Liu. 2026. Process Supervision of Confidence Margin for Calibrated LLM Reasoning.CoRRabs/2604.23333 (2026)

  29. [29]

    Le, Ed H

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InICLR. OpenReview.net

  30. [30]

    Yiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren, Lucas Liu, Baolin Peng, Hao Cheng, Xuehai He, Kuan Wang, Jianfeng Gao, Weizhu Chen, Shuohang Wang, Simon Shaolei Du, and Yelong Shen. 2025. Reinforcement Learning for Reasoning in Large Language Models with One Training Example. arXiv:2504.20571 [cs.LG]

  31. [31]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. InNeurIPS

  32. [32]

    Xuqing Yang, Yi Yuan, Shanzhe Lei, and Xuhong Wang. 2026. Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling. arXiv preprint arXiv:2607.01612(2026)

  33. [33]

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. InNeurIPS

  34. [34]

    Dongkeun Yoon, Seungone Kim, Sohee Yang, Sunkyoung Kim, Soyeon Kim, Yongil Kim, Eunbi Choi, Yireun Kim, and Minjoon Seo. 2025. Reasoning Models Better Express Their Confidence. InNeurIPS

  35. [35]

    Yu Yue, Yufeng Yuan, Qiying Yu, Xiaochen Zuo, Ruofei Zhu, Wenyuan Xu, Jiaze Chen, Cheng-Xiang Wang, Tiantian Fan, Zhengyin Du, Xiangpeng Wei, Xiangyu Yu, Gaohong Liu, Juncai Liu, Lingjun Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Ru Zhang, Xin Liu, Mingxuan Wang, Yonghui Wu, and Lin Yan. 2025. VAPO: Efficient and Re...

  36. [36]

    Zhiyuan Zhai, Bingcong Li, Bingnan Xiao, Ming Li, and Xin Wang. 2026. Adap- tive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization.CoRRabs/2604.14853 (2026)

  37. [37]

    Yuze Zhao, Jintao Huang, Jinghan Hu, Xingjun Wang, Yunlin Mao, Daoze Zhang, Zeyinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, Wenmeng Zhou, and Yingda Chen. 2025. SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning. In AAAI. AAAI Press, 29733–29735

  38. [38]

    Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, Jingren Zhou, and Junyang Lin. 2025. Group Sequence Policy Optimization.CoRRabs/2507.18071 (2025)

  39. [39]

    Shu Zhou, Rui Ling, Junan Chen, Xin Wang, Tao Fan, and Hao Wang. 2026. When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling. InACL (Findings). Association for Computational Linguistics, 23967–23977. Chen et al. A Implementation Details A.1 Training and Generation Hyperparameters All trainable methods use full-parameter reinforcement lea...

  40. [42]

    The maximum completion lengths used at evaluation are 800tokens for Countdown,1200for GSM8K,2048for MATH500 and AMC23,3072for AIME26 and MinervaMath, and4096for Olympiad- Bench

    Adaptive SVR stops at the first non-truncated turn satisfying 𝑣𝑡 = Cand 𝑐𝑡≥ 0.85, subject to a maximum inference budget of 𝑇max = 10. The maximum completion lengths used at evaluation are 800tokens for Countdown,1200for GSM8K,2048for MATH500 and AMC23,3072for AIME26 and MinervaMath, and4096for Olympiad- Bench. Fixed-budget multi-turn evaluations use the s...

  41. [44]

    Adaptive inference uses greedy decoding with𝛾= 0.85and 𝑇max = 10, so the observed variation primarily reflects training stochasticity rather than decoding randomness

    All runs use identical training data, optimization hyperparam- eters, reward coefficients, prompt templates, and evaluation set- tings. Adaptive inference uses greedy decoding with𝛾= 0.85and 𝑇max = 10, so the observed variation primarily reflects training stochasticity rather than decoding randomness. Seed 42 is the check- point used in the main tables an...

  42. [2022]

    Language Models (Mostly) Know What They Know.CoRRabs/2207.05221 (2022)

  43. [2025]

    CoRRabs/2505.15960 (2025)

    Training Step-Level Reasoning Verifiers with Formal Verification Tools. CoRRabs/2505.15960 (2025)