REVIEW 3 major objections 5 minor 43 references
A model can learn to keep or revise its own answers using only self-made verdicts and confidence, and that signal can allocate test-time compute better than fixed budgets.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 06:46 UTC pith:SG6WVWVP
load-bearing objection Solid adaptive test-time systems paper: joint verdict–confidence stopping works on the reported frontier, with real but openly measured stop-error limits. the 3 major comments →
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute: the policy emits a solution plus a discrete verdict and confidence, stops only when the verdict is Correct and confidence clears a threshold, and otherwise refines from its own self-check—without oracle feedback in prompts or at inference—yielding higher aggregate accuracy at far fewer turns than fixed-budget multi-turn baselines.
What carries the argument
Self-Verifying Refinement (SVR) with Joint Verdict–Confidence Reinforcement Learning: fixed-horizon multi-turn GRPO trajectories whose averaged return mixes solve progress, Brier-style calibration, overconfidence penalties, error detection, and stop-readiness, with gradients on the final completion only; at inference a Correct-and-confidence≥γ gate decides whether to keep the answer or continue.
Load-bearing premise
Training only on fixed-length trajectories, without ever rewarding fewer turns, still produces intermediate self-checks reliable enough that one global confidence threshold can safely stop early on many problems.
What would settle it
On the same seven math benchmarks and backbone, run the trained policy with the stated stop rule and check whether adaptive stopping still beats every fixed shared turn budget and the fixed-budget oracle-score baseline on macro accuracy at materially lower average turns; a collapse of that accuracy–compute gap, or PSE that erases the gains, would falsify the claim.
If this is right
- Test-time compute can be allocated turn-by-turn from the model’s own verdict–confidence state rather than from a pre-chosen budget or an external verifier.
- A single trained policy can support different accuracy–cost trade-offs by changing only the deployment confidence threshold, without retraining.
- History-conditioned self-refinement with learned stopping can match multi-sample majority voting accuracy at roughly half the token cost.
- Uniform extra refinement is not just wasteful: without an answer-retention signal it can destroy correct intermediate solutions.
Where Pith is reading between the lines
- If self-verification is the controller, calibration and asymmetric overconfidence become first-class systems problems, not post-hoc reporting niceties.
- The same stop interface could be tried outside contest math wherever a binary correctness check exists only in training (code tests, tool outcomes) but must not be exposed at deployment.
- High premature-stop error on hard sets suggests pairing the gate with a cheap secondary check or deferral action before treating early stop as final.
- Training that never prices realized tokens may still under-teach frugality; adding a mild inference-cost term could tighten the compute side without changing the oracle-free prompt boundary.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Self-Verifying Refinement (SVR), an oracle-free multi-turn RL framework in which a single policy emits a solution plus a discrete verdict and confidence, and uses that self-verification both to condition the next refinement prompt and, at inference only, to decide whether to stop. Training uses fixed-horizon GRPO trajectories with a trajectory-averaged return combining solve, calibration-aware verification (Brier, overconfidence, detection, stop-readiness), and format terms; ground-truth correctness enters only the reward, never refinement prompts or the inference controller. On seven math benchmarks with Qwen3.5-2B, adaptive SVR reports All-7 accuracy 0.563 at 2.99 mean turns, outperforming single-turn GRPO, several multi-turn baselines, and a fixed-budget oracle score-feedback reference while using far fewer tokens than fixed ten-turn inference, and matching Maj@10 GRPO at roughly half the tokens. Supporting analyses include same-policy fixed-budget curves, a global threshold sweep, reward/interface ablations, and three-seed stability.
Significance. If the results hold, the work is a concrete contribution to adaptive test-time compute: it shows that a policy-internal verdict–confidence interface can drive instance-dependent answer retention without an external verifier or separate controller at deployment. Strengths include a clear information boundary (oracle-free refinement), a complete-system baseline suite (Table 1), same-policy fixed-vs-adaptive isolation (Fig. 3 / Table 10), token-matched majority-vote comparison (Table 2), structured ablations (Tables 3–4), and multi-seed reporting (Appendix D.4). These make the accuracy–compute claim more falsifiable than typical multi-turn RL writeups. The practical significance is moderated by a single small backbone, high premature-stop error on harder sets, and a training–inference schedule mismatch that the paper does not fully close analytically.
major comments (3)
- [§3.3–3.4, Eqs. (7)–(10); §4.3] §3.3–3.4 and Eqs. (7)–(10): the central control claim depends on intermediate (v_t, c_t) being reliable enough for early stopping under γ=0.85, yet training forces T_tr=3, never runs the stop rule, and applies the GRPO loss only to final-completion tokens. Most deployment mass is early (All-7 mean 2.99 turns, ESR 86.3%; GSM8K 1.12 turns)—i.e., turns that never receive direct policy gradients. Trajectory averaging and context construction give indirect supervision, but the manuscript should quantify turn-level calibration/verdict quality (Brier, overconf., stop eligibility) at t=1,2 vs t=3 under the trained policy, or provide an ablation with multi-turn credit assignment / stop-on-policy training, so the reader can judge whether early-stop reliability is learned or largely inherited from easy domains and later-turn strength.
- [§4.1–4.2, Table 1] §4.2 / Table 1 and §4.1: the claim that SVR “exceeds … a fixed-budget oracle-guided score-feedback reference” is a complete-system comparison. The oracle baseline receives privileged y_{t−1} in prompts but has no policy-generated stopping signal and is evaluated at a fixed ten-turn budget, while SVR stops adaptively. This confounds feedback source with answer-retention policy and compute schedule. Either reframe the claim strictly as system-level (as the text sometimes does) and avoid implying superior self-verification quality versus oracle feedback, or add a matched controller (e.g., oracle score with the same stop rule, or SVR forced to K=10) so the contribution of internal verification versus adaptive retention is isolated.
- [§4.3, Fig. 4, Table 9] §4.3, Fig. 4, Table 9: Premature Stop Error remains 29.9% All-7 and 37.0% Math-5 (53.3% on MinervaMath) at the reported operating point, so a large fraction of examples commit early and wrong even while aggregate accuracy rises. Fixed-budget prefixes of the same policy top out at 0.450 All-7 versus adaptive 0.563, which supports retention value, but does not by itself show that the control signal is well calibrated on hard instances. The paper should break PSE by whether the trajectory ever contained a correct answer (harmful early stop vs never-solved), report conditional error among early stops (PSE/ESR), and discuss failure modes more prominently in the main text—not only as a future-work sentence—because this directly bounds the “effective internal control signal” claim.
minor comments (5)
- [Figure 1] Figure 1 caption and axis: “Avg. total tokens (×10^3)” is clear, but the main text sometimes switches between turns and tokens without restating that multi-turn prompt growth is included; a one-sentence reminder near Fig. 1 would help.
- [Table 1, §4.2] AIME26 (n=30) and AMC23 (n=40) drive some of the Math-5 narrative; the caution already in §4.2 should also appear near Table 1 when ranking methods on those columns.
- [§4.1; Appendix B.3] Missing-confidence priors (0.8/0.2/0.5 for C/I/U) in Appendix B.3 interact with γ=0.85 (default C without confidence cannot stop). Mention this briefly in the main inference protocol so readers do not treat γ as the only gate parameter.
- [§3.1, §3.4] Notation: T_tr vs T_max is introduced cleanly in §3.1 but occasionally “horizon” is used ambiguously in §3.4; keep the two symbols consistent in prose.
- [§2] Related work cites C3RL/CAS and CoRefine appropriately; a short explicit contrast on closed-loop refinement history vs independent sampling would sharpen positioning without lengthening much.
Circularity Check
No significant circularity: empirical RL method with explicit oracle-free boundary and external held-out evaluation.
full rationale
SVR is an empirical multi-turn RL paper, not a first-principles derivation. Ground-truth correctness enters only reward construction (Eqs. 7–9); refinement prompts and the inference controller never see labels (§3.1–3.2). Adaptive accuracy, turns, tokens, ESR, and PSE are measured on held-out benchmarks against external task evaluators (§4, App. C). The threshold γ and reward coefficients are declared hyperparameters, not fitted quantities re-presented as predictions. Fixed-horizon training with final-completion GRPO and adaptive inference are a design choice with acknowledged limitations (e.g., PSE), not a definitional identity that forces the reported All-7 accuracy. Related-work citations (GRPO, Reflexion, etc.) supply background methods, not load-bearing uniqueness theorems by the same authors. No step reduces the central accuracy–compute claim to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (6)
- deployment confidence threshold γ =
0.85
- self-verification reward coefficients (λ_cal, λ_over, λ_detect, λ_ready) =
0.5, 0.8, 0.2, 0.3
- solve/format reward coefficients (λ_abs, λ_Δ, α, λ_keep, λ_reg, λ_fail, λ_trunc, λ_fmt) =
1.0, 0.3, 0.5, 0.3, 0.5, 0.3, 0.5, 0.4
- training horizon T_tr and group size G =
T_tr=3, G=8
- inference budget T_max =
10
- missing-confidence priors (0.8/0.2/0.5 for C/I/U) =
0.8, 0.2, 0.5
axioms (5)
- domain assumption Group Relative Policy Optimization with clipped objectives is a valid multi-turn trainer when advantages are trajectory-level and gradients are applied only to final-completion tokens.
- domain assumption Binary task evaluators (Countdown expression check; GSM8K/MATH numeric/symbolic equivalence) are reliable ground truth for rewards and accuracy.
- ad hoc to paper A first-order Markov refinement context (problem + truncated previous draft + parsed self-check) is sufficient history for continued improvement and stopping.
- ad hoc to paper Forced fixed-horizon continuation during training yields better controller learning than on-policy early stopping under an immature verifier.
- domain assumption Standard transformer LM + RLVR setup (Qwen3.5-2B, ms-swift, vLLM rollouts) is an adequate testbed for general claims about self-verifying compute control.
invented entities (3)
-
Self-Verifying Refinement (SVR) controller
no independent evidence
-
Joint Verdict–Confidence Reinforcement Learning objective
no independent evidence
-
Stop-readiness reward term r_ready
no independent evidence
read the original abstract
Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues refinement using its own self-verification. Ground-truth correctness is used only to construct training rewards and is never exposed to the policy through refinement prompts or required at inference. SVR is trained with GRPO on fixed-horizon trajectories using rewards that promote solution correctness, calibration-aware self-verification, and stop-ready correct states; adaptive stopping is activated only at inference. On seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves a macro-average accuracy of 0.563 with only 2.99 inference turns on average. In the evaluated complete-system comparison, it exceeds standard GRPO, strong multi-turn baselines, and a fixed-budget oracle-guided score-feedback reference while requiring substantially fewer turns than fixed ten-turn inference. These results demonstrate that learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute allocation.
Figures
Reference graph
Works this paper leans on
-
[1]
Mohammad Ali Alomrani, Yingxue Zhang, Derek Li, Qianyi Sun, Soumyasundar Pal, Zhanguang Zhang, Yaochen Hu, Rohan Deepak Ajwani, Antonios Valkanas, Raika Karimi, Peng Cheng, Yunzhou Wang, Pengyi Liao, Hanrui Huang, Bin Wang, Jianye Hao, and Mark Coates. 2025. Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs.CoRRabs/250...
Pith/arXiv arXiv 2025
-
[2]
Glenn W. Brier. 1950. Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review78, 1 (1950), 1–3. doi:10.1175/1520-0493(1950)078<0001: VOFEIT>2.0.CO;2
-
[3]
Jiefeng Chen, Jie Ren, Xinyun Chen, Chengrun Yang, Ruoxi Sun, Jinsung Yoon, and Sercan Ö. Arik. 2025. SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling.Trans. Mach. Learn. Res.2025 (2025)
2025
-
[4]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training Verifiers to Solve Math Word Problems
2021
-
[5]
Mehul Damani, Idan Shenfeld, Andi Peng, Andreea Bobu, and Jacob Andreas. 2025. Learning How Hard to Think: Input-Adaptive Allocation of LM Computation. In ICLR. OpenReview.net
2025
-
[6]
DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.CoRRabs/2501.12948 (2025)
Pith/arXiv arXiv 2025
-
[7]
Jasper Dekoninck, Nikola Jovanovic, Tim Gehrunger, Kári Rögnvaldsson, Ivo Petrov, Chenhao Sun, and Martin T. Vechev. 2026. Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.CoRR abs/2605.00674 (2026)
Pith/arXiv arXiv 2026
-
[8]
Chanakya Ekbote, Vijay Lingam, Behrooz Omidvar-Tehrani, Jun Huan, Sujay Sanghavi, Anoop Deoras, and Stefano Soatto. 2025. MURPHY: Multi-Turn GRPO for Self Correcting Code Generation.CoRRabs/2511.07833 (2025)
Pith/arXiv arXiv 2025
-
[9]
Weinberger
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. InICML (Proceedings of Machine Learning Research, Vol. 70). PMLR, 1321–1330
2017
-
[10]
Ali Hatamizadeh, Shrimai Prabhumoye, Igor Gitman, Ximing Lu, Seungju Han, Wei Ping, Yejin Choi, and Jan Kautz. 2026. iGRPO: Self-Feedback-Driven LLM Reasoning.CoRRabs/2602.09000 (2026)
arXiv 2026
-
[11]
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024. OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems. InACL (1). Association for Computational Linguistics, 3828–3850
2024
-
[12]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring Mathematical Problem Solving With the MATH Dataset. InNeurIPS Datasets and Benchmarks
2021
-
[13]
Jingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang, Xiangyu Zhang, and Heung- Yeung Shum. 2025. Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model. InNeurIPS
2025
-
[14]
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How Can We KnowWhenLanguage Models Know? On the Calibration of Language Models for Question Answering.Trans. Assoc. Comput. Linguistics9 (2021), 962–977
2021
-
[15]
Chen Jin, Ryutaro Tanno, Tom Diethe, and Philip Teare. 2026. CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute.CoRR abs/2602.08948 (2026)
arXiv 2026
-
[16]
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran- Johnson, Scott Johnston, Sheer El Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec...
-
[17]
Ryo Kamoi, Yusen Zhang, Nan Zhang, Sarkar Snigdha Sarathi Das, and Rui Zhang
-
[18]
Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal M. P. Behbahani, and Aleksandra Faust. 2025. Training Language Models to Self-Correct via Reinforcement Learning. I...
2025
-
[19]
Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving Quantitative Reasoning Problems with Language Models. InNeurIPS
2022
-
[20]
Lianrui Li, Dakuan Lu, Jiawei Shao, Chi Zhang, and Xuelong Li. 2025. ScRPO: From Errors to Insights.CoRRabs/2511.06065 (2025)
arXiv 2025
-
[21]
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. Let’s Verify Step by Step. InICLR. OpenReview.net
2024
-
[22]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.CoRRabs/2402.03300 (2024)
Pith/arXiv arXiv 2024
-
[23]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: language agents with verbal reinforcement learning. InNeurIPS
2023
-
[24]
Utsav Singh, Sidhaarth Sredharan Murali, Souradip Chakraborty, Danush Khanna, Mubarak Shah, and Amrit Singh Bedi. 2026. Multi-Level Multi-Turn RL Outper- forms GRPO: Reasoning with Textual Feedback. InInternational Conference on Learning Representations
2026
-
[25]
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling LLM Test- Time Compute Optimally can be More Effective than Scaling Model Parameters. CoRRabs/2408.03314 (2024)
Pith/arXiv arXiv 2024
-
[26]
Paul Stangel, David Bani-Harouni, Chantal Pellegrini, Ege Özsoy, Kamilia Zaripova, Matthias Keicher, and Nassir Navab. 2025. Rewarding Doubt: A Re- inforcement Learning Approach to Calibrated Confidence Expression of Large Language Models. arXiv:2503.02623 [cs.CL]
arXiv 2025
-
[27]
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Na Zou, Hanjie Chen, and Xia Hu. 2025. Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.Trans. Mach. Learn. Res.2025 (2025)
2025
-
[28]
Liaoyaqi Wang, Chunsheng Zuo, William Jurayj, Benjamin Van Durme, and Anqi Liu. 2026. Process Supervision of Confidence Margin for Calibrated LLM Reasoning.CoRRabs/2604.23333 (2026)
Pith/arXiv arXiv 2026
-
[29]
Le, Ed H
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InICLR. OpenReview.net
2023
-
[30]
Yiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren, Lucas Liu, Baolin Peng, Hao Cheng, Xuehai He, Kuan Wang, Jianfeng Gao, Weizhu Chen, Shuohang Wang, Simon Shaolei Du, and Yelong Shen. 2025. Reinforcement Learning for Reasoning in Large Language Models with One Training Example. arXiv:2504.20571 [cs.LG]
Pith/arXiv arXiv 2025
-
[31]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. InNeurIPS
2022
-
[32]
Xuqing Yang, Yi Yuan, Shanzhe Lei, and Xuhong Wang. 2026. Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling. arXiv preprint arXiv:2607.01612(2026)
Pith/arXiv arXiv 2026
-
[33]
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. InNeurIPS
2023
-
[34]
Dongkeun Yoon, Seungone Kim, Sohee Yang, Sunkyoung Kim, Soyeon Kim, Yongil Kim, Eunbi Choi, Yireun Kim, and Minjoon Seo. 2025. Reasoning Models Better Express Their Confidence. InNeurIPS
2025
-
[35]
Yu Yue, Yufeng Yuan, Qiying Yu, Xiaochen Zuo, Ruofei Zhu, Wenyuan Xu, Jiaze Chen, Cheng-Xiang Wang, Tiantian Fan, Zhengyin Du, Xiangpeng Wei, Xiangyu Yu, Gaohong Liu, Juncai Liu, Lingjun Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Ru Zhang, Xin Liu, Mingxuan Wang, Yonghui Wu, and Lin Yan. 2025. VAPO: Efficient and Re...
Pith/arXiv arXiv 2025
-
[36]
Zhiyuan Zhai, Bingcong Li, Bingnan Xiao, Ming Li, and Xin Wang. 2026. Adap- tive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization.CoRRabs/2604.14853 (2026)
Pith/arXiv arXiv 2026
-
[37]
Yuze Zhao, Jintao Huang, Jinghan Hu, Xingjun Wang, Yunlin Mao, Daoze Zhang, Zeyinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, Wenmeng Zhou, and Yingda Chen. 2025. SWIFT: A Scalable Lightweight Infrastructure for Fine-Tuning. In AAAI. AAAI Press, 29733–29735
2025
-
[38]
Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, Jingren Zhou, and Junyang Lin. 2025. Group Sequence Policy Optimization.CoRRabs/2507.18071 (2025)
Pith/arXiv arXiv 2025
-
[39]
Shu Zhou, Rui Ling, Junan Chen, Xin Wang, Tao Fan, and Hao Wang. 2026. When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling. InACL (Findings). Association for Computational Linguistics, 23967–23977. Chen et al. A Implementation Details A.1 Training and Generation Hyperparameters All trainable methods use full-parameter reinforcement lea...
2026
-
[42]
Adaptive SVR stops at the first non-truncated turn satisfying 𝑣𝑡 = Cand 𝑐𝑡≥ 0.85, subject to a maximum inference budget of 𝑇max = 10. The maximum completion lengths used at evaluation are 800tokens for Countdown,1200for GSM8K,2048for MATH500 and AMC23,3072for AIME26 and MinervaMath, and4096for Olympiad- Bench. Fixed-budget multi-turn evaluations use the s...
-
[44]
Adaptive inference uses greedy decoding with𝛾= 0.85and 𝑇max = 10, so the observed variation primarily reflects training stochasticity rather than decoding randomness
All runs use identical training data, optimization hyperparam- eters, reward coefficients, prompt templates, and evaluation set- tings. Adaptive inference uses greedy decoding with𝛾= 0.85and 𝑇max = 10, so the observed variation primarily reflects training stochasticity rather than decoding randomness. Seed 42 is the check- point used in the main tables an...
-
[2022]
Language Models (Mostly) Know What They Know.CoRRabs/2207.05221 (2022)
Pith/arXiv arXiv 2022
-
[2025]
Training Step-Level Reasoning Verifiers with Formal Verification Tools. CoRRabs/2505.15960 (2025)
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.