Pith. sign in

REVIEW 3 major objections 5 minor 34 references

OS-Pruner frames chain-of-thought pruning as optimal stopping and cuts generation length by 20–60% with little accuracy loss.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 07:07 UTC pith:FKSEIDMH

load-bearing objection Clean optimal-stopping framing for CoT early exit with a real theorem and solid multi-model Pareto curves, but the headline 20–60% cuts live under forced boxed-answer evaluation that free-generation ablations already weaken. the 3 major comments →

arxiv 2607.11089 v1 pith:FKSEIDMH submitted 2026-07-13 cs.AI

OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping

classification cs.AI
keywords chain-of-thoughtoptimal stoppingearly stoppingreasoning efficiencylength reductionplug-in policylarge language models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Reasoning models often keep writing after the answer is already determined, burning tokens without raising accuracy. This paper treats the decision to stop as an optimal stopping problem: after each reasoning step, continue only if the expected accuracy gain is worth the extra length cost. A small plug-in policy is trained on precomputed rewards of the form accuracy-minus-λ-times-tokens, leaving the base model frozen. Across math benchmarks and several reasoning models, the method sits on or near the accuracy–efficiency frontier and beats fixed-threshold early-exit classifiers, while λ gives users direct control of the trade-off.

Core claim

Chain-of-thought pruning is best cast as optimal stopping with reward equal to final-answer accuracy of a prefix minus a length penalty λ. A lightweight policy head trained on that objective, without retraining the base reasoner, yields stopping rules that dominate classification-style early exit on the accuracy–length frontier and deliver roughly 20–60% shorter generations at low accuracy cost.

What carries the argument

The stopping utility r = A(prefix) − λ L(prefix) and the Bellman comparison of immediate accuracy value versus continuation value; a linear head on the last hidden state (with light fine-tuning of the last two layers) learns the stop probability at paragraph boundaries.

Load-bearing premise

Forcing a boxed final answer after every intermediate reasoning prefix must give a faithful accuracy signal that the model does not systematically undo by continuing to think in the answer section.

What would settle it

On the same held-out math sets, run free (unforced) final-answer decoding: if OS-Pruner’s length reductions vanish or accuracy falls well below full traces and fixed-threshold baselines, the optimal-stopping claim fails under realistic decoding.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Sweeping λ traces a controllable accuracy–efficiency frontier without changing the frozen base model.
  • Even models already shortened by model-side compression still overthink and can be further pruned by the same plug-in.
  • Fixed-threshold correctness classifiers can lose arbitrarily large value relative to optimal stopping (Theorem 1), so early-exit training should compare continuation value, not only a confidence cutoff.
  • Easier and medium problems admit large token savings while pass@1 stays nearly intact; hard olympiad problems stay more conservative.
  • Inference cost and latency of reasoning models can drop substantially when the policy is invoked only at paragraph boundaries.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same prefix-level reward pipeline could transfer to coding or tool-use traces whenever an intermediate correctness oracle exists.
  • Progressive annealing of λ is a practical fix for “always continue” local minima and may apply to other sparse sequential decisions.
  • On instruction-weak models, adding training-free entropy or redundancy signals as policy features could stabilize frontiers when forced-answer labels are noisy.
  • Serving systems that already prefill whole paragraphs can host the two-layer policy head with negligible extra latency.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper formulates chain-of-thought pruning as an optimal stopping problem. After each reasoning step the policy decides whether to stop and emit a final answer or continue, maximizing the expected reward r = A(y≤i|x) − λ L(y≤i). A lightweight linear head on the last hidden state (with only the final two transformer layers lightly fine-tuned) is trained offline from precomputed prefix rewards obtained by forcibly terminating each paragraph, boxing the answer, and grading with Math-Verify. Theorem 1 shows that any fixed-threshold correctness classifier can be arbitrarily suboptimal relative to the optimal stopping policy. Empirically, on DeepSeek-R1-Distill-Qwen-7B, GPT-OSS-20B and DRPO-7B across GSM8K, MATH-500 and AIME, OS-Pruner reports 20–60 % length reductions at low accuracy cost and lies on or near the accuracy–efficiency Pareto frontier relative to matched-architecture classification, answer-convergence and FlashThink baselines (Fig. 2, Table 1).

Significance. If the claims hold under realistic generation, the work supplies a clean, theoretically grounded alternative to heuristic early-exit and expensive model-side compression. The explicit λ-controlled utility, the offline reward construction that avoids on-policy rollouts, and the plug-in architecture are practical strengths. Theorem 1 is a short, correct finite-horizon construction that cleanly separates optimal stopping from fixed-threshold classification. The DRPO-7B experiment further shows that the method can still add value on top of a strong model-side compressor. These elements make the paper a useful contribution to efficient reasoning provided the evaluation protocol is shown to be faithful.

major comments (3)
  1. §5.1 and the main experimental protocol force a boxed final answer after every prefix both for training rewards and for the reported tables/figures. Appendix A (and Tables 5–7) demonstrates that when models are allowed free thinking in the answer section, DeepSeek-R1-Distill-Qwen-7B and DRPO-7B systematically compensate: total (CoT+Ans) length reductions shrink, frontiers become less controllable, and baseline accuracies rise. Consequently the headline 20–60 % reductions and AES dominance (Table 1, Fig. 2) are measured under a non-default generation regime. Either primary results under free generation must be elevated, or a stronger argument must be given that forced boxing is the intended deployment setting and that the free-thinking compensation does not reverse the ranking versus baselines.
  2. AES (§6, Eq. 8) raises the accuracy-penalty coefficient from the literature value γ=3 to γ=7. Because best-AES operating points are used to select the reported policies in Table 1, the ranking of methods (especially on AIME and on DRPO-7B) can change under the original coefficients. A short sensitivity table with the literature (α,β,γ) would confirm that the claimed Pareto improvement is not an artifact of the altered metric.
  3. Progressive λ annealing (§5.3, Fig. 1) is essential to escape the trivial “never-stop” local minimum for small λ, yet the schedule (initial value, annealing points, number of steps) is described only qualitatively. Without a precise schedule or an ablation that starts from a cold small-λ initialization, it is hard to judge how much of the reported frontier is due to the optimal-stopping objective versus the curriculum.
minor comments (5)
  1. Notation inconsistency: Eq. (3) writes R_θ(y|x) while Eq. (4) writes R_θ(x,y); unify.
  2. Fig. 2 captions and axis labels would benefit from explicit units (tokens) and a note that length is measured only up to the stop decision under forced answering.
  3. The claim “lightweight during both training and inference” should be qualified by the multi-day offline data-generation cost reported in §6 (4 / 2.5 / 1 day on 4 A100s).
  4. Related-work coverage of training-free entropy/redundancy methods is adequate, but a one-sentence comparison of wall-clock overhead versus HALT-CoT / REFRAIN would help readers place the plug-in cost.
  5. Typos: “Thisoverthinkingbehavior” (p. 1), missing spaces after periods in a few places, and “App. C” referenced before the appendix letter is introduced.

Circularity Check

0 steps flagged

No circularity: optimal-stopping objective, precomputed ground-truth rewards, and held-out length/accuracy measurements are independent of one another by construction.

full rationale

OS-Pruner is an empirical methods paper whose central chain does not reduce to its inputs. The reward r(y≤i|x)=A(y≤i|x)−λL(y≤i) (Eq. 2) is an explicit utility; A is obtained by forcing a final answer from each frozen-model prefix and grading against external ground truth with Math-Verify (§5.1), not from the stopping policy’s own scores. The policy maximizes the expected reward (Eqs. 3–4) via a small head on frozen base weights; reported 20–60% length cuts and AES numbers are then measured on held-out GSM8K/MATH-500/AIME traces under the same protocol (Table 1, Fig. 2). λ is a user-facing trade-off knob swept over a grid, not a parameter fitted to force the headline reduction. Theorem 1 is a self-contained finite-horizon counterexample (Appendix B) showing fixed-threshold correctness classifiers can lose arbitrarily large value relative to the Bellman optimum; it does not import a uniqueness result from the authors’ prior work. Baselines share the same architecture and evaluation protocol, so dominance is comparative measurement, not definitional. Concerns about forced boxed-answer labels versus free thinking (Appendix A) affect external validity of the efficiency claims, not circularity of the derivation. No self-definitional loop, fitted-as-prediction, load-bearing self-citation chain, or renamed known result is present. Score 0 is therefore the correct finding.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 1 invented entities

The central empirical claim rests on a standard sequential-decision framing plus several engineering choices (paragraph steps, forced boxed answers, Math-Verify 0/1 accuracy, λ grids, AES weights, n=2 layers). No new physical entities; free parameters are the trade-off and summary knobs that define operating points rather than hidden constants that manufacture the theorem.

free parameters (4)
  • λ (length penalty in reward)
    Scalar trading accuracy against tokens; swept over model-specific grids and annealed during progressive training. Operating point and best-AES selection depend on this choice.
  • AES coefficients α=1, β=3, γ=7
    Used to pick the reported ‘best’ policy per method; γ raised from literature default to penalize accuracy loss more heavily, affecting which λ/threshold wins Table 1.
  • n=2 fine-tuned final transformer layers
    Architectural hyperparameter fixed for all experiments; controls capacity of the stopping policy.
  • Progressive λ annealing schedule
    Hand-designed schedule to escape the trivial always-continue local minimum at low λ; not derived from theory.
axioms (5)
  • domain assumption CoT can be segmented into discrete reasoning steps (paragraphs) at which stop decisions are made.
    §3.1 defines y as a sequence of steps and invokes the policy only at paragraph boundaries (Algorithm 1).
  • domain assumption Final-answer accuracy after forced early termination is a valid 0/1 reward for that prefix (Math-Verify against ground truth).
    §5.1 preprocessing; load-bearing for precomputed rewards and offline training.
  • domain assumption Last hidden state of the (partially fine-tuned) base model encodes sufficient information for stop/continue value comparison.
    §5.2 cites prior probing work [22] and parameterizes πθ as a linear head on hl.
  • standard math Bellman optimality for finite-horizon stopping with additive length cost (standard optimal stopping).
    §4 value-function view and Theorem 1 construction.
  • ad hoc to paper Base LLM generation policy πLLM is fixed; only the stopping head is optimized.
    Design choice enabling offline reward precomputation (§3–5); excludes joint generator–stopper training.
invented entities (1)
  • OS-Pruner stopping policy head no independent evidence
    purpose: Maps reasoning prefixes to stop probabilities maximizing accuracy-minus-length utility without modifying the frozen generator.
    Method artifact, not a postulated physical entity; independent evidence is the empirical Pareto improvement and Theorem 1, not an external measurement.

pith-pipeline@v1.1.0-grok45 · 27093 in / 3230 out tokens · 38707 ms · 2026-07-14T07:07:54.680921+00:00 · methodology

0 comments
read the original abstract

Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase latency and cost without improving accuracy. Recent studies suggest that CoT trajectories can be significantly pruned, yet existing methods often rely on forcing a static thinking budget, heuristic filtering, sub-optimal early exit via classification, or expensive re-training. In this paper, we introduce OS-Pruner, a lightweight plug-in framework that formulates chain-of-thought pruning as an optimal stopping problem. Given a reasoning prefix, OS-Pruner learns whether further reasoning is worth its token cost by optimizing an explicit utility that trades off final-answer accuracy against generated length. Our novel formulation enables the model to dynamically assess the sufficient point of termination for a reasoning chain. OS-Pruner is designed to be lightweight during both training and inference, and to provide users with fine-grained control over the reasoning-effort vs. accuracy trade-off. On diverse reasoning benchmarks and base models, OS-Pruner achieves 20-60\% reduction in generation length with minimal accuracy sacrifice.

Figures

Figures reproduced from arXiv: 2607.11089 by Adam Jozefiak, Aymane El Gadarri, Ciamac C. Moallemi, Mohammed Ehab, Vivek F. Farias.

Figure 1
Figure 1. Figure 1: Training dynamics for DeepSeek-R1-Distill-Qwen-7B. The evaluations are conducted on a [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Pareto Frontiers of Expected Accuracy vs. Length Reduction across various reasoning [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Pareto Frontiers of Expected Accuracy vs. Length Reduction when the models are allowed [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 15 linked inside Pith

  1. [1]

    Chain-of-thought prompting elicits reasoning in large language models, 2023

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023

  2. [2]

    Self-consistency improves chain of thought reasoning in language models, 2023

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models, 2023

  3. [3]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, and et al. Zhu. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Nature, 645(8081):633–638, September 2025. arXiv:2501.12948 [cs]

  4. [4]

    s1: Simple test-time scaling, March 2025

    Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. s1: Simple test-time scaling, March 2025. arXiv:2501.19393 [cs]

  5. [5]

    Steering LLM Thinking with Budget Guidance, June 2025

    Junyan Li, Wenshuo Zhao, Yang Zhang, and Chuang Gan. Steering LLM Thinking with Budget Guidance, June 2025. arXiv:2506.13752 [cs]

  6. [6]

    O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning, January 2025

    Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao. O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning, January 2025. arXiv:2501.12570 [cs]

  7. [7]

    TokenSkip: Control- lable Chain-of-Thought Compression in LLMs, September 2025

    Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, and Wenjie Li. TokenSkip: Control- lable Chain-of-Thought Compression in LLMs, September 2025. arXiv:2502.12067 [cs]

  8. [8]

    Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy, February 2026

    Zeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen, Zhijian Xu, Yingying Cheng, Fan Zhang, and Qiang Xu. Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy, February 2026. arXiv:2508.03346 [cs]

  9. [9]

    DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization, March 2026

    Gang Li, Yan Chen, Ming Lin, and Tianbao Yang. DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization, March 2026. arXiv:2510.04474 [cs]

  10. [10]

    HALT-CoT: Model-Agnostic Early Stopping for Chain-of-Thought Reasoning via Answer Entropy

    Yassir Laaouach. HALT-CoT: Model-Agnostic Early Stopping for Chain-of-Thought Reasoning via Answer Entropy. June 2025

  11. [11]

    Answer Convergence as a Signal for Early Stopping in Reasoning, September 2025

    Xin Liu and Lu Wang. Answer Convergence as a Signal for Early Stopping in Reasoning, September 2025. arXiv:2506.02536 [cs]

  12. [12]

    Incentivizing llms to self-verify their answers, 2025

    Fuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui, Yang Liu, and Bo An. Incentivizing llms to self-verify their answers, 2025

  13. [13]

    FlashThink: An Early Exit Method For Efficient Reasoning, May 2025

    Guochao Jiang, Guofeng Quan, Zepeng Ding, Ziqin Luo, Dixuan Wang, and Zheng Hu. FlashThink: An Early Exit Method For Efficient Reasoning, May 2025. arXiv:2505.13949 [cs]

  14. [14]

    Early stopping chain-of-thoughts in large language models, 2025

    Minjia Mao, Bowen Yin, Yu Zhu, and Xiao Fang. Early stopping chain-of-thoughts in large language models, 2025

  15. [15]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, ...

  16. [16]

    Compressed Chain of Thought: Efficient Reasoning Through Dense Representations, December 2024

    Jeffrey Cheng and Benjamin Van Durme. Compressed Chain of Thought: Efficient Reasoning Through Dense Representations, December 2024. arXiv:2412.13171 [cs]

  17. [17]

    Training Large Language Models to Reason in a Continuous Latent Space, November

    Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training Large Language Models to Reason in a Continuous Latent Space, November

  18. [18]

    arXiv:2412.06769 [cs]

  19. [19]

    Chain of Draft: Thinking Faster by Writing Less, March 2025

    Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. Chain of Draft: Thinking Faster by Writing Less, March 2025. arXiv:2502.18600 [cs]

  20. [20]

    Think just enough: Sequence-level entropy as a confidence signal for llm reasoning, 2025

    Aman Sharma and Paras Chopra. Think just enough: Sequence-level entropy as a confidence signal for llm reasoning, 2025

  21. [21]

    Dynamic early exit in reasoning models, 2025

    Chenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu, Chenyu Zhu, Qiaowei Li, Minghui Chen, Zheng Lin, and Weiping Wang. Dynamic early exit in reasoning models, 2025

  22. [22]

    Stop when enough: Adaptive early-stopping for chain-of-thought reasoning, 2025

    Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen, and Wei Wang. Stop when enough: Adaptive early-stopping for chain-of-thought reasoning, 2025

  23. [23]

    Reasoning models know when they’re right: Probing hidden states for self-verification, 2025

    Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, and He He. Reasoning models know when they’re right: Probing hidden states for self-verification, 2025

  24. [24]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient Memory Management for Large Language Model Serving with PagedAttention, September 2023. arXiv:2309.06180 [cs]

  25. [25]

    Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement, May 2025

    Xuechen Zhang, Zijian Huang, Chenshun Ni, Ziyang Xiong, Jiasi Chen, and Samet Oymak. Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement, May 2025. arXiv:2505.07961 [cs]

  26. [26]

    Math-Verify: Math Verification Library, April 2026

    Hynek Kydlíˇcek. Math-Verify: Math Verification Library, April 2026. original-date: 2025-01- 17T11:01:36Z

  27. [27]

    OpenAI, Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, Che Chang, Kai Chen, Mark Chen, Enoch Cheung, Aidan Clark, Dan Cook, Marat Dukhan, Casey Dvorak, Kevin Fives, Vlad...

  28. [28]

    arXiv:2508.10925 [cs]

  29. [29]

    Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl

    Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Erran Li, Raluca Ada Popa, and Ion Stoica. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl. https://pretty-radio-b75.notion.site/ DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca30...

  30. [30]

    Omni- MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models, December 2024

    Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. Omni- MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models, December 2024. a...

  31. [31]

    Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems, December 2024

    Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng, Yiwen Hu, Yiru Tang, Jiapeng Wang, Xiaoxue Cheng, Huatong Song, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, and Ji-Rong Wen. Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems, December 2024. arXiv:2412.09413 [cs]

  32. [32]

    Training Verifiers to Solve Math Word Problems, November 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems, November 2021. arXiv:2110.14168 [cs]

  33. [33]

    Measuring Mathematical Problem Solving With the MATH Dataset, November 2021

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring Mathematical Problem Solving With the MATH Dataset, November 2021. arXiv:2103.03874 [cs]

  34. [34]

    Not all thoughts are generated equal: Efficient llm reasoning via multi-turn reinforcement learning, 2025

    Yansong Ning, Wei Li, Jun Fang, Naiqiang Tan, and Hao Liu. Not all thoughts are generated equal: Efficient llm reasoning via multi-turn reinforcement learning, 2025. A Ablation of Final Answer Forcing We ablate the choice of forcing a final answer by appending a box in the answer section and let the model behave freely. We accumulate the length of the rea...