Pith. sign in

REVIEW 3 major objections 5 minor 4 cited by

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Survey argues that overthinking in large reasoning models is solvable and maps two agendas — concise thinking and adaptive thinking — for doing it.

desk verdict A genuinely useful survey of concise and adaptive thinking in LRMs, but its core tables contain verifiable errors that need an audit before the paper can serve as a reliable reference. read the letter →

arxiv 2507.09662 v1 pith:I2ABVH5K submitted 2025-07-13 cs.AI cs.CL

classification cs.AIcs.CL
keywords largereasoningmodelschain-of-thoughtoverthinkingconcisethinkingadaptiveinferenceefficiencyreinforcementlearningsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large reasoning models such as o1 and DeepSeek R1 solve hard math and coding problems by emitting long, step-by-step chains of thought, but they also emit long chains for trivial questions, wasting tokens, time, and money in a pattern the field calls overthinking. This survey sets out to establish that overthinking is a coherent research problem with two complementary solutions: concise thinking, which shortens a given reasoning chain, and adaptive thinking, which decides per query whether to reason at all and how long to reason. The paper organizes the current literature into training-free methods (prompting, pipelines, decoding interventions, model merging) and training-based methods (fine-tuning and reinforcement learning), and reviews the metrics and benchmarks used to measure reasoning efficiency. Its value to a reader is as a map: it shows where each method sits and argues that future work must balance input difficulty, model capability, human preference, and trustworthiness rather than optimize token counts alone.

What carries the argument

The load-bearing structure of the survey is its taxonomy. Methods are split into training-free approaches (prompt-guided strategies, pipeline and router systems, decoding manipulation such as budget forcing, early exiting, logit adjustment, and activation steering, plus model merging) and training-based approaches (variable-length data construction, fine-tuning including implicit and DPO-variant methods, and reinforcement learning with length penalties, GRPO variants, difficulty awareness, and thinking-mode selection). Alongside the taxonomy, the survey assembles a set of efficiency metrics — token counts, latency, speed-up ratio, $E^3 = A^2/T$, Think Density, HCA/SCA/CCA, InfoBias and InfoGain among others — that make the field's goal of accuracy parity with fewer tokens measurable and comparable.

What would settle it

A reader can audit the paper's own tables against its reference list: Table 3 attributes C3oT to Luo et al. (2025b) while the text and bibliography cite Kang et al. (2024). A broader audit of this kind, combined with a uniform reproduction of the cited methods on a fixed benchmark set (for example GSM8K, MATH-500, and AIME24 against a shared long-CoT baseline), would settle whether the taxonomy's attributions are reliable and whether the field's accuracy-parity-with-shorter-chains premise actually holds.

Watch

Extended reading notes

Core claim

The authors' central claim is that the overthinking behavior of large reasoning models is not an unavoidable side effect of their reasoning ability; it can be reduced, and a growing body of recent work already reduces it through two complementary agendas. Concise thinking keeps one reasoning chain but makes it shorter, while adaptive thinking switches between fast direct answers and slow chains based on input difficulty. The survey further claims that the empirical pattern motivating these efforts is real: accuracy does not keep improving with chain length, and different tasks have different optimal reasoning lengths. It supports the map with observations such as RL-trained models frequently reasoning even when explicitly instructed to skip thinking, whereas SFT-based models follow concise-thinking instructions more obediently, which means methods do not transfer uniformly across model families.

Load-bearing premise

The survey's value as a map depends on its taxonomy and citations being accurate and complete; if methods are misattributed or categories misrepresent the literature, readers navigating by this guide will land on the wrong papers or wrong distinctions.

Editorial extensions

If this is right

  • Practitioners can pick an approach by their training budget: prompt engineering and decoding tricks give immediate, training-free savings, while fine-tuning and reinforcement learning produce more persistent changes in reasoning behavior.
  • Because RL-based and SFT-based models respond differently to no-thinking instructions, an efficiency method validated on one model family cannot be assumed to transfer to another; evaluations should name the base model type.
  • The wide spread of evaluation metrics and benchmarks means no single number yet captures the accuracy-efficiency trade-off; until the field standardizes base models, datasets, and hyperparameters, cross-paper comparisons stay provisional.
  • If the overthinking premise holds, product builders can cut inference cost and latency on simple queries, which is the main practical payoff of the surveyed line of work.
  • Defining what counts as a good reasoning chain must include interpretability and safety, not just length, because aggressive compression can remove exactly the transparency that makes reasoning models valuable to users.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A maintainable version of this survey would be a machine-checkable table with one verified row per method; the visible mismatch in Table 3 — C3oT attributed to Luo et al. (2025b) while the text cites Kang et al. (2024) — suggests the current map is not yet that.
  • Router-based methods point toward a natural product pattern: a cheap classifier sends easy queries to fast mode and hard ones to slow mode. An implicit extension, not designed in the paper, is to make the router preference-aware — for example, always showing full reasoning for medical or legal answers even when the query looks easy.
  • The $E^3 = A^2/T$ metric implies an accuracy-versus-token Pareto frontier for each dataset; plotting methods that way would let users select an operating point, which the survey itself does not do.
  • A testable cross-family question follows from the routing work: does a capability-aware router trained on one reasoning model family (say DeepSeek-R1-Distill) generalize to another family (say QwQ)? The paper lists the ingredients but leaves the transfer question open.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper surveys recent work on making large reasoning models (LRMs) more efficient by producing concise or adaptively selected reasoning chains. The survey organizes methods into training-free approaches (prompt-guided, pipeline, decoding manipulation, model merging) and training-based approaches (fine-tuning and reinforcement learning), and it also reviews evaluation benchmarks, metrics, and open challenges. The intended contribution is a comprehensive, up-to-date guide to the field of concise and adaptive thinking for LRMs.

Significance. The survey addresses a timely and practically important problem: reducing the excessive inference cost and latency of LRMs while preserving reasoning accuracy. Its taxonomy is broadly sensible, and it covers a wide range of recent works, including many 2024–2025 preprints. The paper is explicitly scoped and does not attempt to cover all of efficient inference, which is appropriate. Its value as a reference, however, depends on the accuracy of its tables and attributions. The concrete errors identified below (notably the C3oT misattribution in Table 3 and the ambiguous Ling et al. reward in Table 5) undermine confidence in the survey as a reliable indexing tool. These issues are correctable, and the core organizational framework of the survey is a contribution in itself.

major comments (3)
  1. [Table 3 vs. Section 4.1.1] Table 3 lists 'C3oT (Luo et al., 2025b)' with base models LLaMA2-Chat-7B/13B and datasets ECQA/StrategyQA, but the text in Section 4.1.1 and the reference list identify C3oT as Kang et al. (2024), and Figure 1 also cites Kang et al. (2024). Luo et al. (2025b) is Ada-R1 elsewhere in the survey. This misattribution directs a reader who uses Table 3 as an index to the wrong paper and must be corrected.
  2. [Table 5, Ling et al. (2025) row] The printed reward function for Ling et al. is 1 + αL(y)^γ for correct answers. For positive α and γ, this assigns a larger reward to longer correct chains, which is the opposite of a length penalty. Section 4.2.1 describes Ling et al. as adopting a 'powered length penalty (PLP) that is more lenient on longer responses,' implying a penalty that should reduce reward as length grows. The table must either explicitly state that α < 0, show a minus sign, or otherwise reconcile the formula with the textual description; as printed, the table is internally inconsistent and misleading about the method.
  3. [Table 4] The reliability of Table 4 as a guide to open-source variable-length CoT datasets needs verification and correction. The OmniThought sample count appears as '2000,000,' which is a typo for 2,000,000. Additionally, 's1K-mix' is attributed to Yu et al. (2025a), while Section 4.1.2 introduces s1K under Muennighoff et al. (2025); the relationship between s1K and s1K-mix, and the correct attribution, should be clarified. Since the survey's central claim is to be a comprehensive and accurate guide, the tables require a careful audit.
minor comments (5)
  1. [Tables 2 and 3] The dataset name 'SV AMP' appears in Tables 2 and 3; the correct name is SVAMP (Patel et al., 2021), and the formatting should be consistent.
  2. [Table 3, DeGRPO row] The dataset 'MATH-50' in the DeGRPO row appears to be a typo for MATH-500.
  3. [Sections 2.3 and 4.1.2] The metric from Cai et al. (2025) is called 'Reasoning Verbosity' in Section 2.3 and 'Reasoning Veracity (RV)' in Section 4.1.2; the survey should use one consistent term.
  4. [Section 2.3, E^3 metric] The formula for E^3 is written as 'A2T' without formatting; it should be clearly displayed as A^2 divided by T to avoid ambiguity.
  5. [Table 5 caption] The caption defines S(y) and L(y), but the table uses additional symbols (α, γ, λ, L_budget, L_cache, L_max, etc.) that are not defined in the caption. Adding definitions or referencing the original equations would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey derives no quantities, fits no parameters, and contains no load-bearing self-citations.

full rationale

This paper is a literature survey rather than a derivation-based research contribution. Its central claim is that it provides a comprehensive overview of concise and adaptive thinking in large reasoning models, organized into training-free and training-based taxonomies. There are no fitted parameters, no predictive equations derived from inputs, and no claimed first-principles results that could reduce to their own assumptions. The authors do not cite their own prior work in any load-bearing way, and the taxonomy and narrative summaries are external descriptive claims about cited papers rather than circular derivations. The concrete accuracy problems noted by the reader, such as the C3oT attribution in Table 3 pointing to Luo et al. (2025b) while the text and references identify Kang et al. (2024), and the Table 5 display of the Ling et al. (2025) reward as 1 + alpha*L(y)^gamma, which would reward longer correct answers rather than penalize them, are real scholarly-quality concerns. However, misattribution and formula transcription errors are not circularity: the survey is not defining its conclusions in terms of its premises, nor is it renaming a fitted input as a prediction. Because the paper makes no internally derived quantitative claims, there is no circular reasoning chain to expose, and the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

This is a survey paper with no mathematical derivations, fitted parameters, or invented entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey." pith.science (2026). https://pith.science/paper/I2ABVH5K

@misc{pith2026250709662,
  author       = {Pith},
  title        = {Pith review of: Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I2ABVH5K}},
  note         = {Machine review of arXiv:2507.09662}
}
read the original abstract

Large reasoning models (LRMs) like OpenAI o1 and DeepSeek R1 have demonstrated impressive performance on complex reasoning tasks like mathematics and programming with long Chain-of-Thought (CoT) reasoning sequences (slow-thinking), compared with traditional large language models (fast-thinking). However, these reasoning models also face a huge challenge that generating unnecessarily lengthy and redundant reasoning chains even for trivial questions. This phenomenon leads to a significant waste of inference resources, increases the response time for simple queries, and hinders the practical application of LRMs in real-world products. To this end, it is crucial to shorten lengthy reasoning chains and learn adaptive reasoning between fast and slow thinking based on input difficulty. In this survey, we provide a comprehensive overview of recent progress in concise and adaptive thinking for efficient reasoning of LRMs, including methodologies, benchmarks, and challenges for future exploration. We hope this survey can help researchers quickly understand the landscape of this field and inspire novel adaptive thinking ideas to facilitate better usage of LRMs.

Figures

Figures reproduced from arXiv: 2507.09662 by the authors.

Figure 1
Figure 1. Taxonomy of Concise and Adaptive Thinking in LRMs. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Consequence-aware scheduler using an issue-text predictor routes more compute to high-cost failures and cuts cost-weighted loss by 22-33% versus difficulty-based allocation on SWE-bench tasks.

  2. ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute

    cs.CL 2025-08 conditional novelty 6.0 of 10

    ParaThinker trains LLMs for native parallel reasoning and reports 7 to 12 percent higher accuracy on math benchmarks over sequential thinking with modest latency overhead.

  3. BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A control-token insertion and two-stage training method that lets LLMs adhere to user-specified reasoning token budgets while preserving math accuracy.

  4. Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A survey that maps reinforcement learning methods, datasets, benchmarks, and open-source tools across the full training lifecycle of large language models, focusing on verifiable-reward reasoning.

Reference graph

Works this paper leans on

256 extracted references · 8 canonical work pages · cited by 4 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Aradhye Agarwal, Ayan Sengupta, and Tanmoy Chakraborty. 2025. https://arxiv.org/abs/2505.18149 First finish search: Efficient test-time scaling in large language models . Preprint, arXiv:2505.18149

  4. [4]

    Pranjal Aggarwal and Sean Welleck. 2025. https://arxiv.org/abs/2503.04697 L1: Controlling how long a reasoning model thinks with reinforcement learning . Preprint, arXiv:2503.04697

  5. [5]

    Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. https://doi.org/10.18653/v1/2021.acl-long.238 E xplanations for C ommonsense QA : N ew D ataset and M odels . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferen...

  6. [6]

    Syeda Nahida Akter, Shrimai Prabhumoye, Matvei Novikov, Seungju Han, Ying Lin, Evelina Bakhturina, Eric Nyberg, Yejin Choi, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2025. https://arxiv.org/abs/2504.13941 Nemotron-crossthink: Scaling self-learning beyond math reasoning . Preprint, arXiv:2504.13941

  7. [7]

    Mohammad Ali Alomrani, Yingxue Zhang, Derek Li, Qianyi Sun, Soumyasundar Pal, Zhanguang Zhang, Yaochen Hu, Rohan Deepak Ajwani, Antonios Valkanas, Raika Karimi, Peng Cheng, Yunzhou Wang, Pengyi Liao, Hanrui Huang, Bin Wang, Jianye Hao, and Mark Coates. 2025. https://arxiv.org/abs/2507.02076 Reasoning on a budget: A survey of adaptive and controllable test...

  8. [8]

    Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. https://arxiv.org/abs/1905.13319 Mathqa: Towards interpretable math word problem solving with operation-based formalisms . Preprint, arXiv:1905.13319

Show all 256 references
  1. [9]

    Sohyun An, Ruochen Wang, Tianyi Zhou, and Cho-Jui Hsieh. 2025. https://arxiv.org/abs/2505.21765 Don't think longer, think wisely: Optimizing thinking dynamics for large reasoning models . Preprint, arXiv:2505.21765

  2. [10]

    anthropic . 2025. https://www.anthropic.com/news/claude-3-7-sonnet

  3. [11]

    Daman Arora and Andrea Zanette. 2025. https://arxiv.org/abs/2502.04463 Training language models to reason efficiently . Preprint, arXiv:2502.04463

  4. [12]

    Dhananjay Ashok and Jonathan May. 2025. https://arxiv.org/abs/2502.13329 Language models can predict their own behavior . Preprint, arXiv:2502.13329

  5. [13]

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. https://arxiv.org/abs/2108.07732 Program synthesis with large language models . Preprint, arXiv:2108.07732

  6. [14]

    Aytes, Jinheon Baek, and Sung Ju Hwang

    Simon A. Aytes, Jinheon Baek, and Sung Ju Hwang. 2025. https://arxiv.org/abs/2503.05179 Sketch-of-thought: Efficient llm reasoning with adaptive cognitive-inspired sketching . Preprint, arXiv:2503.05179

  7. [15]

    Ayers, Dragomir Radev, and Jeremy Avigad

    Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, and Jeremy Avigad. 2023. https://arxiv.org/abs/2302.12433 Proofnet: Autoformalizing and formally proving undergraduate-level mathematics . Preprint, arXiv:2302.12433

  8. [16]

    Mislav Balunović, Jasper Dekoninck, Ivo Petrov, Nikola Jovanović, and Martin Vechev. 2025. https://arxiv.org/abs/2505.23281 Matharena: Evaluating llms on uncontaminated math competitions . Preprint, arXiv:2505.23281

  9. [17]

    Akhiad Bercovich, Itay Levy, Izik Golan, Mohammad Dabbah, et al. 2025. https://arxiv.org/abs/2505.00949 Llama-nemotron: Efficient reasoning models . Preprint, arXiv:2505.00949

  10. [18]

    Bytedance Seed Team . 2025. https://www.volcengine.com/

  11. [19]

    Wenrui Cai, Chengyu Wang, Junbing Yan, Jun Huang, and Xiangzhong Fang. 2025. https://arxiv.org/abs/2505.10937 Reasoning with omnithought: A large cot dataset with verbosity and cognitive difficulty annotations . Preprint, arXiv:2505.10937

  12. [20]

    Hanting Chen, Yasheng Wang, Kai Han, Dong Li, Lin Li, Zhenni Bi, Jinpeng Li, Haoyu Wang, Fei Mi, Mingjian Zhu, Bin Wang, Kaikai Song, Yifei Fu, Xu He, Yu Luo, Chong Zhu, Quan He, Xueyu Wu, Wei He, Hailin Hu, Yehui Tang, Dacheng Tao, Xinghao Chen, and Yunhe Wang. 2025 a . https...

  13. [21]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, et al. 2021. https://arxiv.org/abs/2107.03374 Evaluating large language models trained on code . Preprint, arXiv:2107.03374

  14. [22]

    Weize Chen, Jiarui Yuan, Tailin Jin, Ning Ding, Huimin Chen, Zhiyuan Liu, and Maosong Sun. 2025 b . https://arxiv.org/abs/2505.19217 The overthinker's diet: Cutting token calories with difficulty-aware training . Preprint, arXiv:2505.19217

  15. [23]

    Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu. 2025 c . https://arxiv.org/abs/2412.21187 Do not think that much for 2+3=? on the overthinking of o1-li...

  16. [24]

    Zigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu, and Xinchao Wang. 2025 d . https://arxiv.org/abs/2505.17941 Verithinker: Learning to verify makes reasoning model efficient . Preprint, arXiv:2505.17941

  17. [25]

    Jeffrey Cheng and Benjamin Van Durme. 2024. https://arxiv.org/abs/2412.13171 Compressed chain of thought: Efficient reasoning through dense representations . Preprint, arXiv:2412.13171

  18. [26]

    Xiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang, Wayne Xin Zhao, Xinyu Kong, and Zhiqiang Zhang. 2025 a . https://arxiv.org/abs/2505.16315 Incentivizing dual process thinking for efficient large language model reasoning . Preprint, arXiv:2505.16315

  19. [27]

    Xiaoxue Cheng, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. 2025 b . https://arxiv.org/abs/2501.01306 Think more, hallucinate less: Mitigating hallucinations via dual process of fast and slow thinking . Preprint, arXiv:2501.01306

  20. [28]

    Zhengxiang Cheng, Dongping Chen, Mingyang Fu, and Tianyi Zhou. 2025 c . https://arxiv.org/abs/2506.14755 Optimizing length compression in large reasoning models . Preprint, arXiv:2506.14755

  21. [29]

    Stephen Chung, Wenyu Du, and Jie Fu. 2025. https://arxiv.org/abs/2505.21097 Thinker: Learning to think fast and slow . Preprint, arXiv:2505.21097

  22. [30]

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://arxiv.org/abs/1803.05457 Think you have solved question answering? try arc, the ai2 reasoning challenge . Preprint, arXiv:1803.05457

  23. [31]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . ...

  24. [32]

    Cognitive Computations. 2025. https://huggingface.co/datasets/cognitivecomputations/dolphin-r1

  25. [33]

    Gonzalez

    Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, Nicholas Thumiger, Aditya Desai, Ion Stoica, Ana Klimovic, Graham Neubig, and Joseph E. Gonzalez. 2025. https://arxiv.org/abs/2502.08235 ...

  26. [34]

    Renfei Dang, Shujian Huang, and Jiajun Chen. 2025. https://arxiv.org/abs/2505.16448 Internal bias in reasoning models leads to overthinking . Preprint, arXiv:2505.16448

  27. [35]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, et al. 2025. https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning . Preprint, arXiv:2501.12948

  28. [36]

    Hexuan Deng, Wenxiang Jiao, Xuebo Liu, Jun Rao, and Min Zhang. 2025. https://arxiv.org/abs/2505.19862 Rea-rl: Reflection-aware online reinforcement learning for efficient large reasoning models . Preprint, arXiv:2505.19862

  29. [37]

    Yuntian Deng, Yejin Choi, and Stuart Shieber. 2024. https://arxiv.org/abs/2405.14838 From explicit cot to implicit cot: Learning to internalize cot step by step . Preprint, arXiv:2405.14838

  30. [38]

    Bowen Ding, Yuhan Chen, Futing Wang, Lingfeng Ming, and Tao Lin. 2025. https://arxiv.org/abs/2506.23840 Do thinking tokens help or trap? towards more efficient large reasoning model . Preprint, arXiv:2506.23840

  31. [39]

    Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Ruhle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah. 2024. https://arxiv.org/abs/2404.14618 Hybrid llm: Cost-efficient and quality-aware query routing . Preprint, arXiv:2404.14618

  32. [40]

    Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. https://arxiv.org/abs/1903.00161 Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs . Preprint, arXiv:1903.00161

  33. [41]

    Hashimoto

    Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2025. https://arxiv.org/abs/2404.04475 Length-controlled alpacaeval: A simple way to debias automatic evaluators . Preprint, arXiv:2404.04475

  34. [42]

    Razvan-Gabriel Dumitru, Darius Peteleaza, Vikas Yadav, and Liangming Pan. 2025. https://arxiv.org/abs/2505.17250 Conciserl: Conciseness-guided reinforcement learning for efficient reasoning models . Preprint, arXiv:2505.17250

  35. [43]

    Roy Eisenstadt, Itamar Zimerman, and Lior Wolf. 2025. https://arxiv.org/abs/2506.07240 Overclocking llm reasoning: Monitoring and controlling thinking path lengths in llms . Preprint, arXiv:2506.07240

  36. [44]

    Chenrui Fan, Ming Li, Lichao Sun, and Tianyi Zhou. 2025 a . https://arxiv.org/abs/2504.06514 Missing premise exacerbates overthinking: Are reasoning models losing critical thinking skill? Preprint, arXiv:2504.06514

  37. [45]

    Siqi Fan, Peng Han, Shuo Shang, Yequan Wang, and Aixin Sun. 2025 b . https://arxiv.org/abs/2505.22017 Cothink: Token-efficient reasoning via instruct models guiding reasoning models . Preprint, arXiv:2505.22017

  38. [46]

    Gongfan Fang, Xinyin Ma, and Xinchao Wang. 2025. https://arxiv.org/abs/2505.13379 Thinkless: Llm learns when to think . Preprint, arXiv:2505.13379

  39. [47]

    Mehdi Fatemi, Banafsheh Rafiee, Mingjie Tang, and Kartik Talamadupula. 2025. https://arxiv.org/abs/2504.05185 Concise reasoning via reinforcement learning . Preprint, arXiv:2504.05185

  40. [48]

    Sicheng Feng, Gongfan Fang, Xinyin Ma, and Xinchao Wang. 2025. https://arxiv.org/abs/2504.10903 Efficient reasoning models: A survey . Preprint, arXiv:2504.10903

  41. [49]

    Tianyu Fu, Yi Ge, Yichen You, Enshu Liu, Zhihang Yuan, Guohao Dai, Shengen Yan, Huazhong Yang, and Yu Wang. 2025 a . https://arxiv.org/abs/2505.21600 R2r: Efficiently navigating divergent reasoning paths with small-large model token routing . Preprint, arXiv:2505.21600

  42. [50]

    Tingchen Fu, Jiawei Gu, Yafu Li, Xiaoye Qu, and Yu Cheng. 2025 b . https://arxiv.org/abs/2505.14810 Scaling reasoning, losing control: Evaluating instruction following in large reasoning models . Preprint, arXiv:2505.14810

  43. [51]

    Yichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu, Zhongdongming Dai, Yonghao Zhuang, Yian Ma, Aurick Qiao, Tajana Rosing, Ion Stoica, and Hao Zhang. 2025 c . https://arxiv.org/abs/2412.20993 Efficiently scaling llm reasoning with certaindex . Preprint, arXiv:2412.20993

  44. [52]

    Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. 2024. https://arxiv.org...

  45. [53]

    Jiaxuan Gao, Shu Yan, Qixin Tan, Lu Yang, Shusheng Xu, Wei Fu, Zhiyu Mei, Kaifeng Lyu, and Yi Wu. 2025. https://arxiv.org/abs/2506.07104 How far are we from optimal reasoning efficiency? Preprint, arXiv:2506.07104

  46. [54]

    Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. https://arxiv.org/abs/2101.02235 Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies . Preprint, arXiv:2101.02235

  47. [55]

    Soumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu, Mengdi Wang, Dinesh Manocha, Furong Huang, Mohammad Ghavamzadeh, and Amrit Singh Bedi. 2025. https://arxiv.org/abs/2506.04210 Does thinking more always help? understanding test-time scaling in reasoning models ...

  48. [56]

    Ruihan Gong, Yue Liu, Wenjie Qu, Mingzhe Du, Yufei He, Yingwei Ma, Yulin Chen, Xiang Liu, Yi Wen, Xinfeng Li, Ruidong Wang, Xinzhong Zhu, Bryan Hooi, and Jiaheng Zhang. 2025. https://arxiv.org/abs/2505.19756 Efficient reasoning via chain of unconscious thought . Preprint, arXi...

  49. [57]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. 2024. https://arxiv.org/abs/2407.21783 The llama 3 herd of models . Preprint, arXiv:2407.21783

  50. [58]

    Merrill, Tatsunori Hashimoto, Yejin Choi, Jenia Jitsev, Reinhard Heckel, Maheswaran Sathiamoorthy, Alexandros G

    Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, Ashima Suvarna, Benjamin Feuer, Liangyu Chen, Zaid Khan, Eric Frankel, Sachin Grover, Caroline Choi, Niklas Muennighoff, Shiye Su, Wanj...

  51. [59]

    Hasan Abed Al Kader Hammoud, Hani Itani, and Bernard Ghanem. 2025. https://arxiv.org/abs/2504.20708 Beyond the last answer: Your reasoning trace uncovers more than you think . Preprint, arXiv:2504.20708

  52. [60]

    Tingxu Han, Zhenting Wang, Chunrong Fang, Shiyu Zhao, Shiqing Ma, and Zhenyu Chen. 2025. https://arxiv.org/abs/2412.18547 Token-budget-aware llm reasoning . Preprint, arXiv:2412.18547

  53. [61]

    Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024. https://arxiv.org/abs/2412.06769 Training large language models to reason in a continuous latent space . Preprint, arXiv:2412.06769

  54. [62]

    Masoud Hashemi, Oluwanifemi Bamgbose, Sathwik Tejaswi Madhusudhan, Jishnu Sethumadhavan Nair, Aman Tiwari, and Vikas Yadav. 2025. https://arxiv.org/abs/2503.15793 Dnr bench: Benchmarking over-reasoning in reasoning llms . Preprint, arXiv:2503.15793

  55. [63]

    Michael Hassid, Gabriel Synnaeve, Yossi Adi, and Roy Schwartz. 2025. https://arxiv.org/abs/2505.17813 Don't overthink it. preferring shorter thinking chains for improved llm reasoning . Preprint, arXiv:2505.17813

  56. [64]

    Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024 a . https://arxiv.org/abs/2402.14008 Olympiadbench: A challenging benchmark for promoting agi with o...

  57. [65]

    Jujie He, Jiacai Liu, Chris Yuhao Liu, Rui Yan, Chaojie Wang, Peng Cheng, Xiaoyu Zhang, Fuxiang Zhang, Jiacheng Xu, Wei Shen, Siyuan Li, Liang Zeng, Tianwen Wei, Cheng Cheng, Bo An, Yang Liu, and Yahui Zhou. 2025 a . https://arxiv.org/abs/2505.22312 Skywork open reasoner 1 tec...

  58. [66]

    Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Xuepeng Liu, Dekai Sun, Shirong Lin, Zhicheng Zheng, Xiaoyong Zhu, Wenbo Su, and Bo Zheng. 2024 b . https://arxiv.org/abs/2411.07140 Chin...

  59. [67]

    Yang He, Xiao Ding, Bibo Cai, Yufei Zhang, Kai Xiong, Zhouhao Sun, Bing Qin, and Ting Liu. 2025 b . https://arxiv.org/abs/2505.20664 Self-route: Automatic mode switching via capability estimation for efficient reasoning . Preprint, arXiv:2505.20664

  60. [68]

    Zhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu, Xingyu Chen, Yue Wang, Linfeng Song, Dian Yu, Zhenwen Liang, Wenxuan Wang, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu. 2025 c . https://arxiv.org/abs/2504.11456 Deepmath-103k: A large-scale, challenging, deconta...

  61. [69]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 a . https://arxiv.org/abs/2009.03300 Measuring massive multitask language understanding . Preprint, arXiv:2009.03300

  62. [70]

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 b . https://arxiv.org/abs/2103.03874 Measuring mathematical problem solving with the math dataset . Preprint, arXiv:2103.03874

  63. [71]

    Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, and Shiyu Chang. 2025. https://arxiv.org/abs/2504.01296 Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning . Preprint, arXiv:2504.01296

  64. [72]

    Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024. https://arxiv.org/abs/2403.12031 Routerbench: A benchmark for multi-llm routing system . Preprint, arXiv:2403.12031

  65. [73]

    Chengyu Huang, Zhengxin Zhang, and Claire Cardie. 2025 a . https://arxiv.org/abs/2505.11225 Hapo: Training language models to reason concisely via history-aware policy optimization . Preprint, arXiv:2505.11225

  66. [74]

    Shijue Huang, Hongru Wang, Wanjun Zhong, Zhaochen Su, Jiazhan Feng, Bowen Cao, and Yi R. Fung. 2025 b . https://arxiv.org/abs/2505.18822 Adactrl: Towards adaptive and controllable reasoning via difficulty-aware budgeting . Preprint, arXiv:2505.18822

  67. [75]

    Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Yichang Xu, and Ling Liu. 2025 c . https://arxiv.org/abs/2503.00555 Safety tax: Safety alignment makes your large reasoning models less reasonable . Preprint, arXiv:2503.00555

  68. [76]

    Yao Huang, Huanran Chen, Shouwei Ruan, Yichi Zhang, Xingxing Wei, and Yinpeng Dong. 2025 d . https://arxiv.org/abs/2505.22411 Mitigating overthinking in large reasoning models via manifold steering . Preprint, arXiv:2505.22411

  69. [77]

    iFLYTEK Xinghuo Team . 2025. https://xinghuo.xfyun.cn/sparkapi

  70. [78]

    Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023. https://arxiv.org/abs/2212.04089 Editing models with task arithmetic . Preprint, arXiv:2212.04089

  71. [79]

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. https://arxiv.org/abs/2403.07974 Livecodebench: Holistic and contamination free evaluation of large language models for code . Preprint, a...

  72. [80]

    Yunjie Ji, Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yiping Peng, Han Zhao, and Xiangang Li. 2025. https://arxiv.org/abs/2505.08311 Am-thinking-v1: Advancing the frontier of reasoning at 32b scale . Preprint, arXiv:2505.08311

  73. [81]

    Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. 2025 a . https://arxiv.org/abs/2502.12025 Safechain: Safety of language models with long chain-of-thought reasoning capabilities . Preprint, arXiv:2502.12025

  74. [82]

    Gangwei Jiang, Yahui Liu, Zhaoyi Li, Qi Wang, Fuzheng Zhang, Linqi Song, Ying Wei, and Defu Lian. 2025 b . https://arxiv.org/abs/2505.22148 What makes a good reasoning chain? uncovering structural patterns in long chain-of-thought reasoning . Preprint, arXiv:2505.22148

  75. [83]

    Guochao Jiang, Guofeng Quan, Zepeng Ding, Ziqin Luo, Dixuan Wang, and Zheng Hu. 2025 c . https://arxiv.org/abs/2505.13949 Flashthink: An early exit method for efficient reasoning . Preprint, arXiv:2505.13949

  76. [84]

    Lingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong, Xingxing Zhang, Tengchao Lv, Lei Cui, and Furu Wei. 2025 d . https://arxiv.org/abs/2505.14631 Think only when you need with large hybrid-reasoning models . Preprint, arXiv:2505.14631

  77. [85]

    Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. 2024. https://arxiv.org/abs/2406.18510 Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer la...

  78. [86]

    Yuxuan Jiang, Dawei Li, and Frank Ferraro. 2025 e . https://arxiv.org/abs/2505.13975 Drp: Distilled reasoning pruning with skill-aware step decomposition for efficient large reasoning models . Preprint, arXiv:2505.13975

  79. [87]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. https://arxiv.org/abs/2310.06770 Swe-bench: Can language models resolve real-world github issues? Preprint, arXiv:2310.06770

  80. [88]

    Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du. 2024. https://arxiv.org/abs/2401.04925 The impact of reasoning step length on large language models . Preprint, arXiv:2401.04925

  81. [89]

    Zhensheng Jin, Xinze Li, Yifan Ji, Chunyi Peng, Zhenghao Liu, Qi Shi, Yukun Yan, Shuo Wang, Furong Peng, and Ge Yu. 2025. https://arxiv.org/abs/2506.10822 Recut: Balancing reasoning length and accuracy in llms via stepwise trails and preference optimization . Preprint, arXiv:2...

  82. [90]

    Yu Kang, Xianghui Sun, Liangyu Chen, and Wei Zou. 2024. https://arxiv.org/abs/2412.11664 C3ot: Generating shorter chain-of-thought without compromising effectiveness . Preprint, arXiv:2412.11664

  83. [91]

    KwaiPilot Team . 2025. https://huggingface.co/kwaipilot/kwaicoder-autothink-preview

  84. [92]

    Ayeong Lee, Ethan Che, and Tianyi Peng. 2025. https://arxiv.org/abs/2503.01141 How well do llms compress their own chain-of-thought? a token complexity approach . Preprint, arXiv:2503.01141

  85. [93]

    Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. https://arxiv.org/abs/2206.14858 Solving quantitative re...

  86. [94]

    Junyan Li, Wenshuo Zhao, Yang Zhang, and Chuang Gan. 2025 a . https://arxiv.org/abs/2506.13752 Steering llm thinking with budget guidance . Preprint, arXiv:2506.13752

  87. [95]

    Gonzalez, and Ion Stoica

    Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024. https://arxiv.org/abs/2406.11939 From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline . Preprint, arXiv:2406.11939

  88. [96]

    Kwok, and Yu Zhang

    Wei Li, Yanbin Wei, Qiushi Huang, Jiangyue Yan, Yang Chen, James T. Kwok, and Yu Zhang. 2025 b . https://arxiv.org/abs/2506.05936 Dynamicmind: A tri-mode thinking system for large language models . Preprint, arXiv:2506.05936

  89. [97]

    Xuying Li, Zhuo Li, Yuji Kosuga, and Victor Bian. 2025 c . https://arxiv.org/abs/2503.01923 Output length effect on deepseek-r1's safety in forced thinking . Preprint, arXiv:2503.01923

  90. [98]

    Zheng Li, Qingxiu Dong, Jingyuan Ma, Di Zhang, and Zhifang Sui. 2025 d . https://arxiv.org/abs/2505.11274 Selfbudgeter: Adaptive token allocation for efficient llm reasoning . Preprint, arXiv:2505.11274

  91. [99]

    Zhiyuan Li, Yi Chang, and Yuan Wu. 2025 e . https://arxiv.org/abs/2505.22113 Think-bench: Evaluating thinking efficiency and chain-of-thought quality of large reasoning models . Preprint, arXiv:2505.22113

  92. [100]

    Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Ying Nian Wu, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, and Cheng-Lin Liu. 2025 f . https://arxiv.org/abs/2506.02678 Tl;dr: Too long, do re-weighting for efficient llm...

  93. [101]

    Zhong-Zhi Li, Duzhen Zhang, Ming-Liang Zhang, Jiaxin Zhang, Zengyan Liu, Yuxuan Yao, Haotian Xu, Junhao Zheng, Pei-Jie Wang, Xiuyi Chen, Yingying Zhang, Fei Yin, Jiahua Dong, Zhiwei Li, Bao-Long Bi, Ling-Rui Mei, Junfeng Fang, Zhijiang Guo, Le Song, and Cheng-Lin Liu. 2025 g ....

  94. [102]

    Guosheng Liang, Longguang Zhong, Ziyi Yang, and Xiaojun Quan. 2025. https://arxiv.org/abs/2505.14183 Thinkswitcher: When to think hard, when to think fast . Preprint, arXiv:2505.14183

  95. [103]

    Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan. 2024. https://arxiv.org/abs/2401.08190 Mario: Math reasoning with code interpreter output -- a reproducible pipeline . Preprint, arXiv:2401.08190

  96. [104]

    Junhong Lin, Xinyue Zeng, Jie Zhu, Song Wang, Julian Shun, Jun Wu, and Dawei Zhou. 2025. https://arxiv.org/abs/2505.16122 Plan and budget: Effective and efficient test-time scaling on large language model reasoning . Preprint, arXiv:2505.16122

  97. [105]

    Zehui Ling, Deshu Chen, Hongwei Zhang, Yifeng Jiao, Xin Guo, and Yuan Cheng. 2025. https://arxiv.org/abs/2506.10446 Fast on the easy, deep on the hard: Efficient reasoning via powered length penalty . Preprint, arXiv:2506.10446

  98. [106]

    Fengyuan Liu, Nouar AlDahoul, Gregory Eady, Yasir Zaki, and Talal Rahwan. 2025 a . https://arxiv.org/abs/2406.10400 Self-reflection makes large language models safer, less biased, and ideologically neutral . Preprint, arXiv:2406.10400

  99. [107]

    Hongwei Liu, Zilong Zheng, Yuxuan Qiao, Haodong Duan, Zhiwei Fei, Fengzhe Zhou, Wenwei Zhang, Songyang Zhang, Dahua Lin, and Kai Chen. 2024 a . https://arxiv.org/abs/2405.12209 Mathbench: Evaluating the theory and application proficiency of llms with a hierarchical mathematics...

  100. [108]

    Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020. https://arxiv.org/abs/2007.08124 Logiqa: A challenge dataset for machine reading comprehension with logical reasoning . Preprint, arXiv:2007.08124

  101. [109]

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. https://arxiv.org/abs/2305.01210 Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation . Preprint, arXiv:2305.01210

  102. [110]

    Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. 2024 b . https://arxiv.org/abs/2408.06450 Evaluating language models for efficient code generation . Preprint, arXiv:2408.06450

  103. [111]

    Kaiyuan Liu, Chen Shen, Zhanwei Zhang, Junjie Liu, Xiaosong Yuan, and Jieping ye. 2025 b . https://arxiv.org/abs/2506.12353 Efficient reasoning through suppression of self-affirmation reflections in large reasoning models . Preprint, arXiv:2506.12353

  104. [112]

    Peijie Liu, Fengli Xu, and Yong Li. 2025 c . https://arxiv.org/abs/2506.06008 Token signature: Predicting chain-of-thought gains with token decoding feature in large language models . Preprint, arXiv:2506.06008

  105. [113]

    Ruikang Liu, Yuxuan Sun, Manyi Zhang, Haoli Bai, Xianzhi Yu, Tiezheng Yu, Chun Yuan, and Lu Hou. 2025 d . https://arxiv.org/abs/2504.04823 Quantization hurts reasoning? an empirical study on quantized reasoning models . Preprint, arXiv:2504.04823

  106. [114]

    Shuqi Liu, Han Wu, Bowei He, Xiongwei Han, Mingxuan Yuan, and Linqi Song. 2025 e . https://arxiv.org/abs/2502.12420 Sens-merging: Sensitivity-guided parameter balancing for merging large language models . Preprint, arXiv:2502.12420

  107. [115]

    Tianqiao Liu, Zui Chen, Zitao Liu, Mi Tian, and Weiqi Luo. 2024 c . https://arxiv.org/abs/2409.08561 Expediting and elevating large language model reasoning via hidden chain-of-thought decoding . Preprint, arXiv:2409.08561

  108. [116]

    Wanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin, Ke Ji, Wenyu Chen, Yan Xu, Yasheng Wang, Lifeng Shang, and Benyou Wang. 2025 f . https://arxiv.org/abs/2506.12860 Qfft, question-free fine-tuning for adaptive reasoning . Preprint, arXiv:2506.12860

  109. [117]

    Xin Liu and Lu Wang. 2025. https://arxiv.org/abs/2506.02536 Answer convergence as a signal for early stopping in reasoning . Preprint, arXiv:2506.02536

  110. [118]

    Yongjiang Liu, Haoxi Li, Xiaosong Ma, Jie Zhang, and Song Guo. 2025 g . https://arxiv.org/abs/2507.02663 Think how to think: Mitigating overthinking with autonomous difficulty cognition in large reasoning models . Preprint, arXiv:2507.02663

  111. [119]

    Yue Liu, Jiaying Wu, Yufei He, Hongcheng Gao, Hongyu Chen, Baolong Bi, Jiaheng Zhang, Zhiqi Huang, and Bryan Hooi. 2025 h . https://arxiv.org/abs/2503.23077 Efficient inference for large reasoning models: A survey . Preprint, arXiv:2503.23077

  112. [120]

    Yule Liu, Jingyi Zheng, Zhen Sun, Zifan Peng, Wenhan Dong, Zeyang Sha, Shiwen Cui, Weiqiang Wang, and Xinlei He. 2025 i . https://arxiv.org/abs/2504.13626 Thought manipulation: External thought can be efficient for large reasoning models . Preprint, arXiv:2504.13626

  113. [121]

    Chenwei Lou, Zewei Sun, Xinnian Liang, Meng Qu, Wei Shen, Wenqi Wang, Yuntao Li, Qingping Yang, and Shuangzhi Wu. 2025. https://arxiv.org/abs/2505.11896 Adacot: Pareto-optimal adaptive chain-of-thought triggering via reinforcement learning . Preprint, arXiv:2505.11896

  114. [122]

    Jinghui Lu, Haiyang Yu, Siliang Xu, Shiwei Ran, Guozhi Tang, Siqi Wang, Bin Shan, Teng Fu, Hao Feng, Jingqun Tang, Han Wang, and Can Huang. 2025. https://arxiv.org/abs/2505.15154 Prolonged reasoning is not all you need: Certainty-based adaptive routing for efficient llm/mllm r...

  115. [123]

    Feng Luo, Yu-Neng Chuang, Guanchu Wang, Hoang Anh Duy Le, Shaochen Zhong, Hongyi Liu, Jiayi Yuan, Yang Sui, Vladimir Braverman, Vipin Chaudhary, and Xia Hu. 2025 a . https://arxiv.org/abs/2505.22662 Autol2s: Auto long-short reasoning for efficient large language models . Prepr...

  116. [124]

    Haotian Luo, Haiying He, Yibo Wang, Jinluan Yang, Rui Liu, Naiqiang Tan, Xiaochun Cao, Dacheng Tao, and Li Shen. 2025 b . https://arxiv.org/abs/2504.21659 Ada-r1: Hybrid-cot via bi-level adaptive reasoning optimization . Preprint, arXiv:2504.21659

  117. [125]

    Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao. 2025 c . https://arxiv.org/abs/2501.12570 O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning . Preprint, arXiv:2501.12570

  118. [126]

    Kaijing Ma, Xinrun Du, Yunran Wang, Haoran Zhang, Zhoufutu Wen, Xingwei Qu, Jian Yang, Jiaheng Liu, Minghao Liu, Xiang Yue, Wenhao Huang, and Ge Zhang. 2025 a . https://arxiv.org/abs/2410.06526 Kor-bench: Benchmarking language models on knowledge-orthogonal reasoning tasks . P...

  119. [127]

    Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, and Matei Zaharia. 2025 b . https://arxiv.org/abs/2504.09858 Reasoning models can be effective without thinking . Preprint, arXiv:2504.09858

  120. [128]

    Xinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang, and Xinchao Wang. 2025 c . https://arxiv.org/abs/2502.09601 Cot-valve: Length-compressible chain-of-thought tuning . Preprint, arXiv:2502.09601

  121. [129]

    Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2023. https://arxiv.org/abs/2212.08410 Teaching small language models to reason . Preprint, arXiv:2212.08410

  122. [130]

    Sadegh Mahdavi, Muchen Li, Kaiwen Liu, Christos Thrampoulidis, Leonid Sigal, and Renjie Liao. 2025. https://arxiv.org/abs/2501.14275 Leveraging online olympiad-level math problems for llms training and contamination-resistant evaluation . Preprint, arXiv:2501.14275

  123. [131]

    Zhiting Mei, Christina Zhang, Tenny Yin, Justin Lidard, Ola Shorinwa, and Anirudha Majumdar. 2025. https://arxiv.org/abs/2506.18183 Reasoning about uncertainty: Do reasoning models know when they don't know? Preprint, arXiv:2506.18183

  124. [132]

    Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2021. https://arxiv.org/abs/2106.15772 A diverse corpus for evaluating and developing english math word problem solvers . Preprint, arXiv:2106.15772

  125. [133]

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. https://arxiv.org/abs/1809.02789 Can a suit of armor conduct electricity? a new dataset for open book question answering . Preprint, arXiv:1809.02789

  126. [134]

    MiniMax, :, Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, et al. 2025. https://arxiv.org/abs/2506.13585 Minimax-m1: Scaling test-time compute efficiently with lightning attention . Preprint, arXiv:2506.13585

  127. [135]

    Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. https://arxiv.org/abs/2104.08773 Cross-task generalization via natural language crowdsourcing instructions . Preprint, arXiv:2104.08773

  128. [136]

    Ivan Moshkov, Darragh Hanley, Ivan Sorokin, Shubham Toshniwal, Christof Henkel, Benedikt Schifferer, Wei Du, and Igor Gitman. 2025. Aimo-2 winning solution: Building state-of-the-art mathematical reasoning models with openmathreasoning dataset. arXiv preprint arXiv:2504.16891

  129. [137]

    Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025. https://arxiv.org/abs/2501.19393 s1: Simple test-time scaling . Preprint, arXiv:2501.19393

  130. [138]

    Tergel Munkhbat, Namgyu Ho, Seo Hyun Kim, Yongjin Yang, Yujin Kim, and Se-Young Yun. 2025. https://arxiv.org/abs/2502.20122 Self-training elicits concise reasoning in large language models . Preprint, arXiv:2502.20122

  131. [139]

    Sania Nayab, Giulio Rossolini, Marco Simoni, Andrea Saracino, Giorgio Buttazzo, Nicolamaria Manes, and Fabrizio Giacomelli. 2025. https://arxiv.org/abs/2407.19825 Concise thoughts: Impact of output length on llm reasoning and cost . Preprint, arXiv:2407.19825

  132. [140]

    Yansong Ning, Wei Li, Jun Fang, Naiqiang Tan, and Hao Liu. 2025. https://arxiv.org/abs/2505.11827 Not all thoughts are generated equal: Efficient llm reasoning via multi-turn reinforcement learning . Preprint, arXiv:2505.11827

  133. [141]

    Amin Heyrani Nobari, Kaveh Alimohammadi, Ali ArjomandBigdeli, Akash Srivastava, Faez Ahmed, and Navid Azizan. 2025. https://arxiv.org/abs/2502.02421 Activation-informed merging of large language models . Preprint, arXiv:2502.02421

  134. [142]

    Gonzalez, M Waleed Kadous, and Ion Stoica

    Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025. https://arxiv.org/abs/2406.18665 Routellm: Learning to route llms with preference data . Preprint, arXiv:2406.18665

  135. [143]

    OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, et al. 2024. https://arxiv.org/abs/2412.16720 Openai o1 system card . Preprint, arXiv:2412.16720

  136. [144]

    Li, Aviv Bick, J

    Daniele Paliotta, Junxiong Wang, Matteo Pagliardini, Kevin Y. Li, Aviv Bick, J. Zico Kolter, Albert Gu, François Fleuret, and Tri Dao. 2025. https://arxiv.org/abs/2502.20339 Thinking slow, fast: Scaling inference compute with distilled reasoners . Preprint, arXiv:2502.20339

  137. [145]

    Jiabao Pan, Yan Zhang, Chen Zhang, Zuozhu Liu, Hongwei Wang, and Haizhou Li. 2024. https://arxiv.org/abs/2407.01009 Dynathink: Fast or slow? a dynamic decision-making framework for large language models . Preprint, arXiv:2407.01009

  138. [146]

    Qianjun Pan, Wenkai Ji, Yuyang Ding, Junsong Li, Shilian Chen, Junyi Wang, Jie Zhou, Qin Chen, Min Zhang, Yulan Wu, and Liang He. 2025 a . https://arxiv.org/abs/2505.02665 A survey of slow thinking-based reasoning llms using reinforced learning and inference-time scaling law ....

  139. [147]

    Zhihong Pan, Kai Zhang, Yuze Zhao, and Yupeng Han. 2025 b . https://arxiv.org/abs/2505.19435 Route to reason: Adaptive routing for llm and reasoning strategy selection . Preprint, arXiv:2505.19435

  140. [148]

    Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2024. https://arxiv.org/abs/2312.06681 Steering llama 2 via contrastive activation addition . Preprint, arXiv:2312.06681

  141. [149]

    Shubham Parashar, Blake Olson, Sambhav Khurana, Eric Li, Hongyi Ling, James Caverlee, and Shuiwang Ji. 2025. https://arxiv.org/abs/2502.12521 Inference-time computations for llm reasoning and planning: A benchmark and insights . Preprint, arXiv:2502.12521

  142. [150]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. https://arxiv.org/abs/2103.07191 Are nlp models really able to solve simple math word problems? Preprint, arXiv:2103.07191

  143. [151]

    Jacob Pfau, William Merrill, and Samuel R. Bowman. 2024. https://arxiv.org/abs/2404.15758 Let's think dot by dot: Hidden computation in transformer language models . Preprint, arXiv:2404.15758

  144. [152]

    Penghui Qi, Zichen Liu, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. 2025. https://arxiv.org/abs/2505.13438 Optimizing anytime reasoning via budget relative policy optimization . Preprint, arXiv:2505.13438

  145. [153]

    Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Fandong Meng, Jie Zhou, Ju Ren, and Yaoxue Zhang. 2025. https://arxiv.org/abs/2505.04881 Concise: Confidence-guided compression in step-by-step efficient reasoning . Preprint, arXiv:2505.04881

  146. [154]

    Yanzhao Qin, Tao Zhang, Tao Zhang, Yanjun Shen, Wenjing Luo, Haoze Sun, Yan Zhang, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang, and Bin Cui. 2024. https://arxiv.org/abs/2408.10943 Sysbench: Can large language models follow system messages? Preprint, arXiv:2408.10943

  147. [155]

    Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, Peng Li, Wei Wei, Jing Shao, Chaochao Lu, Yue Zhang, Xian-Sheng Hua, Bowen Zhou, and Yu Cheng. 2025. https://arxiv.org/abs/2503.21614 A survey of efficient...

  148. [156]

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, et al. 2025. https://arxiv.org/abs/2412.15115 Qwen2.5 technical report . Preprint, arXiv:2412.15115

  149. [157]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2023. https://arxiv.org/abs/2311.12022 Gpqa: A graduate-level google-proof q&a benchmark . Preprint, arXiv:2311.12022

  150. [158]

    Matthew Renze and Erhan Guven. 2024. https://doi.org/10.1109/fllm63129.2024.10852493 The benefits of a concise chain of thought on problem-solving in large language models . In 2024 2nd International Conference on Foundation and Large Language Models (FLLM), page 476–483. IEEE

  151. [159]

    Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J. Reddi. 2025. https://arxiv.org/abs/2502.17416 Reasoning with latent thoughts: On the power of looped transformers . Preprint, arXiv:2502.17416

  152. [160]

    ByteDance Seed, :, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, et al. 2025. https://arxiv.org/abs/2504.13914 Seed1.5-thinking: Advancing superb reasoning models with reinforcement learning . Preprint, arXiv:2504.13914

  153. [161]

    Xuan Shen, Yizhou Wang, Xiangxi Shi, Yanzhi Wang, Pu Zhao, and Jiuxiang Gu. 2025 a . https://arxiv.org/abs/2501.19201 Efficient reasoning with hidden thinking . Preprint, arXiv:2501.19201

  154. [162]

    Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi, Wenjing Zhang, Jiangze Yan, Ning Wang, Kai Wang, and Shiguo Lian. 2025 b . https://arxiv.org/abs/2503.04472 Dast: Difficulty-adaptive slow-thinking for large reasoning models . Preprint, arXiv:2503.04472

  155. [163]

    Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He. 2025 c . https://arxiv.org/abs/2502.21074 Codi: Compressing chain-of-thought into continuous space via self-distillation . Preprint, arXiv:2502.21074

  156. [164]

    Leheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao, Changshuo Shen, Yi Zhang, Xiang Wang, and Tat-Seng Chua. 2025. https://arxiv.org/abs/2506.08390 On reasoning strength planning in large reasoning models . Preprint, arXiv:2506.08390

  157. [165]

    Mingyang Song and Mao Zheng. 2025. https://arxiv.org/abs/2505.21178 Walk before you run! concise llm reasoning via reinforcement learning . Preprint, arXiv:2505.21178

  158. [166]

    DiJia Su, Sainbayar Sukhbaatar, Michael Rabbat, Yuandong Tian, and Qinqing Zheng. 2025 a . https://arxiv.org/abs/2410.09918 Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces . Preprint, arXiv:2410.09918

  159. [167]

    DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, and Qinqing Zheng. 2025 b . https://arxiv.org/abs/2502.03275 Token assorted: Mixing latent and text tokens for improved language model reasoning . Preprint, arXiv:2502.03275

  160. [168]

    Jinyan Su and Claire Cardie. 2025. https://arxiv.org/abs/2505.18298 Thinking fast and right: Balancing accuracy and reasoning length with adaptive rewards . Preprint, arXiv:2505.18298

  161. [169]

    Jinyan Su, Jennifer Healey, Preslav Nakov, and Claire Cardie. 2025 c . https://arxiv.org/abs/2505.00127 Between underthinking and overthinking: An empirical study of reasoning length and correctness in llms . Preprint, arXiv:2505.00127

  162. [170]

    Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu. 2025. https://arxiv.org/abs/2503.16419 Stop overthinking: A survey on efficient reasoning for large language models . Preprint, arXiv...

  163. [171]

    Yi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu, Xiangyu Li, Hao Wen, Yizhen Yuan, Huiwen Zheng, Yan Liang, Yuanchun Li, and Yunxin Liu. 2025 a . https://arxiv.org/abs/2504.14350 An empirical study of llm reasoning ability under strict output length constraint . Preprint, arXiv:2504.14350

  164. [172]

    Zhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang, and Jun Xu. 2025 b . https://arxiv.org/abs/2505.12886 Detection and mitigation of hallucination in large reasoning models: A mechanistic perspective . Preprint, arXiv:2505.12886

  165. [173]

    Le, Ed H

    Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2022. https://arxiv.org/abs/2210.09261 Challenging big-bench tasks and whether chain-of-thought can solve them . ...

  166. [174]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://arxiv.org/abs/1811.00937 Commonsenseqa: A question answering challenge targeting commonsense knowledge . Preprint, arXiv:1811.00937

  167. [175]

    Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, and Ruihua Song. 2025. https://arxiv.org/abs/2505.16552 Think silently, think fast: Dynamic latent compression of llm reasoning chains . Preprint, arXiv:2505.16552

  168. [176]

    Siao Tang, Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2025. https://arxiv.org/abs/2506.18810 Concisehint: Boosting efficient reasoning via continuous concise hints during generation . Preprint, arXiv:2506.18810

  169. [177]

    Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei. 2024. https://arxiv.org/abs/2403.02884 Mathscale: Scaling instruction tuning for mathematical reasoning . Preprint, arXiv:2403.02884

  170. [178]

    Core Team, Bingquan Xia, Bowen Shen, Cici, Dawei Zhu, et al. 2025 a . https://arxiv.org/abs/2505.07608 Mimo: Unlocking the reasoning potential of language model -- from pretraining to posttraining . Preprint, arXiv:2505.07608

  171. [179]

    Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, et al. 2025 b . https://arxiv.org/abs/2501.12599 Kimi k1.5: Scaling reinforcement learning with llms . Preprint, arXiv:2501.12599

  172. [180]

    P Team, Xinrun Du, Yifan Yao, Kaijing Ma, Bingli Wang, et al. 2025 c . https://arxiv.org/abs/2502.14739 Supergpqa: Scaling llm evaluation across 285 graduate disciplines . Preprint, arXiv:2502.14739

  173. [181]

    Qwen Team. 2025. https://qwenlm.github.io/blog/qwq-32b/ Qwq-32b: Embracing the power of reinforcement learning

  174. [182]

    Tencent Hunyuan Team, Ao Liu, Botong Zhou, Can Xu, Chayse Zhou, et al. 2025 d . https://arxiv.org/abs/2505.15431 Hunyuan-turbos: Advancing large language models through mamba-transformer synergy and adaptive chain-of-thought . Preprint, arXiv:2505.15431

  175. [183]

    Tencent Hunyuan . 2025. https://github.com/tencent-hunyuan/hunyuan-a13b

  176. [184]

    Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yiping Peng, Yunjie Ji, Han Zhao, and Xiangang Li. 2025. https://arxiv.org/abs/2504.17565 Deepdistill: Enhancing llm reasoning capabilities via large-scale difficulty-graded data training . Preprint, arXiv:2504.17565

  177. [185]

    transluce . 2025. https://transluce.org/investigating-o3-truthfulness

  178. [186]

    Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, Linjing Li, Xiangyuan Lan, and Dongbin Zhao. 2025. https://arxiv.org/abs/2505.10832 Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl . Preprint, arXiv:2505.10832

  179. [187]

    vectara . 2025. https://www.vectara.com/blog/deepseek-r1-hallucinates-more-than-deepseek-v3

  180. [188]

    Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna, and Tianyi Zhou. 2025 a . https://arxiv.org/abs/2506.08343 Wait, we don't need to "wait"! removing thinking tokens improves reasoning efficiency . Preprint, arXiv:2506.08343

  181. [189]

    Jikai Wang, Juntao Li, Jianye Hou, Bowen Yan, Lijun Wu, and Min Zhang. 2025 b . https://arxiv.org/abs/2504.19095 Efficient reasoning for llms through speculative chain-of-thought . Preprint, arXiv:2504.19095

  182. [190]

    Jingyao Wang, Wenwen Qiang, Zeen Song, Changwen Zheng, and Hui Xiong. 2025 c . https://arxiv.org/abs/2505.10425 Learning to think: Information-theoretic reinforcement fine-tuning for llms . Preprint, arXiv:2505.10425

  183. [191]

    Rush, and Tri Dao

    Junxiong Wang, Wen-Ding Li, Daniele Paliotta, Daniel Ritter, Alexander M. Rush, and Tri Dao. 2025 d . https://arxiv.org/abs/2504.10449 M1: Towards scalable test-time compute with mamba reasoning models . Preprint, arXiv:2504.10449

  184. [192]

    Minzheng Wang, Yongbin Li, Haobo Wang, Xinghua Zhang, Nan Xu, Bingli Wu, Fei Huang, Haiyang Yu, and Wenji Mao. 2025 e . https://arxiv.org/abs/2505.02156 Adaptive thinking via mode policy optimization for social language agents . Preprint, arXiv:2505.02156

  185. [193]

    Xiaoqiang Wang, Suyuchen Wang, Yun Zhu, and Bang Liu. 2025 f . https://arxiv.org/abs/2505.18962 System-1.5 reasoning: Traversal in language and latent spaces with dynamic shortcuts . Preprint, arXiv:2505.18962

  186. [194]

    Xinyi Wang, Lucas Caccia, Oleksiy Ostapenko, Xingdi Yuan, William Yang Wang, and Alessandro Sordoni. 2024 a . https://arxiv.org/abs/2310.05707 Guiding language model reasoning with planning tokens . Preprint, arXiv:2310.05707

  187. [195]

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024 b . https://arxiv.org/abs/2406.01574 Mmlu-pro: A more robust ...

  188. [196]

    Yuhang Wang, Youhe Jiang, Bin Cui, and Fangcheng Fu. 2025 g . https://arxiv.org/abs/2505.13326 Thinking short and right over thinking long: Serving llm reasoning efficiently and accurately . Preprint, arXiv:2505.13326

  189. [197]

    Yunhao Wang, Yuhao Zhang, Tinghao Yu, Can Xu, Feng Zhang, and Fengzong Lian. 2025 h . https://arxiv.org/abs/2505.20101 Adaptive deep reasoning: Triggering deep thinking when needed . Preprint, arXiv:2505.20101

  190. [198]

    Bartoldson, Bhavya Kailkhura, and Cihang Xie

    Zijun Wang, Haoqin Tu, Yuhan Wang, Juncheng Wu, Jieru Mei, Brian R. Bartoldson, Bhavya Kailkhura, and Cihang Xie. 2025 i . https://arxiv.org/abs/2504.01903 Star-1: Safer alignment of reasoning llms with 1k data . Preprint, arXiv:2504.01903

  191. [199]

    Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024. https://arxiv.org/abs/2411.04368 Measuring short-form factuality in large language models . Preprint, arXiv:2411.04368

  192. [200]

    Liang Wen, Yunke Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Junchen Liu, Lifu Tang, Xiaowei Lv, Haosheng Zou, Yongchao Deng, Shousheng Jia, and Xiangzheng Zhang. 2025 a . https://arxiv.org/abs/2503.10460 Light-r1: Curriculum sft, dpo and rl for long cot from scrat...

  193. [201]

    Pengcheng Wen, Jiaming Ji, Chi-Min Chan, Juntao Dai, Donghai Hong, Yaodong Yang, Sirui Han, and Yike Guo. 2025 b . https://arxiv.org/abs/2503.12918 Thinkpatterns-21k: A systematic study on the impact of thinking patterns in llms . Preprint, arXiv:2503.12918

  194. [202]

    Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblu...

  195. [203]

    Han Wu, Yuxuan Yao, Shuqi Liu, Zehua Liu, Xiaojin Fu, Xiongwei Han, Xing Li, Hui-Ling Zhen, Tao Zhong, and Mingxuan Yuan. 2025 a . https://arxiv.org/abs/2503.20641 Unlocking efficient long-to-short llm reasoning with model merging . Preprint, arXiv:2503.20641

  196. [204]

    Siye Wu, Jian Xie, Yikai Zhang, Aili Chen, Kai Zhang, Yu Su, and Yanghua Xiao. 2025 b . https://arxiv.org/abs/2505.20258 Arm: Adaptive reasoning model . Preprint, arXiv:2505.20258

  197. [205]

    Yifan Wu, Jingze Shi, Bingheng Wu, Jiayi Zhang, Xiaotian Lin, Nan Tang, and Yuyu Luo. 2025 c . https://arxiv.org/abs/2505.19716 Concise reasoning, big gains: Pruning long reasoning trace with difficulty-aware prompting . Preprint, arXiv:2505.19716

  198. [206]

    Yuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du, Stefanie Jegelka, and Yisen Wang. 2025 d . https://arxiv.org/abs/2502.07266 When more is less: Understanding chain-of-thought length in llms . Preprint, arXiv:2502.07266

  199. [207]

    Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, and Wenjie Li. 2025. https://arxiv.org/abs/2502.12067 Tokenskip: Controllable chain-of-thought compression in llms . Preprint, arXiv:2502.12067

  200. [208]

    Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile, Sang Truong, Chelsea Finn, and Nick Haber. 2025. https://arxiv.org/abs/2506.05256 Just enough thinking: Efficient reasoning with adaptive length penalties reinforcement learning . Preprint, arXiv:2506.05256

  201. [209]

    Wenyi Xiao, Leilei Gan, Weilong Dai, Wanggui He, Ziwei Huang, Haoyuan Li, Fangxun Shu, Zhelun Yu, Peng Zhang, Hao Jiang, and Fei Wu. 2025. https://arxiv.org/abs/2504.18458 Fast-slow thinking for large vision-language model reasoning . Preprint, arXiv:2504.18458

  202. [210]

    Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. 2024. https://arxiv.org/abs/2402.01622 Travelplanner: A benchmark for real-world planning with language agents . Preprint, arXiv:2402.01622

  203. [211]

    Roy Xie, David Qiu, Deepak Gopinath, Dong Lin, Yanchao Sun, Chong Wang, Saloni Potdar, and Bhuwan Dhingra. 2025. https://arxiv.org/abs/2505.19640 Interleaved reasoning for large language models via reinforcement learning . Preprint, arXiv:2505.19640

  204. [212]

    Tao Xiong, Xavier Hu, Wenyan Fan, and Shengyu Zhang. 2025. https://arxiv.org/abs/2507.00606 Mixture of reasonings: Teach large language models to reason with adaptive strategies . Preprint, arXiv:2507.00606

  205. [213]

    Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. 2025 a . https://arxiv.org/abs/2502.18600 Chain of draft: Thinking faster by writing less . Preprint, arXiv:2502.18600

  206. [214]

    Xiaoang Xu, Shuo Wang, Xu Han, Zhenghao Liu, Huijia Wu, Peipei Li, Zhiyuan Liu, Maosong Sun, and Zhaofeng He. 2025 b . https://arxiv.org/abs/2505.24550 A*-thought: Efficient reasoning via bidirectional compression for low-resource settings . Preprint, arXiv:2505.24550

  207. [215]

    Yige Xu, Xu Guo, Zhiwei Zeng, and Chunyan Miao. 2025 c . https://arxiv.org/abs/2502.12134 Softcot: Soft chain-of-thought for efficient reasoning with llms . Preprint, arXiv:2502.12134

  208. [216]

    Yuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo, Junnan Li, and Caiming Xiong. 2025 d . https://arxiv.org/abs/2505.05315 Scalable chain of thoughts via elastic reasoning . Preprint, arXiv:2505.05315

  209. [217]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, et al. 2025 a . https://arxiv.org/abs/2505.09388 Qwen3 technical report . Preprint, arXiv:2505.09388

  210. [218]

    Chenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu, Chenyu Zhu, Qiaowei Li, Zheng Lin, Li Cao, and Weiping Wang. 2025 b . https://arxiv.org/abs/2504.15895 Dynamic early exit in reasoning models . Preprint, arXiv:2504.15895

  211. [219]

    Junjie Yang, Ke Lin, and Xing Yu. 2025 c . https://arxiv.org/abs/2504.03234 Think when you need: Self-adaptive chain-of-thought learning . Preprint, arXiv:2504.03234

  212. [220]

    Shu Yang, Junchao Wu, Xuansheng Wu, Derek Wong, Ninhao Liu, and Di Wang. 2025 d . https://arxiv.org/abs/2506.19492 Is long-to-short a free lunch? investigating inconsistency and reasoning efficiency in lrms . Preprint, arXiv:2506.19492

  213. [221]

    Wenkai Yang, Shuming Ma, Yankai Lin, and Furu Wei. 2025 e . https://arxiv.org/abs/2502.18080 Towards thinking-optimal scaling of test-time compute for llm reasoning . Preprint, arXiv:2502.18080

  214. [222]

    Cohen, Ruslan Salakhutdinov, and Christopher D

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://arxiv.org/abs/1809.09600 Hotpotqa: A dataset for diverse, explainable multi-hop question answering . Preprint, arXiv:1809.09600

  215. [223]

    Wenlin Yao, Haitao Mi, and Dong Yu. 2024. https://arxiv.org/abs/2409.17433 Hdflow: Enhancing llm complex problem-solving with hybrid thinking and dynamic workflows . Preprint, arXiv:2409.17433

  216. [224]

    Yuxuan Yao, Shuqi Liu, Zehua Liu, Qintong Li, Mingyang Liu, Xiongwei Han, Zhijiang Guo, Han Wu, and Linqi Song. 2025 a . https://arxiv.org/abs/2505.14009 Activation-guided consensus merging for large language models . Preprint, arXiv:2505.14009

  217. [225]

    Zijun Yao, Yantao Liu, Yanxu Chen, Jianhui Chen, Junfeng Fang, Lei Hou, Juanzi Li, and Tat-Seng Chua. 2025 b . https://arxiv.org/abs/2505.23646 Are reasoning models more prone to hallucination? Preprint, arXiv:2505.23646

  218. [226]

    Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue. 2025. https://arxiv.org/abs/2502.03373 Demystifying long chain-of-thought reasoning in llms . Preprint, arXiv:2502.03373

  219. [227]

    Jingyang Yi, Jiazheng Wang, and Sida Li. 2025. https://arxiv.org/abs/2504.21370 Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning . Preprint, arXiv:2504.21370

  220. [228]

    Xixian Yong, Xiao Zhou, Yingying Zhang, Jinlin Li, Yefeng Zheng, and Xian Wu. 2025. https://arxiv.org/abs/2505.18237 Think or not? exploring thinking efficiency in large reasoning models via an information-theoretic lens . Preprint, arXiv:2505.18237

  221. [229]

    Bin Yu, Hang Yuan, Haotian Li, Xueyin Xu, Yuliang Wei, Bailing Wang, Weizhen Qi, and Kai Chen. 2025 a . https://arxiv.org/abs/2505.03469 Long-short chain-of-thought mixture supervised fine-tuning eliciting efficient reasoning in large language models . Preprint, arXiv:2505.03469

  222. [230]

    Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2024 a . https://arxiv.org/abs/2311.03099 Language models are super mario: Absorbing abilities from homologous models as a free lunch . Preprint, arXiv:2311.03099

  223. [231]

    Ping Yu, Jing Xu, Jason Weston, and Ilia Kulikov. 2024 b . https://arxiv.org/abs/2407.06023 Distilling system 2 into system 1 . Preprint, arXiv:2407.06023

  224. [232]

    Qifan Yu, Zhenyu He, Sijie Li, Xun Zhou, Jun Zhang, Jingjing Xu, and Di He. 2025 b . https://arxiv.org/abs/2502.08482 Enhancing auto-regressive chain-of-thought through loop-aligned reasoning . Preprint, arXiv:2502.08482

  225. [233]

    Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, et al. 2025 c . https://arxiv.org/abs/2503.14476 Dapo: An open-source llm reinforcement learning system at scale . Preprint, arXiv:2503.14476

  226. [234]

    Danlong Yuan, Tian Xie, Shaohan Huang, Zhuocheng Gong, Huishuai Zhang, Chong Luo, Furu Wei, and Dongyan Zhao. 2025. https://arxiv.org/abs/2505.12284 Efficient rl training for reasoning models via length-aware optimization . Preprint, arXiv:2505.12284

  227. [235]

    Zhiyuan Zeng, Qinyuan Cheng, Zhangyue Yin, Yunhua Zhou, and Xipeng Qiu. 2025 a . https://arxiv.org/abs/2502.12215 Revisiting the test-time scaling of o1-like models: Do they truly possess test-time scaling capabilities? Preprint, arXiv:2502.12215

  228. [236]

    Zihao Zeng, Xuyao Huang, Boxiu Li, Hao Zhang, and Zhijie Deng. 2025 b . https://arxiv.org/abs/2505.19788 Done is better than perfect: Unlocking efficient reasoning by structured multi-turn decomposition . Preprint, arXiv:2505.19788

  229. [237]

    Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, and He He. 2025 a . https://arxiv.org/abs/2504.05419 Reasoning models know when they're right: Probing hidden states for self-verification . Preprint, arXiv:2504.05419

  230. [238]

    Chong Zhang, Yue Deng, Xiang Lin, Bin Wang, Dianwen Ng, Hai Ye, Xingxuan Li, Yao Xiao, Zhanfeng Mo, Qi Zhang, and Lidong Bing. 2025 b . https://arxiv.org/abs/2505.00551 100 days after deepseek-r1: A survey on replication studies and more directions for reasoning language model...

  231. [239]

    Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, and Juanzi Li. 2025 c . https://arxiv.org/abs/2505.13417 Adaptthink: Reasoning models can learn when to think . Preprint, arXiv:2505.13417

  232. [240]

    Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, and Ningyu Zhang. 2025 d . https://arxiv.org/abs/2502.15589 Lightthinker: Thinking step-by-step compression . Preprint, arXiv:2502.15589

  233. [241]

    Junyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, and Huan Zhang. 2025 e . https://arxiv.org/abs/2505.24863 Alphaone: Reasoning models thinking slow and fast at test time . Preprint, arXiv:2505.24863

  234. [242]

    Ruiqi Zhang, Changyi Xiao, and Yixin Cao. 2025 f . https://arxiv.org/abs/2506.04182 Long or short cot? investigating instance-level switch of large reasoning models . Preprint, arXiv:2506.04182

  235. [243]

    Shengjia Zhang, Junjie Wu, Jiawei Chen, Changwang Zhang, Xingyu Lou, Wangchunshu Zhou, Sheng Zhou, Can Wang, and Jun Wang. 2025 g . https://arxiv.org/abs/2506.02397 Othink-r1: Intrinsic fast/slow thinking mode switching for over-reasoning mitigation . Preprint, arXiv:2506.02397

  236. [244]

    Xiaoyun Zhang, Jingqing Ruan, Xing Ma, Yawen Zhu, Haodong Zhao, Hao Li, Jiansong Chen, Ke Zeng, and Xunliang Cai. 2025 h . https://arxiv.org/abs/2505.15400 When to continue thinking: Adaptive thinking mode switching for efficient reasoning . Preprint, arXiv:2505.15400

  237. [245]

    Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, and Minlie Huang. 2025 i . https://arxiv.org/abs/2505.15404 How should we enhance the safety of large reasoning models: An empirical study ....

  238. [246]

    Han Zhao, Haotian Wang, Yiping Peng, Sitong Zhao, Xiaoyu Tian, Shuaiting Chen, Yunjie Ji, and Xiangang Li. 2025 a . https://arxiv.org/abs/2503.19633 1.4 million open-source distilled reasoning dataset to empower large language model training . Preprint, arXiv:2503.19633

  239. [247]

    Haoran Zhao, Yuchen Yan, Yongliang Shen, Haolei Xu, Wenqi Zhang, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, and Yueting Zhuang. 2025 b . https://arxiv.org/abs/2505.14604 Let llms break free from overthinking via self-braking tuning . Preprint, arXiv:2505.14604

  240. [248]

    Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che, Tat-Seng Chua, and Ting Liu. 2025 c . https://arxiv.org/abs/2503.17979 Trade-offs in large reasoning models: An empirical analysis of deliberative and adaptive reasoning over founda...

  241. [249]

    Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, and Xian Wu. 2025 d . https://arxiv.org/abs/2505.17427 T ^2 : An adaptive test-time scaling strategy for contextual question answering . Preprint, arXiv:2505.17427

  242. [250]

    Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. 2022. https://arxiv.org/abs/2109.00110 Minif2f: a cross-system benchmark for formal olympiad-level mathematics . Preprint, arXiv:2109.00110

  243. [251]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. https://arxiv.org/abs/2306.05685 Judging llm-as-a-judge with mt-bench and chatbot arena . P...

  244. [252]

    Rongzhi Zhu, Yi Liu, Zequn Sun, Yiwei Wang, and Wei Hu. 2025 a . https://arxiv.org/abs/2505.15276 When can large reasoning models save thinking? mechanistic analysis of behavioral divergence in reasoning . Preprint, arXiv:2505.15276

  245. [253]

    Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng, Xingwei Qu, Jinfa Huang, Dawei Zhu, Hao Wang, Kaiwen Xue, Xuanliang Zhang, Yong Shan, Tianle Cai, Taylor Kergan, Assel Kembay, Andrew Smith, Chenghua Lin, Binh Nguyen, Yuqi Pan, Yuhong Chou, Zefan Cai, Zhenhe Wu, Yongchi Zhao, Tianyu L...

  246. [254]

    Zihao Zhu, Hongbao Zhang, Ruotong Wang, Ke Xu, Siwei Lyu, and Baoyuan Wu. 2025 c . https://arxiv.org/abs/2502.12202 To think or not to think: Exploring the unthinking vulnerability in large reasoning models . Preprint, arXiv:2502.12202

  247. [255]

    Ren Zhuang, Ben Wang, and Shuifa Sun. 2025. https://arxiv.org/abs/2505.08392 Accelerating chain-of-thought reasoning: When goal-gradient importance meets dynamic skipping . Preprint, arXiv:2505.08392

  248. [256]

    Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, et al. 2025. https://arxiv.org/abs/2406.15877 Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions . Preprint, arXiv:2406.15877

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.