REVIEW 3 major objections 5 minor 4 cited by
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Survey argues that overthinking in large reasoning models is solvable and maps two agendas — concise thinking and adaptive thinking — for doing it.
desk verdict A genuinely useful survey of concise and adaptive thinking in LRMs, but its core tables contain verifiable errors that need an audit before the paper can serve as a reliable reference. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing structure of the survey is its taxonomy. Methods are split into training-free approaches (prompt-guided strategies, pipeline and router systems, decoding manipulation such as budget forcing, early exiting, logit adjustment, and activation steering, plus model merging) and training-based approaches (variable-length data construction, fine-tuning including implicit and DPO-variant methods, and reinforcement learning with length penalties, GRPO variants, difficulty awareness, and thinking-mode selection). Alongside the taxonomy, the survey assembles a set of efficiency metrics — token counts, latency, speed-up ratio, $E^3 = A^2/T$, Think Density, HCA/SCA/CCA, InfoBias and InfoGain among others — that make the field's goal of accuracy parity with fewer tokens measurable and comparable.
What would settle it
A reader can audit the paper's own tables against its reference list: Table 3 attributes C3oT to Luo et al. (2025b) while the text and bibliography cite Kang et al. (2024). A broader audit of this kind, combined with a uniform reproduction of the cited methods on a fixed benchmark set (for example GSM8K, MATH-500, and AIME24 against a shared long-CoT baseline), would settle whether the taxonomy's attributions are reliable and whether the field's accuracy-parity-with-shorter-chains premise actually holds.
Extended reading notes
Core claim
The authors' central claim is that the overthinking behavior of large reasoning models is not an unavoidable side effect of their reasoning ability; it can be reduced, and a growing body of recent work already reduces it through two complementary agendas. Concise thinking keeps one reasoning chain but makes it shorter, while adaptive thinking switches between fast direct answers and slow chains based on input difficulty. The survey further claims that the empirical pattern motivating these efforts is real: accuracy does not keep improving with chain length, and different tasks have different optimal reasoning lengths. It supports the map with observations such as RL-trained models frequently reasoning even when explicitly instructed to skip thinking, whereas SFT-based models follow concise-thinking instructions more obediently, which means methods do not transfer uniformly across model families.
Load-bearing premise
The survey's value as a map depends on its taxonomy and citations being accurate and complete; if methods are misattributed or categories misrepresent the literature, readers navigating by this guide will land on the wrong papers or wrong distinctions.
Editorial extensions
If this is right
- Practitioners can pick an approach by their training budget: prompt engineering and decoding tricks give immediate, training-free savings, while fine-tuning and reinforcement learning produce more persistent changes in reasoning behavior.
- Because RL-based and SFT-based models respond differently to no-thinking instructions, an efficiency method validated on one model family cannot be assumed to transfer to another; evaluations should name the base model type.
- The wide spread of evaluation metrics and benchmarks means no single number yet captures the accuracy-efficiency trade-off; until the field standardizes base models, datasets, and hyperparameters, cross-paper comparisons stay provisional.
- If the overthinking premise holds, product builders can cut inference cost and latency on simple queries, which is the main practical payoff of the surveyed line of work.
- Defining what counts as a good reasoning chain must include interpretability and safety, not just length, because aggressive compression can remove exactly the transparency that makes reasoning models valuable to users.
Reading between the lines
- A maintainable version of this survey would be a machine-checkable table with one verified row per method; the visible mismatch in Table 3 — C3oT attributed to Luo et al. (2025b) while the text cites Kang et al. (2024) — suggests the current map is not yet that.
- Router-based methods point toward a natural product pattern: a cheap classifier sends easy queries to fast mode and hard ones to slow mode. An implicit extension, not designed in the paper, is to make the router preference-aware — for example, always showing full reasoning for medical or legal answers even when the query looks easy.
- The $E^3 = A^2/T$ metric implies an accuracy-versus-token Pareto frontier for each dataset; plotting methods that way would let users select an operating point, which the survey itself does not do.
- A testable cross-family question follows from the routing work: does a capability-aware router trained on one reasoning model family (say DeepSeek-R1-Distill) generalize to another family (say QwQ)? The paper lists the ingredients but leaves the transfer question open.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper surveys recent work on making large reasoning models (LRMs) more efficient by producing concise or adaptively selected reasoning chains. The survey organizes methods into training-free approaches (prompt-guided, pipeline, decoding manipulation, model merging) and training-based approaches (fine-tuning and reinforcement learning), and it also reviews evaluation benchmarks, metrics, and open challenges. The intended contribution is a comprehensive, up-to-date guide to the field of concise and adaptive thinking for LRMs.
Significance. The survey addresses a timely and practically important problem: reducing the excessive inference cost and latency of LRMs while preserving reasoning accuracy. Its taxonomy is broadly sensible, and it covers a wide range of recent works, including many 2024–2025 preprints. The paper is explicitly scoped and does not attempt to cover all of efficient inference, which is appropriate. Its value as a reference, however, depends on the accuracy of its tables and attributions. The concrete errors identified below (notably the C3oT misattribution in Table 3 and the ambiguous Ling et al. reward in Table 5) undermine confidence in the survey as a reliable indexing tool. These issues are correctable, and the core organizational framework of the survey is a contribution in itself.
major comments (3)
- [Table 3 vs. Section 4.1.1] Table 3 lists 'C3oT (Luo et al., 2025b)' with base models LLaMA2-Chat-7B/13B and datasets ECQA/StrategyQA, but the text in Section 4.1.1 and the reference list identify C3oT as Kang et al. (2024), and Figure 1 also cites Kang et al. (2024). Luo et al. (2025b) is Ada-R1 elsewhere in the survey. This misattribution directs a reader who uses Table 3 as an index to the wrong paper and must be corrected.
- [Table 5, Ling et al. (2025) row] The printed reward function for Ling et al. is 1 + αL(y)^γ for correct answers. For positive α and γ, this assigns a larger reward to longer correct chains, which is the opposite of a length penalty. Section 4.2.1 describes Ling et al. as adopting a 'powered length penalty (PLP) that is more lenient on longer responses,' implying a penalty that should reduce reward as length grows. The table must either explicitly state that α < 0, show a minus sign, or otherwise reconcile the formula with the textual description; as printed, the table is internally inconsistent and misleading about the method.
- [Table 4] The reliability of Table 4 as a guide to open-source variable-length CoT datasets needs verification and correction. The OmniThought sample count appears as '2000,000,' which is a typo for 2,000,000. Additionally, 's1K-mix' is attributed to Yu et al. (2025a), while Section 4.1.2 introduces s1K under Muennighoff et al. (2025); the relationship between s1K and s1K-mix, and the correct attribution, should be clarified. Since the survey's central claim is to be a comprehensive and accurate guide, the tables require a careful audit.
minor comments (5)
- [Tables 2 and 3] The dataset name 'SV AMP' appears in Tables 2 and 3; the correct name is SVAMP (Patel et al., 2021), and the formatting should be consistent.
- [Table 3, DeGRPO row] The dataset 'MATH-50' in the DeGRPO row appears to be a typo for MATH-500.
- [Sections 2.3 and 4.1.2] The metric from Cai et al. (2025) is called 'Reasoning Verbosity' in Section 2.3 and 'Reasoning Veracity (RV)' in Section 4.1.2; the survey should use one consistent term.
- [Section 2.3, E^3 metric] The formula for E^3 is written as 'A2T' without formatting; it should be clearly displayed as A^2 divided by T to avoid ambiguity.
- [Table 5 caption] The caption defines S(y) and L(y), but the table uses additional symbols (α, γ, λ, L_budget, L_cache, L_max, etc.) that are not defined in the caption. Adding definitions or referencing the original equations would improve readability.
Circularity Check
No circularity: the survey derives no quantities, fits no parameters, and contains no load-bearing self-citations.
full rationale
This paper is a literature survey rather than a derivation-based research contribution. Its central claim is that it provides a comprehensive overview of concise and adaptive thinking in large reasoning models, organized into training-free and training-based taxonomies. There are no fitted parameters, no predictive equations derived from inputs, and no claimed first-principles results that could reduce to their own assumptions. The authors do not cite their own prior work in any load-bearing way, and the taxonomy and narrative summaries are external descriptive claims about cited papers rather than circular derivations. The concrete accuracy problems noted by the reader, such as the C3oT attribution in Table 3 pointing to Luo et al. (2025b) while the text and references identify Kang et al. (2024), and the Table 5 display of the Ling et al. (2025) reward as 1 + alpha*L(y)^gamma, which would reward longer correct answers rather than penalize them, are real scholarly-quality concerns. However, misattribution and formula transcription errors are not circularity: the survey is not defining its conclusions in terms of its premises, nor is it renaming a fitted input as a prediction. Because the paper makes no internally derived quantitative claims, there is no circular reasoning chain to expose, and the appropriate circularity score is 0.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey." pith.science (2026). https://pith.science/paper/I2ABVH5K
@misc{pith2026250709662,
author = {Pith},
title = {Pith review of: Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/I2ABVH5K}},
note = {Machine review of arXiv:2507.09662}
}
read the original abstract
Large reasoning models (LRMs) like OpenAI o1 and DeepSeek R1 have demonstrated impressive performance on complex reasoning tasks like mathematics and programming with long Chain-of-Thought (CoT) reasoning sequences (slow-thinking), compared with traditional large language models (fast-thinking). However, these reasoning models also face a huge challenge that generating unnecessarily lengthy and redundant reasoning chains even for trivial questions. This phenomenon leads to a significant waste of inference resources, increases the response time for simple queries, and hinders the practical application of LRMs in real-world products. To this end, it is crucial to shorten lengthy reasoning chains and learn adaptive reasoning between fast and slow thinking based on input difficulty. In this survey, we provide a comprehensive overview of recent progress in concise and adaptive thinking for efficient reasoning of LRMs, including methodologies, benchmarks, and challenges for future exploration. We hope this survey can help researchers quickly understand the landscape of this field and inspire novel adaptive thinking ideas to facilitate better usage of LRMs.
Figures
Forward citations
Cited by 4 Pith papers
-
Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation
Consequence-aware scheduler using an issue-text predictor routes more compute to high-cost failures and cuts cost-weighted loss by 22-33% versus difficulty-based allocation on SWE-bench tasks.
-
ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
ParaThinker trains LLMs for native parallel reasoning and reports 7 to 12 percent higher accuracy on math benchmarks over sequential thinking with modest latency overhead.
-
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
A control-token insertion and two-stage training method that lets LLMs adhere to user-specified reasoning token budgets while preserving math accuracy.
-
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
A survey that maps reinforcement learning methods, datasets, benchmarks, and open-source tools across the full training lifecycle of large language models, focusing on verifiable-reward reasoning.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Aradhye Agarwal, Ayan Sengupta, and Tanmoy Chakraborty. 2025. https://arxiv.org/abs/2505.18149 First finish search: Efficient test-time scaling in large language models . Preprint, arXiv:2505.18149
arXiv 2025
-
[4]
Pranjal Aggarwal and Sean Welleck. 2025. https://arxiv.org/abs/2503.04697 L1: Controlling how long a reasoning model thinks with reinforcement learning . Preprint, arXiv:2503.04697
arXiv 2025
-
[5]
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021. https://doi.org/10.18653/v1/2021.acl-long.238 E xplanations for C ommonsense QA : N ew D ataset and M odels . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferen...
-
[6]
Syeda Nahida Akter, Shrimai Prabhumoye, Matvei Novikov, Seungju Han, Ying Lin, Evelina Bakhturina, Eric Nyberg, Yejin Choi, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2025. https://arxiv.org/abs/2504.13941 Nemotron-crossthink: Scaling self-learning beyond math reasoning . Preprint, arXiv:2504.13941
arXiv 2025
-
[7]
Mohammad Ali Alomrani, Yingxue Zhang, Derek Li, Qianyi Sun, Soumyasundar Pal, Zhanguang Zhang, Yaochen Hu, Rohan Deepak Ajwani, Antonios Valkanas, Raika Karimi, Peng Cheng, Yunzhou Wang, Pengyi Liao, Hanrui Huang, Bin Wang, Jianye Hao, and Mark Coates. 2025. https://arxiv.org/abs/2507.02076 Reasoning on a budget: A survey of adaptive and controllable test...
arXiv 2025
-
[8]
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. https://arxiv.org/abs/1905.13319 Mathqa: Towards interpretable math word problem solving with operation-based formalisms . Preprint, arXiv:1905.13319
arXiv 2019
Show all 256 references
-
[9]
Sohyun An, Ruochen Wang, Tianyi Zhou, and Cho-Jui Hsieh. 2025. https://arxiv.org/abs/2505.21765 Don't think longer, think wisely: Optimizing thinking dynamics for large reasoning models . Preprint, arXiv:2505.21765
2025 arXiv
-
[10]
anthropic . 2025. https://www.anthropic.com/news/claude-3-7-sonnet
2025
-
[11]
Daman Arora and Andrea Zanette. 2025. https://arxiv.org/abs/2502.04463 Training language models to reason efficiently . Preprint, arXiv:2502.04463
2025
-
[12]
Dhananjay Ashok and Jonathan May. 2025. https://arxiv.org/abs/2502.13329 Language models can predict their own behavior . Preprint, arXiv:2502.13329
2025
-
[13]
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. https://arxiv.org/abs/2108.07732 Program synthesis with large language models . Preprint, arXiv:2108.07732
2021 arXiv
-
[14]
Aytes, Jinheon Baek, and Sung Ju Hwang
Simon A. Aytes, Jinheon Baek, and Sung Ju Hwang. 2025. https://arxiv.org/abs/2503.05179 Sketch-of-thought: Efficient llm reasoning with adaptive cognitive-inspired sketching . Preprint, arXiv:2503.05179
2025
-
[15]
Ayers, Dragomir Radev, and Jeremy Avigad
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, and Jeremy Avigad. 2023. https://arxiv.org/abs/2302.12433 Proofnet: Autoformalizing and formally proving undergraduate-level mathematics . Preprint, arXiv:2302.12433
2023 arXiv
-
[16]
Mislav Balunović, Jasper Dekoninck, Ivo Petrov, Nikola Jovanović, and Martin Vechev. 2025. https://arxiv.org/abs/2505.23281 Matharena: Evaluating llms on uncontaminated math competitions . Preprint, arXiv:2505.23281
2025 arXiv
-
[17]
Akhiad Bercovich, Itay Levy, Izik Golan, Mohammad Dabbah, et al. 2025. https://arxiv.org/abs/2505.00949 Llama-nemotron: Efficient reasoning models . Preprint, arXiv:2505.00949
2025
-
[18]
Bytedance Seed Team . 2025. https://www.volcengine.com/
2025
-
[19]
Wenrui Cai, Chengyu Wang, Junbing Yan, Jun Huang, and Xiangzhong Fang. 2025. https://arxiv.org/abs/2505.10937 Reasoning with omnithought: A large cot dataset with verbosity and cognitive difficulty annotations . Preprint, arXiv:2505.10937
2025 arXiv
-
[20]
Hanting Chen, Yasheng Wang, Kai Han, Dong Li, Lin Li, Zhenni Bi, Jinpeng Li, Haoyu Wang, Fei Mi, Mingjian Zhu, Bin Wang, Kaikai Song, Yifei Fu, Xu He, Yu Luo, Chong Zhu, Quan He, Xueyu Wu, Wei He, Hailin Hu, Yehui Tang, Dacheng Tao, Xinghao Chen, and Yunhe Wang. 2025 a . https...
2025 arXiv
-
[21]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, et al. 2021. https://arxiv.org/abs/2107.03374 Evaluating large language models trained on code . Preprint, arXiv:2107.03374
2021 arXiv
-
[22]
Weize Chen, Jiarui Yuan, Tailin Jin, Ning Ding, Huimin Chen, Zhiyuan Liu, and Maosong Sun. 2025 b . https://arxiv.org/abs/2505.19217 The overthinker's diet: Cutting token calories with difficulty-aware training . Preprint, arXiv:2505.19217
2025 arXiv
-
[23]
Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu. 2025 c . https://arxiv.org/abs/2412.21187 Do not think that much for 2+3=? on the overthinking of o1-li...
2025 arXiv
-
[24]
Zigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu, and Xinchao Wang. 2025 d . https://arxiv.org/abs/2505.17941 Verithinker: Learning to verify makes reasoning model efficient . Preprint, arXiv:2505.17941
2025 arXiv
-
[25]
Jeffrey Cheng and Benjamin Van Durme. 2024. https://arxiv.org/abs/2412.13171 Compressed chain of thought: Efficient reasoning through dense representations . Preprint, arXiv:2412.13171
2024 arXiv
-
[26]
Xiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang, Wayne Xin Zhao, Xinyu Kong, and Zhiqiang Zhang. 2025 a . https://arxiv.org/abs/2505.16315 Incentivizing dual process thinking for efficient large language model reasoning . Preprint, arXiv:2505.16315
2025 arXiv
-
[27]
Xiaoxue Cheng, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. 2025 b . https://arxiv.org/abs/2501.01306 Think more, hallucinate less: Mitigating hallucinations via dual process of fast and slow thinking . Preprint, arXiv:2501.01306
2025 arXiv
-
[28]
Zhengxiang Cheng, Dongping Chen, Mingyang Fu, and Tianyi Zhou. 2025 c . https://arxiv.org/abs/2506.14755 Optimizing length compression in large reasoning models . Preprint, arXiv:2506.14755
2025 arXiv
-
[29]
Stephen Chung, Wenyu Du, and Jie Fu. 2025. https://arxiv.org/abs/2505.21097 Thinker: Learning to think fast and slow . Preprint, arXiv:2505.21097
2025
-
[30]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. https://arxiv.org/abs/1803.05457 Think you have solved question answering? try arc, the ai2 reasoning challenge . Preprint, arXiv:1803.05457
2018 arXiv
-
[31]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://arxiv.org/abs/2110.14168 Training verifiers to solve math word problems . ...
2021 arXiv
-
[32]
Cognitive Computations. 2025. https://huggingface.co/datasets/cognitivecomputations/dolphin-r1
2025
-
[33]
Gonzalez
Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, Nicholas Thumiger, Aditya Desai, Ion Stoica, Ana Klimovic, Graham Neubig, and Joseph E. Gonzalez. 2025. https://arxiv.org/abs/2502.08235 ...
2025 arXiv
-
[34]
Renfei Dang, Shujian Huang, and Jiajun Chen. 2025. https://arxiv.org/abs/2505.16448 Internal bias in reasoning models leads to overthinking . Preprint, arXiv:2505.16448
2025
-
[35]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, et al. 2025. https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning . Preprint, arXiv:2501.12948
2025 arXiv
-
[36]
Hexuan Deng, Wenxiang Jiao, Xuebo Liu, Jun Rao, and Min Zhang. 2025. https://arxiv.org/abs/2505.19862 Rea-rl: Reflection-aware online reinforcement learning for efficient large reasoning models . Preprint, arXiv:2505.19862
2025
-
[37]
Yuntian Deng, Yejin Choi, and Stuart Shieber. 2024. https://arxiv.org/abs/2405.14838 From explicit cot to implicit cot: Learning to internalize cot step by step . Preprint, arXiv:2405.14838
2024 arXiv
-
[38]
Bowen Ding, Yuhan Chen, Futing Wang, Lingfeng Ming, and Tao Lin. 2025. https://arxiv.org/abs/2506.23840 Do thinking tokens help or trap? towards more efficient large reasoning model . Preprint, arXiv:2506.23840
2025 arXiv
-
[39]
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Ruhle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah. 2024. https://arxiv.org/abs/2404.14618 Hybrid llm: Cost-efficient and quality-aware query routing . Preprint, arXiv:2404.14618
2024 arXiv
-
[40]
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. https://arxiv.org/abs/1903.00161 Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs . Preprint, arXiv:1903.00161
2019 arXiv
-
[41]
Hashimoto
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2025. https://arxiv.org/abs/2404.04475 Length-controlled alpacaeval: A simple way to debias automatic evaluators . Preprint, arXiv:2404.04475
2025 arXiv
-
[42]
Razvan-Gabriel Dumitru, Darius Peteleaza, Vikas Yadav, and Liangming Pan. 2025. https://arxiv.org/abs/2505.17250 Conciserl: Conciseness-guided reinforcement learning for efficient reasoning models . Preprint, arXiv:2505.17250
2025 arXiv
-
[43]
Roy Eisenstadt, Itamar Zimerman, and Lior Wolf. 2025. https://arxiv.org/abs/2506.07240 Overclocking llm reasoning: Monitoring and controlling thinking path lengths in llms . Preprint, arXiv:2506.07240
2025 arXiv
-
[44]
Chenrui Fan, Ming Li, Lichao Sun, and Tianyi Zhou. 2025 a . https://arxiv.org/abs/2504.06514 Missing premise exacerbates overthinking: Are reasoning models losing critical thinking skill? Preprint, arXiv:2504.06514
2025 arXiv
-
[45]
Siqi Fan, Peng Han, Shuo Shang, Yequan Wang, and Aixin Sun. 2025 b . https://arxiv.org/abs/2505.22017 Cothink: Token-efficient reasoning via instruct models guiding reasoning models . Preprint, arXiv:2505.22017
2025
-
[46]
Gongfan Fang, Xinyin Ma, and Xinchao Wang. 2025. https://arxiv.org/abs/2505.13379 Thinkless: Llm learns when to think . Preprint, arXiv:2505.13379
2025 arXiv
-
[47]
Mehdi Fatemi, Banafsheh Rafiee, Mingjie Tang, and Kartik Talamadupula. 2025. https://arxiv.org/abs/2504.05185 Concise reasoning via reinforcement learning . Preprint, arXiv:2504.05185
2025
-
[48]
Sicheng Feng, Gongfan Fang, Xinyin Ma, and Xinchao Wang. 2025. https://arxiv.org/abs/2504.10903 Efficient reasoning models: A survey . Preprint, arXiv:2504.10903
2025
-
[49]
Tianyu Fu, Yi Ge, Yichen You, Enshu Liu, Zhihang Yuan, Guohao Dai, Shengen Yan, Huazhong Yang, and Yu Wang. 2025 a . https://arxiv.org/abs/2505.21600 R2r: Efficiently navigating divergent reasoning paths with small-large model token routing . Preprint, arXiv:2505.21600
2025
-
[50]
Tingchen Fu, Jiawei Gu, Yafu Li, Xiaoye Qu, and Yu Cheng. 2025 b . https://arxiv.org/abs/2505.14810 Scaling reasoning, losing control: Evaluating instruction following in large reasoning models . Preprint, arXiv:2505.14810
2025 arXiv
-
[51]
Yichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu, Zhongdongming Dai, Yonghao Zhuang, Yian Ma, Aurick Qiao, Tajana Rosing, Ion Stoica, and Hao Zhang. 2025 c . https://arxiv.org/abs/2412.20993 Efficiently scaling llm reasoning with certaindex . Preprint, arXiv:2412.20993
2025 arXiv
-
[52]
Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. 2024. https://arxiv.org...
2024 arXiv
-
[53]
Jiaxuan Gao, Shu Yan, Qixin Tan, Lu Yang, Shusheng Xu, Wei Fu, Zhiyu Mei, Kaifeng Lyu, and Yi Wu. 2025. https://arxiv.org/abs/2506.07104 How far are we from optimal reasoning efficiency? Preprint, arXiv:2506.07104
2025 arXiv
-
[54]
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. https://arxiv.org/abs/2101.02235 Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies . Preprint, arXiv:2101.02235
2021 arXiv
-
[55]
Soumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu, Mengdi Wang, Dinesh Manocha, Furong Huang, Mohammad Ghavamzadeh, and Amrit Singh Bedi. 2025. https://arxiv.org/abs/2506.04210 Does thinking more always help? understanding test-time scaling in reasoning models ...
2025
-
[56]
Ruihan Gong, Yue Liu, Wenjie Qu, Mingzhe Du, Yufei He, Yingwei Ma, Yulin Chen, Xiang Liu, Yi Wen, Xinfeng Li, Ruidong Wang, Xinzhong Zhu, Bryan Hooi, and Jiaheng Zhang. 2025. https://arxiv.org/abs/2505.19756 Efficient reasoning via chain of unconscious thought . Preprint, arXi...
2025 arXiv
-
[57]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. 2024. https://arxiv.org/abs/2407.21783 The llama 3 herd of models . Preprint, arXiv:2407.21783
2024 arXiv
-
[58]
Merrill, Tatsunori Hashimoto, Yejin Choi, Jenia Jitsev, Reinhard Heckel, Maheswaran Sathiamoorthy, Alexandros G
Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, Ashima Suvarna, Benjamin Feuer, Liangyu Chen, Zaid Khan, Eric Frankel, Sachin Grover, Caroline Choi, Niklas Muennighoff, Shiye Su, Wanj...
2025 arXiv
-
[59]
Hasan Abed Al Kader Hammoud, Hani Itani, and Bernard Ghanem. 2025. https://arxiv.org/abs/2504.20708 Beyond the last answer: Your reasoning trace uncovers more than you think . Preprint, arXiv:2504.20708
2025 arXiv
-
[60]
Tingxu Han, Zhenting Wang, Chunrong Fang, Shiyu Zhao, Shiqing Ma, and Zhenyu Chen. 2025. https://arxiv.org/abs/2412.18547 Token-budget-aware llm reasoning . Preprint, arXiv:2412.18547
2025 arXiv
-
[61]
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024. https://arxiv.org/abs/2412.06769 Training large language models to reason in a continuous latent space . Preprint, arXiv:2412.06769
2024 arXiv
-
[62]
Masoud Hashemi, Oluwanifemi Bamgbose, Sathwik Tejaswi Madhusudhan, Jishnu Sethumadhavan Nair, Aman Tiwari, and Vikas Yadav. 2025. https://arxiv.org/abs/2503.15793 Dnr bench: Benchmarking over-reasoning in reasoning llms . Preprint, arXiv:2503.15793
2025 arXiv
-
[63]
Michael Hassid, Gabriel Synnaeve, Yossi Adi, and Roy Schwartz. 2025. https://arxiv.org/abs/2505.17813 Don't overthink it. preferring shorter thinking chains for improved llm reasoning . Preprint, arXiv:2505.17813
2025
-
[64]
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. 2024 a . https://arxiv.org/abs/2402.14008 Olympiadbench: A challenging benchmark for promoting agi with o...
2024 arXiv
-
[65]
Jujie He, Jiacai Liu, Chris Yuhao Liu, Rui Yan, Chaojie Wang, Peng Cheng, Xiaoyu Zhang, Fuxiang Zhang, Jiacheng Xu, Wei Shen, Siyuan Li, Liang Zeng, Tianwen Wei, Cheng Cheng, Bo An, Yang Liu, and Yahui Zhou. 2025 a . https://arxiv.org/abs/2505.22312 Skywork open reasoner 1 tec...
2025 arXiv
-
[66]
Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Xuepeng Liu, Dekai Sun, Shirong Lin, Zhicheng Zheng, Xiaoyong Zhu, Wenbo Su, and Bo Zheng. 2024 b . https://arxiv.org/abs/2411.07140 Chin...
2024 arXiv
-
[67]
Yang He, Xiao Ding, Bibo Cai, Yufei Zhang, Kai Xiong, Zhouhao Sun, Bing Qin, and Ting Liu. 2025 b . https://arxiv.org/abs/2505.20664 Self-route: Automatic mode switching via capability estimation for efficient reasoning . Preprint, arXiv:2505.20664
2025 arXiv
-
[68]
Zhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu, Xingyu Chen, Yue Wang, Linfeng Song, Dian Yu, Zhenwen Liang, Wenxuan Wang, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu. 2025 c . https://arxiv.org/abs/2504.11456 Deepmath-103k: A large-scale, challenging, deconta...
2025 arXiv
-
[69]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 a . https://arxiv.org/abs/2009.03300 Measuring massive multitask language understanding . Preprint, arXiv:2009.03300
2021 arXiv
-
[70]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 b . https://arxiv.org/abs/2103.03874 Measuring mathematical problem solving with the math dataset . Preprint, arXiv:2103.03874
2021 arXiv
-
[71]
Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, and Shiyu Chang. 2025. https://arxiv.org/abs/2504.01296 Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning . Preprint, arXiv:2504.01296
2025 arXiv
-
[72]
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024. https://arxiv.org/abs/2403.12031 Routerbench: A benchmark for multi-llm routing system . Preprint, arXiv:2403.12031
2024 arXiv
-
[73]
Chengyu Huang, Zhengxin Zhang, and Claire Cardie. 2025 a . https://arxiv.org/abs/2505.11225 Hapo: Training language models to reason concisely via history-aware policy optimization . Preprint, arXiv:2505.11225
2025
-
[74]
Shijue Huang, Hongru Wang, Wanjun Zhong, Zhaochen Su, Jiazhan Feng, Bowen Cao, and Yi R. Fung. 2025 b . https://arxiv.org/abs/2505.18822 Adactrl: Towards adaptive and controllable reasoning via difficulty-aware budgeting . Preprint, arXiv:2505.18822
2025
-
[75]
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Yichang Xu, and Ling Liu. 2025 c . https://arxiv.org/abs/2503.00555 Safety tax: Safety alignment makes your large reasoning models less reasonable . Preprint, arXiv:2503.00555
2025 arXiv
-
[76]
Yao Huang, Huanran Chen, Shouwei Ruan, Yichi Zhang, Xingxing Wei, and Yinpeng Dong. 2025 d . https://arxiv.org/abs/2505.22411 Mitigating overthinking in large reasoning models via manifold steering . Preprint, arXiv:2505.22411
2025
-
[77]
iFLYTEK Xinghuo Team . 2025. https://xinghuo.xfyun.cn/sparkapi
2025
-
[78]
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023. https://arxiv.org/abs/2212.04089 Editing models with task arithmetic . Preprint, arXiv:2212.04089
2023 arXiv
-
[79]
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. https://arxiv.org/abs/2403.07974 Livecodebench: Holistic and contamination free evaluation of large language models for code . Preprint, a...
2024 arXiv
-
[80]
Yunjie Ji, Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yiping Peng, Han Zhao, and Xiangang Li. 2025. https://arxiv.org/abs/2505.08311 Am-thinking-v1: Advancing the frontier of reasoning at 32b scale . Preprint, arXiv:2505.08311
2025 arXiv
-
[81]
Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. 2025 a . https://arxiv.org/abs/2502.12025 Safechain: Safety of language models with long chain-of-thought reasoning capabilities . Preprint, arXiv:2502.12025
2025 arXiv
-
[82]
Gangwei Jiang, Yahui Liu, Zhaoyi Li, Qi Wang, Fuzheng Zhang, Linqi Song, Ying Wei, and Defu Lian. 2025 b . https://arxiv.org/abs/2505.22148 What makes a good reasoning chain? uncovering structural patterns in long chain-of-thought reasoning . Preprint, arXiv:2505.22148
2025 arXiv
-
[83]
Guochao Jiang, Guofeng Quan, Zepeng Ding, Ziqin Luo, Dixuan Wang, and Zheng Hu. 2025 c . https://arxiv.org/abs/2505.13949 Flashthink: An early exit method for efficient reasoning . Preprint, arXiv:2505.13949
2025 arXiv
-
[84]
Lingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong, Xingxing Zhang, Tengchao Lv, Lei Cui, and Furu Wei. 2025 d . https://arxiv.org/abs/2505.14631 Think only when you need with large hybrid-reasoning models . Preprint, arXiv:2505.14631
2025 arXiv
-
[85]
Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. 2024. https://arxiv.org/abs/2406.18510 Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer la...
2024 arXiv
-
[86]
Yuxuan Jiang, Dawei Li, and Frank Ferraro. 2025 e . https://arxiv.org/abs/2505.13975 Drp: Distilled reasoning pruning with skill-aware step decomposition for efficient large reasoning models . Preprint, arXiv:2505.13975
2025 arXiv
-
[87]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. https://arxiv.org/abs/2310.06770 Swe-bench: Can language models resolve real-world github issues? Preprint, arXiv:2310.06770
2024 arXiv
-
[88]
Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du. 2024. https://arxiv.org/abs/2401.04925 The impact of reasoning step length on large language models . Preprint, arXiv:2401.04925
2024 arXiv
-
[89]
Zhensheng Jin, Xinze Li, Yifan Ji, Chunyi Peng, Zhenghao Liu, Qi Shi, Yukun Yan, Shuo Wang, Furong Peng, and Ge Yu. 2025. https://arxiv.org/abs/2506.10822 Recut: Balancing reasoning length and accuracy in llms via stepwise trails and preference optimization . Preprint, arXiv:2...
2025 arXiv
-
[90]
Yu Kang, Xianghui Sun, Liangyu Chen, and Wei Zou. 2024. https://arxiv.org/abs/2412.11664 C3ot: Generating shorter chain-of-thought without compromising effectiveness . Preprint, arXiv:2412.11664
2024 arXiv
-
[91]
KwaiPilot Team . 2025. https://huggingface.co/kwaipilot/kwaicoder-autothink-preview
2025
-
[92]
Ayeong Lee, Ethan Che, and Tianyi Peng. 2025. https://arxiv.org/abs/2503.01141 How well do llms compress their own chain-of-thought? a token complexity approach . Preprint, arXiv:2503.01141
2025 arXiv
-
[93]
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. https://arxiv.org/abs/2206.14858 Solving quantitative re...
2022 arXiv
-
[94]
Junyan Li, Wenshuo Zhao, Yang Zhang, and Chuang Gan. 2025 a . https://arxiv.org/abs/2506.13752 Steering llm thinking with budget guidance . Preprint, arXiv:2506.13752
2025 arXiv
-
[95]
Gonzalez, and Ion Stoica
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024. https://arxiv.org/abs/2406.11939 From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline . Preprint, arXiv:2406.11939
2024 arXiv
-
[96]
Kwok, and Yu Zhang
Wei Li, Yanbin Wei, Qiushi Huang, Jiangyue Yan, Yang Chen, James T. Kwok, and Yu Zhang. 2025 b . https://arxiv.org/abs/2506.05936 Dynamicmind: A tri-mode thinking system for large language models . Preprint, arXiv:2506.05936
2025 arXiv
-
[97]
Xuying Li, Zhuo Li, Yuji Kosuga, and Victor Bian. 2025 c . https://arxiv.org/abs/2503.01923 Output length effect on deepseek-r1's safety in forced thinking . Preprint, arXiv:2503.01923
2025 arXiv
-
[98]
Zheng Li, Qingxiu Dong, Jingyuan Ma, Di Zhang, and Zhifang Sui. 2025 d . https://arxiv.org/abs/2505.11274 Selfbudgeter: Adaptive token allocation for efficient llm reasoning . Preprint, arXiv:2505.11274
2025 arXiv
-
[99]
Zhiyuan Li, Yi Chang, and Yuan Wu. 2025 e . https://arxiv.org/abs/2505.22113 Think-bench: Evaluating thinking efficiency and chain-of-thought quality of large reasoning models . Preprint, arXiv:2505.22113
2025 arXiv
-
[100]
Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Ying Nian Wu, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, and Cheng-Lin Liu. 2025 f . https://arxiv.org/abs/2506.02678 Tl;dr: Too long, do re-weighting for efficient llm...
2025 arXiv
-
[101]
Zhong-Zhi Li, Duzhen Zhang, Ming-Liang Zhang, Jiaxin Zhang, Zengyan Liu, Yuxuan Yao, Haotian Xu, Junhao Zheng, Pei-Jie Wang, Xiuyi Chen, Yingying Zhang, Fei Yin, Jiahua Dong, Zhiwei Li, Bao-Long Bi, Ling-Rui Mei, Junfeng Fang, Zhijiang Guo, Le Song, and Cheng-Lin Liu. 2025 g ....
2025 arXiv
-
[102]
Guosheng Liang, Longguang Zhong, Ziyi Yang, and Xiaojun Quan. 2025. https://arxiv.org/abs/2505.14183 Thinkswitcher: When to think hard, when to think fast . Preprint, arXiv:2505.14183
2025 arXiv
-
[103]
Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan. 2024. https://arxiv.org/abs/2401.08190 Mario: Math reasoning with code interpreter output -- a reproducible pipeline . Preprint, arXiv:2401.08190
2024 arXiv
-
[104]
Junhong Lin, Xinyue Zeng, Jie Zhu, Song Wang, Julian Shun, Jun Wu, and Dawei Zhou. 2025. https://arxiv.org/abs/2505.16122 Plan and budget: Effective and efficient test-time scaling on large language model reasoning . Preprint, arXiv:2505.16122
2025
-
[105]
Zehui Ling, Deshu Chen, Hongwei Zhang, Yifeng Jiao, Xin Guo, and Yuan Cheng. 2025. https://arxiv.org/abs/2506.10446 Fast on the easy, deep on the hard: Efficient reasoning via powered length penalty . Preprint, arXiv:2506.10446
2025 arXiv
-
[106]
Fengyuan Liu, Nouar AlDahoul, Gregory Eady, Yasir Zaki, and Talal Rahwan. 2025 a . https://arxiv.org/abs/2406.10400 Self-reflection makes large language models safer, less biased, and ideologically neutral . Preprint, arXiv:2406.10400
2025 arXiv
-
[107]
Hongwei Liu, Zilong Zheng, Yuxuan Qiao, Haodong Duan, Zhiwei Fei, Fengzhe Zhou, Wenwei Zhang, Songyang Zhang, Dahua Lin, and Kai Chen. 2024 a . https://arxiv.org/abs/2405.12209 Mathbench: Evaluating the theory and application proficiency of llms with a hierarchical mathematics...
2024 arXiv
-
[108]
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020. https://arxiv.org/abs/2007.08124 Logiqa: A challenge dataset for machine reading comprehension with logical reasoning . Preprint, arXiv:2007.08124
2020 arXiv
-
[109]
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. https://arxiv.org/abs/2305.01210 Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation . Preprint, arXiv:2305.01210
2023 arXiv
-
[110]
Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. 2024 b . https://arxiv.org/abs/2408.06450 Evaluating language models for efficient code generation . Preprint, arXiv:2408.06450
2024 arXiv
-
[111]
Kaiyuan Liu, Chen Shen, Zhanwei Zhang, Junjie Liu, Xiaosong Yuan, and Jieping ye. 2025 b . https://arxiv.org/abs/2506.12353 Efficient reasoning through suppression of self-affirmation reflections in large reasoning models . Preprint, arXiv:2506.12353
2025 arXiv
-
[112]
Peijie Liu, Fengli Xu, and Yong Li. 2025 c . https://arxiv.org/abs/2506.06008 Token signature: Predicting chain-of-thought gains with token decoding feature in large language models . Preprint, arXiv:2506.06008
2025 arXiv
-
[113]
Ruikang Liu, Yuxuan Sun, Manyi Zhang, Haoli Bai, Xianzhi Yu, Tiezheng Yu, Chun Yuan, and Lu Hou. 2025 d . https://arxiv.org/abs/2504.04823 Quantization hurts reasoning? an empirical study on quantized reasoning models . Preprint, arXiv:2504.04823
2025 arXiv
-
[114]
Shuqi Liu, Han Wu, Bowei He, Xiongwei Han, Mingxuan Yuan, and Linqi Song. 2025 e . https://arxiv.org/abs/2502.12420 Sens-merging: Sensitivity-guided parameter balancing for merging large language models . Preprint, arXiv:2502.12420
2025 arXiv
-
[115]
Tianqiao Liu, Zui Chen, Zitao Liu, Mi Tian, and Weiqi Luo. 2024 c . https://arxiv.org/abs/2409.08561 Expediting and elevating large language model reasoning via hidden chain-of-thought decoding . Preprint, arXiv:2409.08561
2024 arXiv
-
[116]
Wanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin, Ke Ji, Wenyu Chen, Yan Xu, Yasheng Wang, Lifeng Shang, and Benyou Wang. 2025 f . https://arxiv.org/abs/2506.12860 Qfft, question-free fine-tuning for adaptive reasoning . Preprint, arXiv:2506.12860
2025 arXiv
-
[117]
Xin Liu and Lu Wang. 2025. https://arxiv.org/abs/2506.02536 Answer convergence as a signal for early stopping in reasoning . Preprint, arXiv:2506.02536
2025
-
[118]
Yongjiang Liu, Haoxi Li, Xiaosong Ma, Jie Zhang, and Song Guo. 2025 g . https://arxiv.org/abs/2507.02663 Think how to think: Mitigating overthinking with autonomous difficulty cognition in large reasoning models . Preprint, arXiv:2507.02663
2025
-
[119]
Yue Liu, Jiaying Wu, Yufei He, Hongcheng Gao, Hongyu Chen, Baolong Bi, Jiaheng Zhang, Zhiqi Huang, and Bryan Hooi. 2025 h . https://arxiv.org/abs/2503.23077 Efficient inference for large reasoning models: A survey . Preprint, arXiv:2503.23077
2025 arXiv
-
[120]
Yule Liu, Jingyi Zheng, Zhen Sun, Zifan Peng, Wenhan Dong, Zeyang Sha, Shiwen Cui, Weiqiang Wang, and Xinlei He. 2025 i . https://arxiv.org/abs/2504.13626 Thought manipulation: External thought can be efficient for large reasoning models . Preprint, arXiv:2504.13626
2025 arXiv
-
[121]
Chenwei Lou, Zewei Sun, Xinnian Liang, Meng Qu, Wei Shen, Wenqi Wang, Yuntao Li, Qingping Yang, and Shuangzhi Wu. 2025. https://arxiv.org/abs/2505.11896 Adacot: Pareto-optimal adaptive chain-of-thought triggering via reinforcement learning . Preprint, arXiv:2505.11896
2025 arXiv
-
[122]
Jinghui Lu, Haiyang Yu, Siliang Xu, Shiwei Ran, Guozhi Tang, Siqi Wang, Bin Shan, Teng Fu, Hao Feng, Jingqun Tang, Han Wang, and Can Huang. 2025. https://arxiv.org/abs/2505.15154 Prolonged reasoning is not all you need: Certainty-based adaptive routing for efficient llm/mllm r...
2025 arXiv
-
[123]
Feng Luo, Yu-Neng Chuang, Guanchu Wang, Hoang Anh Duy Le, Shaochen Zhong, Hongyi Liu, Jiayi Yuan, Yang Sui, Vladimir Braverman, Vipin Chaudhary, and Xia Hu. 2025 a . https://arxiv.org/abs/2505.22662 Autol2s: Auto long-short reasoning for efficient large language models . Prepr...
2025
-
[124]
Haotian Luo, Haiying He, Yibo Wang, Jinluan Yang, Rui Liu, Naiqiang Tan, Xiaochun Cao, Dacheng Tao, and Li Shen. 2025 b . https://arxiv.org/abs/2504.21659 Ada-r1: Hybrid-cot via bi-level adaptive reasoning optimization . Preprint, arXiv:2504.21659
2025 arXiv
-
[125]
Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao. 2025 c . https://arxiv.org/abs/2501.12570 O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning . Preprint, arXiv:2501.12570
2025 arXiv
-
[126]
Kaijing Ma, Xinrun Du, Yunran Wang, Haoran Zhang, Zhoufutu Wen, Xingwei Qu, Jian Yang, Jiaheng Liu, Minghao Liu, Xiang Yue, Wenhao Huang, and Ge Zhang. 2025 a . https://arxiv.org/abs/2410.06526 Kor-bench: Benchmarking language models on knowledge-orthogonal reasoning tasks . P...
2025 arXiv
-
[127]
Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, and Matei Zaharia. 2025 b . https://arxiv.org/abs/2504.09858 Reasoning models can be effective without thinking . Preprint, arXiv:2504.09858
2025 arXiv
-
[128]
Xinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang, and Xinchao Wang. 2025 c . https://arxiv.org/abs/2502.09601 Cot-valve: Length-compressible chain-of-thought tuning . Preprint, arXiv:2502.09601
2025 arXiv
-
[129]
Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2023. https://arxiv.org/abs/2212.08410 Teaching small language models to reason . Preprint, arXiv:2212.08410
2023 arXiv
-
[130]
Sadegh Mahdavi, Muchen Li, Kaiwen Liu, Christos Thrampoulidis, Leonid Sigal, and Renjie Liao. 2025. https://arxiv.org/abs/2501.14275 Leveraging online olympiad-level math problems for llms training and contamination-resistant evaluation . Preprint, arXiv:2501.14275
2025 arXiv
-
[131]
Zhiting Mei, Christina Zhang, Tenny Yin, Justin Lidard, Ola Shorinwa, and Anirudha Majumdar. 2025. https://arxiv.org/abs/2506.18183 Reasoning about uncertainty: Do reasoning models know when they don't know? Preprint, arXiv:2506.18183
2025 arXiv
-
[132]
Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2021. https://arxiv.org/abs/2106.15772 A diverse corpus for evaluating and developing english math word problem solvers . Preprint, arXiv:2106.15772
2021 arXiv
-
[133]
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. https://arxiv.org/abs/1809.02789 Can a suit of armor conduct electricity? a new dataset for open book question answering . Preprint, arXiv:1809.02789
2018 arXiv
-
[134]
MiniMax, :, Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, et al. 2025. https://arxiv.org/abs/2506.13585 Minimax-m1: Scaling test-time compute efficiently with lightning attention . Preprint, arXiv:2506.13585
2025 arXiv
-
[135]
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. https://arxiv.org/abs/2104.08773 Cross-task generalization via natural language crowdsourcing instructions . Preprint, arXiv:2104.08773
2022 arXiv
-
[136]
Ivan Moshkov, Darragh Hanley, Ivan Sorokin, Shubham Toshniwal, Christof Henkel, Benedikt Schifferer, Wei Du, and Igor Gitman. 2025. Aimo-2 winning solution: Building state-of-the-art mathematical reasoning models with openmathreasoning dataset. arXiv preprint arXiv:2504.16891
2025 arXiv
-
[137]
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025. https://arxiv.org/abs/2501.19393 s1: Simple test-time scaling . Preprint, arXiv:2501.19393
2025 arXiv
-
[138]
Tergel Munkhbat, Namgyu Ho, Seo Hyun Kim, Yongjin Yang, Yujin Kim, and Se-Young Yun. 2025. https://arxiv.org/abs/2502.20122 Self-training elicits concise reasoning in large language models . Preprint, arXiv:2502.20122
2025 arXiv
-
[139]
Sania Nayab, Giulio Rossolini, Marco Simoni, Andrea Saracino, Giorgio Buttazzo, Nicolamaria Manes, and Fabrizio Giacomelli. 2025. https://arxiv.org/abs/2407.19825 Concise thoughts: Impact of output length on llm reasoning and cost . Preprint, arXiv:2407.19825
2025 arXiv
-
[140]
Yansong Ning, Wei Li, Jun Fang, Naiqiang Tan, and Hao Liu. 2025. https://arxiv.org/abs/2505.11827 Not all thoughts are generated equal: Efficient llm reasoning via multi-turn reinforcement learning . Preprint, arXiv:2505.11827
2025 arXiv
-
[141]
Amin Heyrani Nobari, Kaveh Alimohammadi, Ali ArjomandBigdeli, Akash Srivastava, Faez Ahmed, and Navid Azizan. 2025. https://arxiv.org/abs/2502.02421 Activation-informed merging of large language models . Preprint, arXiv:2502.02421
2025
-
[142]
Gonzalez, M Waleed Kadous, and Ion Stoica
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025. https://arxiv.org/abs/2406.18665 Routellm: Learning to route llms with preference data . Preprint, arXiv:2406.18665
2025 arXiv
-
[143]
OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, et al. 2024. https://arxiv.org/abs/2412.16720 Openai o1 system card . Preprint, arXiv:2412.16720
2024 arXiv
-
[144]
Li, Aviv Bick, J
Daniele Paliotta, Junxiong Wang, Matteo Pagliardini, Kevin Y. Li, Aviv Bick, J. Zico Kolter, Albert Gu, François Fleuret, and Tri Dao. 2025. https://arxiv.org/abs/2502.20339 Thinking slow, fast: Scaling inference compute with distilled reasoners . Preprint, arXiv:2502.20339
2025 arXiv
-
[145]
Jiabao Pan, Yan Zhang, Chen Zhang, Zuozhu Liu, Hongwei Wang, and Haizhou Li. 2024. https://arxiv.org/abs/2407.01009 Dynathink: Fast or slow? a dynamic decision-making framework for large language models . Preprint, arXiv:2407.01009
2024 arXiv
-
[146]
Qianjun Pan, Wenkai Ji, Yuyang Ding, Junsong Li, Shilian Chen, Junyi Wang, Jie Zhou, Qin Chen, Min Zhang, Yulan Wu, and Liang He. 2025 a . https://arxiv.org/abs/2505.02665 A survey of slow thinking-based reasoning llms using reinforced learning and inference-time scaling law ....
2025 arXiv
-
[147]
Zhihong Pan, Kai Zhang, Yuze Zhao, and Yupeng Han. 2025 b . https://arxiv.org/abs/2505.19435 Route to reason: Adaptive routing for llm and reasoning strategy selection . Preprint, arXiv:2505.19435
2025 arXiv
-
[148]
Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2024. https://arxiv.org/abs/2312.06681 Steering llama 2 via contrastive activation addition . Preprint, arXiv:2312.06681
2024 arXiv
-
[149]
Shubham Parashar, Blake Olson, Sambhav Khurana, Eric Li, Hongyi Ling, James Caverlee, and Shuiwang Ji. 2025. https://arxiv.org/abs/2502.12521 Inference-time computations for llm reasoning and planning: A benchmark and insights . Preprint, arXiv:2502.12521
2025 arXiv
-
[150]
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. https://arxiv.org/abs/2103.07191 Are nlp models really able to solve simple math word problems? Preprint, arXiv:2103.07191
2021 arXiv
-
[151]
Jacob Pfau, William Merrill, and Samuel R. Bowman. 2024. https://arxiv.org/abs/2404.15758 Let's think dot by dot: Hidden computation in transformer language models . Preprint, arXiv:2404.15758
2024 arXiv
-
[152]
Penghui Qi, Zichen Liu, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. 2025. https://arxiv.org/abs/2505.13438 Optimizing anytime reasoning via budget relative policy optimization . Preprint, arXiv:2505.13438
2025
-
[153]
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Fandong Meng, Jie Zhou, Ju Ren, and Yaoxue Zhang. 2025. https://arxiv.org/abs/2505.04881 Concise: Confidence-guided compression in step-by-step efficient reasoning . Preprint, arXiv:2505.04881
2025
-
[154]
Yanzhao Qin, Tao Zhang, Tao Zhang, Yanjun Shen, Wenjing Luo, Haoze Sun, Yan Zhang, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang, and Bin Cui. 2024. https://arxiv.org/abs/2408.10943 Sysbench: Can large language models follow system messages? Preprint, arXiv:2408.10943
2024
-
[155]
Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, Peng Li, Wei Wei, Jing Shao, Chaochao Lu, Yue Zhang, Xian-Sheng Hua, Bowen Zhou, and Yu Cheng. 2025. https://arxiv.org/abs/2503.21614 A survey of efficient...
2025
-
[156]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, et al. 2025. https://arxiv.org/abs/2412.15115 Qwen2.5 technical report . Preprint, arXiv:2412.15115
2025 arXiv
-
[157]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2023. https://arxiv.org/abs/2311.12022 Gpqa: A graduate-level google-proof q&a benchmark . Preprint, arXiv:2311.12022
2023 arXiv
-
[158]
Matthew Renze and Erhan Guven. 2024. https://doi.org/10.1109/fllm63129.2024.10852493 The benefits of a concise chain of thought on problem-solving in large language models . In 2024 2nd International Conference on Foundation and Large Language Models (FLLM), page 476–483. IEEE
2024
-
[159]
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J. Reddi. 2025. https://arxiv.org/abs/2502.17416 Reasoning with latent thoughts: On the power of looped transformers . Preprint, arXiv:2502.17416
2025 arXiv
-
[160]
ByteDance Seed, :, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, et al. 2025. https://arxiv.org/abs/2504.13914 Seed1.5-thinking: Advancing superb reasoning models with reinforcement learning . Preprint, arXiv:2504.13914
2025
-
[161]
Xuan Shen, Yizhou Wang, Xiangxi Shi, Yanzhi Wang, Pu Zhao, and Jiuxiang Gu. 2025 a . https://arxiv.org/abs/2501.19201 Efficient reasoning with hidden thinking . Preprint, arXiv:2501.19201
2025 arXiv
-
[162]
Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi, Wenjing Zhang, Jiangze Yan, Ning Wang, Kai Wang, and Shiguo Lian. 2025 b . https://arxiv.org/abs/2503.04472 Dast: Difficulty-adaptive slow-thinking for large reasoning models . Preprint, arXiv:2503.04472
2025
-
[163]
Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He. 2025 c . https://arxiv.org/abs/2502.21074 Codi: Compressing chain-of-thought into continuous space via self-distillation . Preprint, arXiv:2502.21074
2025 arXiv
-
[164]
Leheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao, Changshuo Shen, Yi Zhang, Xiang Wang, and Tat-Seng Chua. 2025. https://arxiv.org/abs/2506.08390 On reasoning strength planning in large reasoning models . Preprint, arXiv:2506.08390
2025
-
[165]
Mingyang Song and Mao Zheng. 2025. https://arxiv.org/abs/2505.21178 Walk before you run! concise llm reasoning via reinforcement learning . Preprint, arXiv:2505.21178
2025 arXiv
-
[166]
DiJia Su, Sainbayar Sukhbaatar, Michael Rabbat, Yuandong Tian, and Qinqing Zheng. 2025 a . https://arxiv.org/abs/2410.09918 Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces . Preprint, arXiv:2410.09918
2025 arXiv
-
[167]
DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, and Qinqing Zheng. 2025 b . https://arxiv.org/abs/2502.03275 Token assorted: Mixing latent and text tokens for improved language model reasoning . Preprint, arXiv:2502.03275
2025 arXiv
-
[168]
Jinyan Su and Claire Cardie. 2025. https://arxiv.org/abs/2505.18298 Thinking fast and right: Balancing accuracy and reasoning length with adaptive rewards . Preprint, arXiv:2505.18298
2025 arXiv
-
[169]
Jinyan Su, Jennifer Healey, Preslav Nakov, and Claire Cardie. 2025 c . https://arxiv.org/abs/2505.00127 Between underthinking and overthinking: An empirical study of reasoning length and correctness in llms . Preprint, arXiv:2505.00127
2025 arXiv
-
[170]
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu. 2025. https://arxiv.org/abs/2503.16419 Stop overthinking: A survey on efficient reasoning for large language models . Preprint, arXiv...
2025 arXiv
-
[171]
Yi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu, Xiangyu Li, Hao Wen, Yizhen Yuan, Huiwen Zheng, Yan Liang, Yuanchun Li, and Yunxin Liu. 2025 a . https://arxiv.org/abs/2504.14350 An empirical study of llm reasoning ability under strict output length constraint . Preprint, arXiv:2504.14350
2025 arXiv
-
[172]
Zhongxiang Sun, Qipeng Wang, Haoyu Wang, Xiao Zhang, and Jun Xu. 2025 b . https://arxiv.org/abs/2505.12886 Detection and mitigation of hallucination in large reasoning models: A mechanistic perspective . Preprint, arXiv:2505.12886
2025 arXiv
-
[173]
Le, Ed H
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2022. https://arxiv.org/abs/2210.09261 Challenging big-bench tasks and whether chain-of-thought can solve them . ...
2022 arXiv
-
[174]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://arxiv.org/abs/1811.00937 Commonsenseqa: A question answering challenge targeting commonsense knowledge . Preprint, arXiv:1811.00937
2019 arXiv
-
[175]
Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, and Ruihua Song. 2025. https://arxiv.org/abs/2505.16552 Think silently, think fast: Dynamic latent compression of llm reasoning chains . Preprint, arXiv:2505.16552
2025
-
[176]
Siao Tang, Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2025. https://arxiv.org/abs/2506.18810 Concisehint: Boosting efficient reasoning via continuous concise hints during generation . Preprint, arXiv:2506.18810
2025
-
[177]
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei. 2024. https://arxiv.org/abs/2403.02884 Mathscale: Scaling instruction tuning for mathematical reasoning . Preprint, arXiv:2403.02884
2024 arXiv
-
[178]
Core Team, Bingquan Xia, Bowen Shen, Cici, Dawei Zhu, et al. 2025 a . https://arxiv.org/abs/2505.07608 Mimo: Unlocking the reasoning potential of language model -- from pretraining to posttraining . Preprint, arXiv:2505.07608
2025 arXiv
-
[179]
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, et al. 2025 b . https://arxiv.org/abs/2501.12599 Kimi k1.5: Scaling reinforcement learning with llms . Preprint, arXiv:2501.12599
2025 arXiv
-
[180]
P Team, Xinrun Du, Yifan Yao, Kaijing Ma, Bingli Wang, et al. 2025 c . https://arxiv.org/abs/2502.14739 Supergpqa: Scaling llm evaluation across 285 graduate disciplines . Preprint, arXiv:2502.14739
2025 arXiv
-
[181]
Qwen Team. 2025. https://qwenlm.github.io/blog/qwq-32b/ Qwq-32b: Embracing the power of reinforcement learning
2025
-
[182]
Tencent Hunyuan Team, Ao Liu, Botong Zhou, Can Xu, Chayse Zhou, et al. 2025 d . https://arxiv.org/abs/2505.15431 Hunyuan-turbos: Advancing large language models through mamba-transformer synergy and adaptive chain-of-thought . Preprint, arXiv:2505.15431
2025
-
[183]
Tencent Hunyuan . 2025. https://github.com/tencent-hunyuan/hunyuan-a13b
2025
-
[184]
Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yiping Peng, Yunjie Ji, Han Zhao, and Xiangang Li. 2025. https://arxiv.org/abs/2504.17565 Deepdistill: Enhancing llm reasoning capabilities via large-scale difficulty-graded data training . Preprint, arXiv:2504.17565
2025 arXiv
-
[185]
transluce . 2025. https://transluce.org/investigating-o3-truthfulness
2025
-
[186]
Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, Linjing Li, Xiangyuan Lan, and Dongbin Zhao. 2025. https://arxiv.org/abs/2505.10832 Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl . Preprint, arXiv:2505.10832
2025
-
[187]
vectara . 2025. https://www.vectara.com/blog/deepseek-r1-hallucinates-more-than-deepseek-v3
2025
-
[188]
Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna, and Tianyi Zhou. 2025 a . https://arxiv.org/abs/2506.08343 Wait, we don't need to "wait"! removing thinking tokens improves reasoning efficiency . Preprint, arXiv:2506.08343
2025 arXiv
-
[189]
Jikai Wang, Juntao Li, Jianye Hou, Bowen Yan, Lijun Wu, and Min Zhang. 2025 b . https://arxiv.org/abs/2504.19095 Efficient reasoning for llms through speculative chain-of-thought . Preprint, arXiv:2504.19095
2025 arXiv
-
[190]
Jingyao Wang, Wenwen Qiang, Zeen Song, Changwen Zheng, and Hui Xiong. 2025 c . https://arxiv.org/abs/2505.10425 Learning to think: Information-theoretic reinforcement fine-tuning for llms . Preprint, arXiv:2505.10425
2025
-
[191]
Rush, and Tri Dao
Junxiong Wang, Wen-Ding Li, Daniele Paliotta, Daniel Ritter, Alexander M. Rush, and Tri Dao. 2025 d . https://arxiv.org/abs/2504.10449 M1: Towards scalable test-time compute with mamba reasoning models . Preprint, arXiv:2504.10449
2025 arXiv
-
[192]
Minzheng Wang, Yongbin Li, Haobo Wang, Xinghua Zhang, Nan Xu, Bingli Wu, Fei Huang, Haiyang Yu, and Wenji Mao. 2025 e . https://arxiv.org/abs/2505.02156 Adaptive thinking via mode policy optimization for social language agents . Preprint, arXiv:2505.02156
2025
-
[193]
Xiaoqiang Wang, Suyuchen Wang, Yun Zhu, and Bang Liu. 2025 f . https://arxiv.org/abs/2505.18962 System-1.5 reasoning: Traversal in language and latent spaces with dynamic shortcuts . Preprint, arXiv:2505.18962
2025 arXiv
-
[194]
Xinyi Wang, Lucas Caccia, Oleksiy Ostapenko, Xingdi Yuan, William Yang Wang, and Alessandro Sordoni. 2024 a . https://arxiv.org/abs/2310.05707 Guiding language model reasoning with planning tokens . Preprint, arXiv:2310.05707
2024 arXiv
-
[195]
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024 b . https://arxiv.org/abs/2406.01574 Mmlu-pro: A more robust ...
2024 arXiv
-
[196]
Yuhang Wang, Youhe Jiang, Bin Cui, and Fangcheng Fu. 2025 g . https://arxiv.org/abs/2505.13326 Thinking short and right over thinking long: Serving llm reasoning efficiently and accurately . Preprint, arXiv:2505.13326
2025 arXiv
-
[197]
Yunhao Wang, Yuhao Zhang, Tinghao Yu, Can Xu, Feng Zhang, and Fengzong Lian. 2025 h . https://arxiv.org/abs/2505.20101 Adaptive deep reasoning: Triggering deep thinking when needed . Preprint, arXiv:2505.20101
2025 arXiv
-
[198]
Bartoldson, Bhavya Kailkhura, and Cihang Xie
Zijun Wang, Haoqin Tu, Yuhan Wang, Juncheng Wu, Jieru Mei, Brian R. Bartoldson, Bhavya Kailkhura, and Cihang Xie. 2025 i . https://arxiv.org/abs/2504.01903 Star-1: Safer alignment of reasoning llms with 1k data . Preprint, arXiv:2504.01903
2025
-
[199]
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024. https://arxiv.org/abs/2411.04368 Measuring short-form factuality in large language models . Preprint, arXiv:2411.04368
2024 arXiv
-
[200]
Liang Wen, Yunke Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Junchen Liu, Lifu Tang, Xiaowei Lv, Haosheng Zou, Yongchao Deng, Shousheng Jia, and Xiangzheng Zhang. 2025 a . https://arxiv.org/abs/2503.10460 Light-r1: Curriculum sft, dpo and rl for long cot from scrat...
2025 arXiv
-
[201]
Pengcheng Wen, Jiaming Ji, Chi-Min Chan, Juntao Dai, Donghai Hong, Yaodong Yang, Sirui Han, and Yike Guo. 2025 b . https://arxiv.org/abs/2503.12918 Thinkpatterns-21k: A systematic study on the impact of thinking patterns in llms . Preprint, arXiv:2503.12918
2025
-
[202]
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblu...
2025 arXiv
-
[203]
Han Wu, Yuxuan Yao, Shuqi Liu, Zehua Liu, Xiaojin Fu, Xiongwei Han, Xing Li, Hui-Ling Zhen, Tao Zhong, and Mingxuan Yuan. 2025 a . https://arxiv.org/abs/2503.20641 Unlocking efficient long-to-short llm reasoning with model merging . Preprint, arXiv:2503.20641
2025 arXiv
-
[204]
Siye Wu, Jian Xie, Yikai Zhang, Aili Chen, Kai Zhang, Yu Su, and Yanghua Xiao. 2025 b . https://arxiv.org/abs/2505.20258 Arm: Adaptive reasoning model . Preprint, arXiv:2505.20258
2025
-
[205]
Yifan Wu, Jingze Shi, Bingheng Wu, Jiayi Zhang, Xiaotian Lin, Nan Tang, and Yuyu Luo. 2025 c . https://arxiv.org/abs/2505.19716 Concise reasoning, big gains: Pruning long reasoning trace with difficulty-aware prompting . Preprint, arXiv:2505.19716
2025 arXiv
-
[206]
Yuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du, Stefanie Jegelka, and Yisen Wang. 2025 d . https://arxiv.org/abs/2502.07266 When more is less: Understanding chain-of-thought length in llms . Preprint, arXiv:2502.07266
2025 arXiv
-
[207]
Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, and Wenjie Li. 2025. https://arxiv.org/abs/2502.12067 Tokenskip: Controllable chain-of-thought compression in llms . Preprint, arXiv:2502.12067
2025
-
[208]
Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile, Sang Truong, Chelsea Finn, and Nick Haber. 2025. https://arxiv.org/abs/2506.05256 Just enough thinking: Efficient reasoning with adaptive length penalties reinforcement learning . Preprint, arXiv:2506.05256
2025 arXiv
-
[209]
Wenyi Xiao, Leilei Gan, Weilong Dai, Wanggui He, Ziwei Huang, Haoyuan Li, Fangxun Shu, Zhelun Yu, Peng Zhang, Hao Jiang, and Fei Wu. 2025. https://arxiv.org/abs/2504.18458 Fast-slow thinking for large vision-language model reasoning . Preprint, arXiv:2504.18458
2025
-
[210]
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. 2024. https://arxiv.org/abs/2402.01622 Travelplanner: A benchmark for real-world planning with language agents . Preprint, arXiv:2402.01622
2024 arXiv
-
[211]
Roy Xie, David Qiu, Deepak Gopinath, Dong Lin, Yanchao Sun, Chong Wang, Saloni Potdar, and Bhuwan Dhingra. 2025. https://arxiv.org/abs/2505.19640 Interleaved reasoning for large language models via reinforcement learning . Preprint, arXiv:2505.19640
2025
-
[212]
Tao Xiong, Xavier Hu, Wenyan Fan, and Shengyu Zhang. 2025. https://arxiv.org/abs/2507.00606 Mixture of reasonings: Teach large language models to reason with adaptive strategies . Preprint, arXiv:2507.00606
2025 arXiv
-
[213]
Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. 2025 a . https://arxiv.org/abs/2502.18600 Chain of draft: Thinking faster by writing less . Preprint, arXiv:2502.18600
2025 arXiv
-
[214]
Xiaoang Xu, Shuo Wang, Xu Han, Zhenghao Liu, Huijia Wu, Peipei Li, Zhiyuan Liu, Maosong Sun, and Zhaofeng He. 2025 b . https://arxiv.org/abs/2505.24550 A*-thought: Efficient reasoning via bidirectional compression for low-resource settings . Preprint, arXiv:2505.24550
2025
-
[215]
Yige Xu, Xu Guo, Zhiwei Zeng, and Chunyan Miao. 2025 c . https://arxiv.org/abs/2502.12134 Softcot: Soft chain-of-thought for efficient reasoning with llms . Preprint, arXiv:2502.12134
2025 arXiv
-
[216]
Yuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo, Junnan Li, and Caiming Xiong. 2025 d . https://arxiv.org/abs/2505.05315 Scalable chain of thoughts via elastic reasoning . Preprint, arXiv:2505.05315
2025 arXiv
-
[217]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, et al. 2025 a . https://arxiv.org/abs/2505.09388 Qwen3 technical report . Preprint, arXiv:2505.09388
2025 arXiv
-
[218]
Chenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu, Chenyu Zhu, Qiaowei Li, Zheng Lin, Li Cao, and Weiping Wang. 2025 b . https://arxiv.org/abs/2504.15895 Dynamic early exit in reasoning models . Preprint, arXiv:2504.15895
2025
-
[219]
Junjie Yang, Ke Lin, and Xing Yu. 2025 c . https://arxiv.org/abs/2504.03234 Think when you need: Self-adaptive chain-of-thought learning . Preprint, arXiv:2504.03234
2025 arXiv
-
[220]
Shu Yang, Junchao Wu, Xuansheng Wu, Derek Wong, Ninhao Liu, and Di Wang. 2025 d . https://arxiv.org/abs/2506.19492 Is long-to-short a free lunch? investigating inconsistency and reasoning efficiency in lrms . Preprint, arXiv:2506.19492
2025 arXiv
-
[221]
Wenkai Yang, Shuming Ma, Yankai Lin, and Furu Wei. 2025 e . https://arxiv.org/abs/2502.18080 Towards thinking-optimal scaling of test-time compute for llm reasoning . Preprint, arXiv:2502.18080
2025
-
[222]
Cohen, Ruslan Salakhutdinov, and Christopher D
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://arxiv.org/abs/1809.09600 Hotpotqa: A dataset for diverse, explainable multi-hop question answering . Preprint, arXiv:1809.09600
2018 arXiv
-
[223]
Wenlin Yao, Haitao Mi, and Dong Yu. 2024. https://arxiv.org/abs/2409.17433 Hdflow: Enhancing llm complex problem-solving with hybrid thinking and dynamic workflows . Preprint, arXiv:2409.17433
2024 arXiv
-
[224]
Yuxuan Yao, Shuqi Liu, Zehua Liu, Qintong Li, Mingyang Liu, Xiongwei Han, Zhijiang Guo, Han Wu, and Linqi Song. 2025 a . https://arxiv.org/abs/2505.14009 Activation-guided consensus merging for large language models . Preprint, arXiv:2505.14009
2025
-
[225]
Zijun Yao, Yantao Liu, Yanxu Chen, Jianhui Chen, Junfeng Fang, Lei Hou, Juanzi Li, and Tat-Seng Chua. 2025 b . https://arxiv.org/abs/2505.23646 Are reasoning models more prone to hallucination? Preprint, arXiv:2505.23646
2025 arXiv
-
[226]
Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue. 2025. https://arxiv.org/abs/2502.03373 Demystifying long chain-of-thought reasoning in llms . Preprint, arXiv:2502.03373
2025 arXiv
-
[227]
Jingyang Yi, Jiazheng Wang, and Sida Li. 2025. https://arxiv.org/abs/2504.21370 Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning . Preprint, arXiv:2504.21370
2025
-
[228]
Xixian Yong, Xiao Zhou, Yingying Zhang, Jinlin Li, Yefeng Zheng, and Xian Wu. 2025. https://arxiv.org/abs/2505.18237 Think or not? exploring thinking efficiency in large reasoning models via an information-theoretic lens . Preprint, arXiv:2505.18237
2025 arXiv
-
[229]
Bin Yu, Hang Yuan, Haotian Li, Xueyin Xu, Yuliang Wei, Bailing Wang, Weizhen Qi, and Kai Chen. 2025 a . https://arxiv.org/abs/2505.03469 Long-short chain-of-thought mixture supervised fine-tuning eliciting efficient reasoning in large language models . Preprint, arXiv:2505.03469
2025 arXiv
-
[230]
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2024 a . https://arxiv.org/abs/2311.03099 Language models are super mario: Absorbing abilities from homologous models as a free lunch . Preprint, arXiv:2311.03099
2024 arXiv
-
[231]
Ping Yu, Jing Xu, Jason Weston, and Ilia Kulikov. 2024 b . https://arxiv.org/abs/2407.06023 Distilling system 2 into system 1 . Preprint, arXiv:2407.06023
2024 arXiv
-
[232]
Qifan Yu, Zhenyu He, Sijie Li, Xun Zhou, Jun Zhang, Jingjing Xu, and Di He. 2025 b . https://arxiv.org/abs/2502.08482 Enhancing auto-regressive chain-of-thought through loop-aligned reasoning . Preprint, arXiv:2502.08482
2025 arXiv
-
[233]
Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, et al. 2025 c . https://arxiv.org/abs/2503.14476 Dapo: An open-source llm reinforcement learning system at scale . Preprint, arXiv:2503.14476
2025 arXiv
-
[234]
Danlong Yuan, Tian Xie, Shaohan Huang, Zhuocheng Gong, Huishuai Zhang, Chong Luo, Furu Wei, and Dongyan Zhao. 2025. https://arxiv.org/abs/2505.12284 Efficient rl training for reasoning models via length-aware optimization . Preprint, arXiv:2505.12284
2025
-
[235]
Zhiyuan Zeng, Qinyuan Cheng, Zhangyue Yin, Yunhua Zhou, and Xipeng Qiu. 2025 a . https://arxiv.org/abs/2502.12215 Revisiting the test-time scaling of o1-like models: Do they truly possess test-time scaling capabilities? Preprint, arXiv:2502.12215
2025 arXiv
-
[236]
Zihao Zeng, Xuyao Huang, Boxiu Li, Hao Zhang, and Zhijie Deng. 2025 b . https://arxiv.org/abs/2505.19788 Done is better than perfect: Unlocking efficient reasoning by structured multi-turn decomposition . Preprint, arXiv:2505.19788
2025 arXiv
-
[237]
Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, and He He. 2025 a . https://arxiv.org/abs/2504.05419 Reasoning models know when they're right: Probing hidden states for self-verification . Preprint, arXiv:2504.05419
2025 arXiv
-
[238]
Chong Zhang, Yue Deng, Xiang Lin, Bin Wang, Dianwen Ng, Hai Ye, Xingxuan Li, Yao Xiao, Zhanfeng Mo, Qi Zhang, and Lidong Bing. 2025 b . https://arxiv.org/abs/2505.00551 100 days after deepseek-r1: A survey on replication studies and more directions for reasoning language model...
2025 arXiv
-
[239]
Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, and Juanzi Li. 2025 c . https://arxiv.org/abs/2505.13417 Adaptthink: Reasoning models can learn when to think . Preprint, arXiv:2505.13417
2025 arXiv
-
[240]
Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, and Ningyu Zhang. 2025 d . https://arxiv.org/abs/2502.15589 Lightthinker: Thinking step-by-step compression . Preprint, arXiv:2502.15589
2025
-
[241]
Junyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, and Huan Zhang. 2025 e . https://arxiv.org/abs/2505.24863 Alphaone: Reasoning models thinking slow and fast at test time . Preprint, arXiv:2505.24863
2025 arXiv
-
[242]
Ruiqi Zhang, Changyi Xiao, and Yixin Cao. 2025 f . https://arxiv.org/abs/2506.04182 Long or short cot? investigating instance-level switch of large reasoning models . Preprint, arXiv:2506.04182
2025 arXiv
-
[243]
Shengjia Zhang, Junjie Wu, Jiawei Chen, Changwang Zhang, Xingyu Lou, Wangchunshu Zhou, Sheng Zhou, Can Wang, and Jun Wang. 2025 g . https://arxiv.org/abs/2506.02397 Othink-r1: Intrinsic fast/slow thinking mode switching for over-reasoning mitigation . Preprint, arXiv:2506.02397
2025
-
[244]
Xiaoyun Zhang, Jingqing Ruan, Xing Ma, Yawen Zhu, Haodong Zhao, Hao Li, Jiansong Chen, Ke Zeng, and Xunliang Cai. 2025 h . https://arxiv.org/abs/2505.15400 When to continue thinking: Adaptive thinking mode switching for efficient reasoning . Preprint, arXiv:2505.15400
2025 arXiv
-
[245]
Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, and Minlie Huang. 2025 i . https://arxiv.org/abs/2505.15404 How should we enhance the safety of large reasoning models: An empirical study ....
2025 arXiv
-
[246]
Han Zhao, Haotian Wang, Yiping Peng, Sitong Zhao, Xiaoyu Tian, Shuaiting Chen, Yunjie Ji, and Xiangang Li. 2025 a . https://arxiv.org/abs/2503.19633 1.4 million open-source distilled reasoning dataset to empower large language model training . Preprint, arXiv:2503.19633
2025 arXiv
-
[247]
Haoran Zhao, Yuchen Yan, Yongliang Shen, Haolei Xu, Wenqi Zhang, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, and Yueting Zhuang. 2025 b . https://arxiv.org/abs/2505.14604 Let llms break free from overthinking via self-braking tuning . Preprint, arXiv:2505.14604
2025
-
[248]
Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che, Tat-Seng Chua, and Ting Liu. 2025 c . https://arxiv.org/abs/2503.17979 Trade-offs in large reasoning models: An empirical analysis of deliberative and adaptive reasoning over founda...
2025
-
[249]
Zhengyi Zhao, Shubo Zhang, Zezhong Wang, Huimin Wang, Yutian Zhao, Bin Liang, Yefeng Zheng, Binyang Li, Kam-Fai Wong, and Xian Wu. 2025 d . https://arxiv.org/abs/2505.17427 T ^2 : An adaptive test-time scaling strategy for contextual question answering . Preprint, arXiv:2505.17427
2025 arXiv
-
[250]
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. 2022. https://arxiv.org/abs/2109.00110 Minif2f: a cross-system benchmark for formal olympiad-level mathematics . Preprint, arXiv:2109.00110
2022 arXiv
-
[251]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. https://arxiv.org/abs/2306.05685 Judging llm-as-a-judge with mt-bench and chatbot arena . P...
2023 arXiv
-
[252]
Rongzhi Zhu, Yi Liu, Zequn Sun, Yiwei Wang, and Wei Hu. 2025 a . https://arxiv.org/abs/2505.15276 When can large reasoning models save thinking? mechanistic analysis of behavioral divergence in reasoning . Preprint, arXiv:2505.15276
2025 arXiv
-
[253]
Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng, Xingwei Qu, Jinfa Huang, Dawei Zhu, Hao Wang, Kaiwen Xue, Xuanliang Zhang, Yong Shan, Tianle Cai, Taylor Kergan, Assel Kembay, Andrew Smith, Chenghua Lin, Binh Nguyen, Yuqi Pan, Yuhong Chou, Zefan Cai, Zhenhe Wu, Yongchi Zhao, Tianyu L...
2025 arXiv
-
[254]
Zihao Zhu, Hongbao Zhang, Ruotong Wang, Ke Xu, Siwei Lyu, and Baoyuan Wu. 2025 c . https://arxiv.org/abs/2502.12202 To think or not to think: Exploring the unthinking vulnerability in large reasoning models . Preprint, arXiv:2502.12202
2025 arXiv
-
[255]
Ren Zhuang, Ben Wang, and Shuifa Sun. 2025. https://arxiv.org/abs/2505.08392 Accelerating chain-of-thought reasoning: When goal-gradient importance meets dynamic skipping . Preprint, arXiv:2505.08392
2025 arXiv
-
[256]
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, et al. 2025. https://arxiv.org/abs/2406.15877 Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions . Preprint, arXiv:2406.15877
2025 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.