REVIEW 4 major objections 4 minor 80 references
Compressing the intermediate reasoning text in a spoken language model improves spoken math accuracy while cutting text-token cost to roughly 40% of the full-reasoning baseline.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 11:14 UTC pith:33KBDUXW
load-bearing objection Genuinely new combination—compressed reasoning as joint speech-guidance and reasoning carrier in an SLM—but the headline 3% gain is confounded by an extra training stage and test-set-selected hyperparameters. the 4 major comments →
Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that the text modality in a spoken language model can be repurposed from a full transcript-plus-reasoning scaffold into a single compact reasoning carrier that simultaneously conditions speech generation and encodes the solution path. Concretely, ECoM Reasoning removes the user-text segment from the output sequence and compresses the reasoning span to 40% of its tokens while leaving the assistant's answer text intact. Trained with Progressive Compression—standard CoM, then CoM with explicit full reasoning, then compressed reasoning—the model reaches 60.66% average accuracy on four spoken arithmetic test sets, versus 57.74% for full-reasoning CoM and 39.58% for standa
What carries the argument
The load-bearing mechanism is a compressed reasoning representation, denoted ~Rt, built from two complementary compressions: sentence-level removal of the user's transcribed speech (the model learns to understand speech without emitting a transcript) and token-level pruning of the reasoning trace down to a chosen ratio (40% is best) using a trained token-importance scorer. Training then proceeds through a three-stage curriculum—spoken dialogue, full explicit reasoning, compressed reasoning—so that the model first masters the full chain and later learns to pack the same information into fewer tokens. The assistant's final text is deliberately not compressed, because the ablations show it is t
Load-bearing premise
The 40% compression ratio and the three-stage curriculum were chosen by ranking accuracy on the same four benchmarks used to report the final gains, so the claimed advantage presupposes those choices are not overfits to those test sets.
What would settle it
Choose the compression ratio and training schedule on a held-out development set, then run the same comparison on spoken-math benchmarks not used for model selection; the central claim fails if the ~3-point accuracy advantage over full-reasoning CoM does not survive.
If this is right
- A spoken math model trained with compressed reasoning beats both no-reasoning and full-reasoning baselines, so explicit reasoning and token savings are not in conflict.
- Cutting text tokens from ~96.6 to ~30.2 per response reduces first-speech-token latency from 4.90 s to 1.60 s, bringing compressed reasoning closer to real-time spoken interaction.
- A compression ratio around 40% is a sweet spot: 0% (no reasoning) and 20% (aggressive) both lose accuracy, while 100% wastes tokens; the reasoning traces contain removable redundancy.
- Progressive training is necessary: training ECoM directly from scratch drops average accuracy to 27.36%, and hidden-state analysis shows progressive models retain full-chain information in their representations.
- Leaving the assistant text uncompressed is essential; compressing Yt causes severe accuracy loss in both trained and inference-only settings.
Where Pith is reading between the lines
- If the hidden-state similarity evidence generalizes, the compressed text acts as a sparse 'pointer' to a latent full reasoning chain; a testable extension is probing whether the model can reconstruct the full reasoning from the compressed token when prompted.
- The fixed 40% ratio is a global choice; an adaptive per-problem compression budget—allocating more tokens to harder questions—might push accuracy further and is a natural next experiment.
- The same 'compress the intermediate carrier' recipe may transfer to other modalities, e.g., compressing chain-of-thought in image or audio understanding tasks where intermediate representations are expensive, not just in speech output.
- The authors' sentence-level removal of the user transcript suggests that explicit transcription is not needed for speech understanding once trained; if true, SLMs could omit user-text generation entirely in many settings, not just math.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ECoM Reasoning, a framework for spoken language models that compresses the intermediate textual component so that it serves both as speech guidance and as a compact reasoning carrier. The method removes the user transcript X_t and applies token-level compression to the reasoning trace R_t using LLMLingua-2 importance scores. Training uses a three-stage progressive curriculum: standard CoM, CoM with full reasoning, then ECoM with compressed reasoning. On four spoken math benchmarks (AddSub, MultiArith, SingleEq, SVAMP), the authors report 60.66% average accuracy versus 57.74% for CoM Reasoning, with average generated text tokens reduced from 96.56 to 30.21, and an improved Acc/#Tok of 2.05 versus 0.61. Appendices include comparisons with other efficient-reasoning methods and speech-quality metrics.
Significance. If the result holds, this is a useful contribution: it extends efficient-reasoning methods to spoken language models and shows that a compressed textual representation can carry both reasoning and speech-guidance information. The paper is well-situated in the literature, and the authors provide a fair number of ablations, a latency measurement, speech-quality evaluation, and extended baseline comparisons. The core idea is interesting and timely. However, the central empirical claim is presently not isolated from training-budget effects or from test-set-based hyperparameter selection, so the reported 3-point accuracy gain over CoM Reasoning is not yet established as a property of compression.
major comments (4)
- [Section 3.2 and Section 4.2, Table 2] There is a training-compute confound. ECoM Reasoning is the Stage 3 model, while the CoM Reasoning baseline is the Stage 2 checkpoint. The paper never trains a compute-matched CoM Reasoning model for the same additional number of steps on full-form reasoning data with X_t retained. Table 3 shows that the 100% compression row (which removes X_t but keeps full R_t) reaches only 54.69%, 3.05 points below CoM Reasoning, so removing X_t actually hurts without the additional Stage 3 training. A control that gives CoM Reasoning the same extra training budget is needed to attribute the 60.66 vs. 57.74 gain to compression rather than to additional optimization steps.
- [Section 4.3.1, Table 3 and Section 4.3.2, Table 5] The 40% compression ratio and the three-stage curriculum are selected by ranking average accuracy on the same four test sets used to report the headline numbers. There is no held-out model-selection split, and all numbers come from a single seed. The final 60.66% accuracy and the 2.05 Acc/#Tok efficiency are therefore at risk of being tuned-configuration artifacts. The authors should either report results on a held-out split of the evaluation data, or provide multi-seed variance estimates so the 3-point gain over CoM Reasoning can be assessed statistically.
- [Section 4.2, Table 2 and Section 4.3.1, Table 3] The paper reports no variance or significance testing for any accuracy or #Tok number. The main comparison, 60.66 vs. 57.74, is a 2.92-point difference, and with a single training seed it is impossible to tell whether this exceeds run-to-run noise. Additionally, the #Tok metric uses an IQR-based outlier removal with threshold coefficient 10, which is an arbitrary choice not reported with sensitivity analysis. Given that token efficiency is a central claim, the authors should provide standard deviations across seeds and, ideally, report raw or median token counts.
- [Section 3.1.2 and Table 3, 100% row] The paper states that progressive training makes full removal of X_t harmless, but the evidence does not support this. In Table 3, the 100% compression row (R_t uncompressed, X_t removed) is 54.69%, which is substantially below CoM Reasoning (57.74%). Without an ablation that keeps X_t present under the same Stage 3 training budget and compression ratio, the claim that X_t removal is harmless is not established. This matters because the efficiency gain partly comes from removing the user transcript.
minor comments (4)
- [Abstract] "Improves accuracy by 21%" and "by 3%" should be stated as percentage points: 60.66 vs. 39.58 is +21.08 points absolute, and 60.66 vs. 57.74 is +2.92 points absolute. The current wording overstates the relative improvements.
- [Abstract and Section 4.2] "Using only 40% of the text tokens" is ambiguous. The retained ratio of R_t is 40%, but the measured #Tok is 30.21 vs. 96.56 for CoM Reasoning, i.e., about 31% of the token count. Please clarify which comparison the 40% refers to.
- [Appendix A, Table 7] The training configuration is inconsistent: the text says 2 epochs with each epoch approximately 63,900 steps (about 127,800 total steps), while Table 7 lists Max Steps = 300,000. Clarify whether 300,000 steps is the total across all stages, per stage, or a maximum not reached.
- [Figure 3a] The caption reports non-overlapping 95% confidence intervals for hidden-state similarity, but the paper does not describe how these intervals were computed. Please specify the bootstrap procedure or number of samples.
Circularity Check
Formal probabilistic derivation is self-contained, but the headline empirical results are partly circular because the 40% compression ratio and the three-stage training schedule are selected on the same evaluation benchmarks used to report the final accuracy/token-efficiency claims.
specific steps
-
fitted input called prediction
[Section 4.2 'Main Results' (Table 2) and Section 4.3.1 'Effect of Compression Ratio' (Table 3)]
"For ECoM Reasoning, we report the variant with reasoning compressed to 40% of the original length. ... We find that a compression ratio of 40% achieves the best overall performance, yielding the highest response accuracy and high token efficiency."
The 40% ratio is not an independent, pre-specified design choice; it is the highest-accuracy point selected from a sweep over 100/80/60/40/20/0% on the same four benchmarks (AddSub, MultiArith, SingleEq, SVAMP) that are then used in Table 2 to report the headline 60.66 Acc, 30.21 #Tok, 2.05 Acc/#Tok and the abstract's claim of '3% over CoM with full reasoning traces while using only 40% of the text tokens.' The reported gain is therefore the best point from a model-selection sweep, not an out-of-sample prediction, so the central efficiency/accuracy claim is partly forced by that selection.
-
fitted input called prediction
[Section 4.3.2 'Impact of Training Strategy' (Table 5)]
"Our default three-stage strategy consistently achieves the best response accuracy, suggesting that it provides the most effective curriculum for transferring the benefits of textual guidance and explicit reasoning into the compressed format."
The Progressive Compression training strategy is validated by comparing one-, two-, three-, and four-stage schedules and selecting the 'default three-stage' schedule that performs best on the same evaluation benchmarks used for the final reported numbers. Thus the claim that progressive compression is responsible for the accuracy gain is supported by a schedule selected on the test benchmarks, not by a held-out model-selection split or a pre-registered comparison.
full rationale
The probabilistic formulation in Eqs. (1)-(11) is a modeling expansion, not a derivation whose output equals its input: no equation is defined in terms of the headline result, and no fitted constant is disguised as a predicted variable. The self-citations in the paper (DrVoice, Fun-Audio-Chat, SLAM-LLM, URO-Bench) are tooling, data-construction, and evaluation dependencies, not load-bearing premises from which the compression claim is derived, so they do not constitute circularity. The main circularity-adjacent issue is empirical: the 40% compression ratio and the three-stage curriculum are selected by ranking accuracy on the exact four benchmarks that subsequently produce the abstract's '21%' and '3%' improvements and the 2.05 Acc/#Tok efficiency number. This makes the headline outcome partly a test-set-selected configuration rather than an independent validation. Separately, the absence of a compute-matched CoM Reasoning control is a training-budget confound, which I treat as a correctness risk rather than circularity. Overall, the derivation is self-contained, but the central empirical claim is partially selection-driven, giving a moderate circularity score.
Axiom & Free-Parameter Ledger
free parameters (3)
- Reasoning-token retention ratio for R_t =
0.40
- Curriculum schedule =
3-stage: CoM -> CoM Reasoning -> ECoM
- IQR outlier threshold coefficient for #Tok =
10
axioms (5)
- domain assumption Viterbi-style approximation (Eqs. 2-3): the single most probable user transcript, reasoning text, and assistant text suffice for speech response generation.
- domain assumption LLMLingua-2 token-importance scores trained on text data transfer to spoken math reasoning traces.
- ad hoc to paper Full removal of the user transcript X_t is harmless after progressive training.
- domain assumption Synthesized speech (GPT-4o-mini-TTS and CosyVoice3) adequately represents spoken math questions for training and evaluation.
- domain assumption ASR transcription plus LLM judge correctly measures reasoning accuracy of generated speech.
read the original abstract
Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One important reason is that SLMs reason over purely verbalized mathematical expressions, which are harder to interpret than symbolic text. However, directly transferring text-based reasoning to SLMs is nontrivial due to architectural constraints and the additional computational requirements. To address this challenge, we propose Efficient Chain-of-Modality Reasoning (ECoM Reasoning), the first framework to introduce compressed reasoning into SLMs. By compressing the textual component so that it jointly serves as speech guidance and reasoning representation, ECoM Reasoning improves reasoning accuracy while using a smaller token budget than the standard Chain-of-Modality (CoM) architecture, which generates intermediate text before speech. To train this capability, we further propose Progressive Compression, a curriculum-based strategy that gradually trains the model from full-form reasoning to compressed reasoning. Experiments on spoken mathematical question answering benchmarks show that ECoM Reasoning improves accuracy by 21% over standard CoM without explicit reasoning, and by 3% over CoM with full reasoning traces while using only 40% of the text tokens, demonstrating that it enhances SLM reasoning while remaining inference-efficient.
Figures
Reference graph
Works this paper leans on
-
[1]
Liquid AI. 2025. LFM2 Technical Report.arXiv preprint arXiv:2511.23404(2025)
arXiv 2025
-
[2]
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Lon...
2019
-
[3]
Siddhant Arora, Jinchuan Tian, Hayato Futami, Jee-weon Jung, Jiatong Shi, Yosuke Kashiwagi, Emiru Tsunoo, and Shinji Watanabe. 2025. Chain-of-thought training for open e2e spoken dialogue systems.arXiv preprint arXiv:2506.00722 (2025)
Pith/arXiv arXiv 2025
-
[4]
Siddhant Arora, Jinchuan Tian, Hayato Futami, Jiatong Shi, Yosuke Kashiwagi, Emiru Tsunoo, and Shinji Watanabe. 2025. Chain-of-thought reasoning in streaming full-duplex end-to-end spoken dialogue systems.arXiv preprint arXiv:2510.02066(2025)
arXiv 2025
-
[5]
Vidhisha Balachandran, Jingya Chen, Lingjiao Chen, Shivam Garg, Neel Joshi, Yash Lara, John Langford, Besmira Nushi, Vibhav Vineet, Yue Wu, et al. 2025. Inference-time scaling for complex tasks: Where we stand and what lies ahead. arXiv preprint arXiv:2504.00294(2025)
Pith/arXiv arXiv 2025
-
[6]
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. InProceedings of the 26th annual international conference on machine learning. 41–48
2009
-
[7]
Qiguang Chen, Dengyun Peng, Jinhao Liu, HuiKang Su, Jiannan Guan, Libo Qin, and Wanxiang Che. 2025. Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models. arXiv preprint arXiv:2508.11582(2025)
Pith/arXiv arXiv 2025
-
[8]
Wenxi Chen, Ziyang Ma, Ruiqi Yan, Yuzhe Liang, Xiquan Li, Ruiyang Xu, Zhikang Niu, Yanqiao Zhu, Yifan Yang, Zhanxun Liu, et al . 2025. Slam-omni: Timbre- controllable voice interaction system with single-stage training. InFindings of the Association for Computational Linguistics: ACL 2025. 2262–2282
2025
-
[9]
Yushen Chen et al. 2025. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2025
-
[10]
Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, and Lijuan Wang. 2025. SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models. arXiv preprint arXiv:2510.06917(2025)
arXiv 2025
-
[11]
Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, and Lijuan Wang. 2025. Stitch: Simultaneous thinking and talking with chunked reasoning for spoken language models.arXiv preprint arXiv:2507.15375(2025)
arXiv 2025
-
[12]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168(2021)
Pith/arXiv arXiv 2021
-
[13]
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024. Moshi: a speech-text foundation model for real-time dialogue.arXiv preprint arXiv:2410.00037(2024)
Pith/arXiv arXiv 2024
-
[14]
Yuntian Deng, Yejin Choi, and Stuart Shieber. 2024. From explicit cot to implicit cot: Learning to internalize cot step by step.arXiv preprint arXiv:2405.14838 (2024)
Pith/arXiv arXiv 2024
-
[15]
Yuntian Deng, Kiran Prasad, Roland Fernandez, Paul Smolensky, Vishrav Chaud- hary, and Stuart Shieber. 2023. Implicit chain of thought reasoning via knowledge distillation.arXiv preprint arXiv:2311.01460(2023)
Pith/arXiv arXiv 2023
-
[16]
Ding Ding, Zeqian Ju, Yichong Leng, Songxiang Liu, Tong Liu, Zeyu Shang, Kai Shen, Wei Song, Xu Tan, Heyi Tang, et al . 2025. Kimi-audio technical report. arXiv preprint arXiv:2504.18425(2025)
Pith/arXiv arXiv 2025
-
[17]
Yexing Du, Ziyang Ma, Yifan Yang, Keqi Deng, Xie Chen, Bo Yang, Yang Xiang, Ming Liu, and Bing Qin. 2024. Cot-st: Enhancing llm-based speech translation with multimodal chain-of-thought.arXiv preprint arXiv:2409.19510(2024)
Pith/arXiv arXiv 2024
-
[18]
Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu, Tianyu Zhao, Hao Wang, Xiang Lv, Hui Wang, Chongjia Ni, Xian Shi, et al. 2025. Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training.arXiv preprint arXiv:2505.17589(2025)
Pith/arXiv arXiv 2025
-
[19]
Qingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang, and Yang Feng. 2025. LLaMA-omni 2: LLM-based real-time spoken chatbot with autoregressive stream- ing speech synthesis. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 18617–18629
2025
-
[20]
Hasan Abed Al Kader Hammoud, Kumail Alhamoud, Abed Hammoud, Elie Bou- Zeid, Marzyeh Ghassemi, and Bernard Ghanem. 2025. Train long, think short: Curriculum learning for efficient reasoning.arXiv preprint arXiv:2508.08940 (2025)
Pith/arXiv arXiv 2025
-
[21]
Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kush- man. 2014. Learning to solve arithmetic word problems with verb categorization. InProceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 523–533
2014
-
[22]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card.arXiv preprint arXiv:2410.21276(2024)
Pith/arXiv arXiv 2024
-
[23]
Sieun Hyeon, Kyudan Jung, Jaehee Won, Nam-Joon Kim, Hyun Gon Ryu, Hyuk- Jae Lee, and Jaeyoung Do. 2025. Mathspeech: Leveraging small lms for accurate conversion in mathematical speech-to-formula. InProceedings of the AAAI Con- ference on Artificial Intelligence, Vol. 39. 24194–24202
2025
-
[24]
Huiqiang Jiang et al. 2023. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[25]
Huiqiang Jiang et al. 2024. LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2024
-
[26]
Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics3 (2015), 585–597
2015
-
[27]
Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, et al. 2024. Numi- namath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository13, 9 (2024), 9
2024
-
[28]
Jijie Li, Li Du, Hanyu Zhao, Bo-wen Zhang, Liangdong Wang, Boyan Gao, Guang Liu, and Yonghua Lin. 2025. Infinity instruct: Scaling instruction selection and synthesis to enhance language models.arXiv preprint arXiv:2506.11116(2025)
Pith/arXiv arXiv 2025
-
[29]
Jindong Li, Yali Fu, Li Fan, Jiahong Liu, Yao Shu, Chengwei Qin, Menglin Yang, Irwin King, and Rex Ying. 2025. Implicit reasoning in large language models: A comprehensive survey.arXiv preprint arXiv:2509.02350(2025)
Pith/arXiv arXiv 2025
-
[30]
Shufan Li and Aditya Grover. 2025. Predgen: Accelerated inference of large language models through input-time speculation for real-time speech interaction. arXiv preprint arXiv:2506.15556(2025)
arXiv 2025
-
[31]
Tianpeng Li, Jun Liu, Tao Zhang, Yuanbo Fang, Da Pan, Mingrui Wang, Zheng Liang, Zehuan Li, Mingan Lin, Guosheng Dong, et al. 2025. Baichuan-audio: A uni- fied framework for end-to-end speech interaction.arXiv preprint arXiv:2502.17239 (2025)
Pith/arXiv arXiv 2025
-
[32]
Zeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen, Zhijian Xu, Yingying Cheng, Fan Zhang, and Qiang Xu. 2026. Making Slow Thinking Faster: Com- pressing LLM Chain-of-Thought via Step Entropy. InThe Fourteenth International Conference on Learning Representations
2026
-
[33]
Junhong Lin, Xinyue Zeng, Jie Zhu, Song Wang, Julian Shun, Jun Wu, and Dawei Zhou. 2025. Plan and budget: Effective and efficient test-time scaling on large language model reasoning.arXiv preprint arXiv:2505.16122(2025)
arXiv 2025
-
[34]
Yueqian Lin, Zhengmian Hu, Qinsi Wang, Yudong Liu, Hengfan Zhang, Jayaku- mar Subramanian, Nikos Vlassis, Hai Helen Li, and Yiran Chen. 2025. Voice evaluation of reasoning ability: Diagnosing the modality-induced performance gap.arXiv preprint arXiv:2509.26542(2025)
arXiv 2025
-
[35]
Siyuan Liu, Jiahui Xu, Feng Jiang, Kuang Wang, Zefeng Zhao, Chu-Ren Huang, Jinghang Gu, Changqing Yin, and Haizhou Li. 2026. Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue Systems.arXiv preprint arXiv:2602.23266(2026)
arXiv 2026
-
[36]
Ziyang Ma, Zhuo Chen, Yuping Wang, Eng Siong Chng, and Xie Chen. 2025. Audio-cot: Exploring chain-of-thought reasoning in large audio language model. arXiv preprint arXiv:2501.07246(2025)
Pith/arXiv arXiv 2025
-
[37]
Ziyang Ma, Guanrou Yang, Wenxi Chen, Zhifu Gao, Yexing Du, Xiquan Li, Zhisheng Zheng, Haina Zhu, Jianheng Zhuo, Zheshu Song, et al. 2026. SLAM- LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing.IEEE Journal of Selected Topics in Signal Processing(2026)
2026
-
[38]
Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaro- engchai, Soroosh Mariooryad, Ehud Rivlin, RJ Skerry-Ryan, and Michelle Tadmor Ramanovich. 2023. Spoken question answering and speech continuation using spectrogram-powered llm.arXiv preprint arXiv:2305.15255(2023)
Pith/arXiv arXiv 2023
-
[39]
Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, et al. 2024. Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression. In Findings of the Association for Computational Linguistics: ACL 2024. 963–981. MM ’26, November 10–14, 2026, Rio de Jan...
2024
-
[40]
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems?. InProceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies. 2080–2094
2021
-
[41]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. InInternational conference on machine learning. PMLR, 28492–28518
2023
-
[42]
Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. InProceedings of the 2015 conference on empirical methods in natural language processing. 1743–1752
2015
-
[43]
Hejian Sang, Yuanda Xu, Zhengze Zhou, Ran He, Zhipeng Wang, and Jiachen Sun. 2026. On-Policy Self-Distillation for Reasoning Compression.arXiv preprint arXiv:2603.05433(2026)
Pith/arXiv arXiv 2026
-
[44]
Xuan Shen et al . 2025. Efficient Reasoning with Hidden Thinking.CoRR abs/2501.19201 (2025). arXiv:2501.19201 doi:10.48550/ARXIV.2501.19201
-
[45]
Qundong Shi, Jie Zhou, Biyuan Lin, Junbo Cui, Guoyang Zeng, Yixuan Zhou, Ziyang Wang, Xin Liu, Zhen Luo, Yudong Wang, et al. 2026. UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models. arXiv preprint arXiv:2601.01373(2026)
arXiv 2026
-
[46]
Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou, SK Bong, Yashesh Gaur, Jay Mahadeokar, Ozlem Kalinli, and Mike Seltzer. 2025. Can Speech LLMs Think while Listening?arXiv preprint arXiv:2510.07497(2025)
arXiv 2025
-
[47]
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Na Zou, et al . 2025. Stop overthinking: A survey on efficient reasoning for large language models.arXiv preprint arXiv:2503.16419(2025)
Pith/arXiv arXiv 2025
-
[48]
Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye, Huadai Liu, Honggang Zhang, Wei Xue, and Yike Guo. 2024. Both ears wide open: Towards language-driven spatial audio generation.arXiv preprint arXiv:2410.10676(2024)
Pith/arXiv arXiv 2024
-
[49]
Chao-Hong Tan, Qian Chen, Wen Wang, Chong Deng, Qinglin Zhang, Luyao Cheng, Hai Yu, Xin Zhang, Xiang Lv, Tianyu Zhao, et al. 2025. DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representa- tions.arXiv preprint arXiv:2506.09349(2025)
arXiv 2025
-
[50]
Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Ruihua Song, and Jian Luan
-
[51]
Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. https://qwenlm. github.io/blog/qwen2.5/
2024
-
[52]
Tongyi Fun Team, Qian Chen, Luyao Cheng, Chong Deng, Xiangang Li, Jiaqing Liu, Chao-Hong Tan, Wen Wang, Junhao Xu, Jieping Ye, et al. 2025. Fun-Audio- Chat Technical Report.arXiv preprint arXiv:2512.20156(2025)
arXiv 2025
-
[53]
Chen Wang, Minpeng Liao, Zhongqiang Huang, Junhong Wu, Chengqing Zong, and Jiajun Zhang. 2024. Blsp-emo: Towards empathetic large speech-language models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 19186–19199
2024
-
[54]
Chaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu, Yan Lu, Jinyu Li, and Zhizheng Wu. 2026. Closing the Modality Reasoning Gap for Speech Large Language Models.arXiv preprint arXiv:2601.05543(2026)
Pith/arXiv arXiv 2026
-
[55]
Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, and Hao Fei. 2025. Multimodal chain-of-thought reasoning: A comprehensive survey.arXiv preprint arXiv:2503.12605(2025)
Pith/arXiv arXiv 2025
-
[56]
Chengwei Wei, Bin Wang, Jung-jae Kim, and Nancy F Chen. 2025. Towards spoken mathematical reasoning: Benchmarking speech-based models over multi- faceted math problems.arXiv preprint arXiv:2505.15000(2025)
Pith/arXiv arXiv 2025
-
[57]
Donghang Wu, Haoyang Zhang, Jun Chen, Hexin Liu, Eng Siong Chng, Fei Tian, Xuerui Yang, Xiangyu Zhang, Daxin Jiang, Gang Yu, et al . 2025. Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models.arXiv preprint arXiv:2510.09592(2025)
Pith/arXiv arXiv 2025
-
[58]
Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, and Wenjie Li. 2025. Tokenskip: Controllable chain-of-thought compression in llms. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 3351–3363
2025
-
[59]
Jingran Xie, Shun Lei, Yue Yu, Yang Xiang, Hui Wang, Xixin Wu, and Zhiyong Wu
-
[60]
Zhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu, Shuicheng Yan, and Chun- yan Miao. 2025. Audio-reasoner: Improving reasoning capability in large audio language models.arXiv preprint arXiv:2503.02318(2025)
arXiv 2025
-
[61]
InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Leveraging chain of thought towards empathetic spoken dialogue without corresponding question-answering data. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2025
-
[62]
Zhifei Xie and Changqiao Wu. 2024. Mini-omni: Language models can hear, talk while thinking in streaming.arXiv preprint arXiv:2408.16725(2024)
Pith/arXiv arXiv 2024
-
[63]
Zhifei Xie, Ziyang Ma, Zihang Liu, Kaiyu Pang, Hongyu Li, Jialin Zhang, Yue Liao, Deheng Ye, Chunyan Miao, and Shuicheng Yan. 2025. Mini-omni-reasoner: Token- level thinking-in-speaking in large speech models.arXiv preprint arXiv:2508.15827 (2025)
arXiv 2025
-
[64]
Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. 2025. Chain of draft: Thinking faster by writing less.arXiv preprint arXiv:2502.18600(2025)
Pith/arXiv arXiv 2025
-
[65]
Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, et al. 2025. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765(2025)
Pith/arXiv arXiv 2025
-
[66]
Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Mengzhe Chen, Qian Chen, and Lei Xie. 2024. E-chat: Emotion-sensitive spoken dialogue system with large language models. In2024 IEEE 14th International Symposium on Chinese Spoken Language Processing (ISCSLP). IEEE, 586–590
2024
-
[67]
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. 2024. Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing.arXiv preprint arXiv:2406.08464 (2024)
Pith/arXiv arXiv 2024
-
[68]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Cheng- peng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfen...
Pith/arXiv arXiv 2024
-
[69]
Ruiqi Yan, Xiquan Li, Wenxi Chen, Zhikang Niu, Chen Yang, Ziyang Ma, Kai Yu, and Xie Chen. 2025. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models.arXiv preprint arXiv:2502.17810(2025)
Pith/arXiv arXiv 2025
-
[70]
Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang, Shengmin Jiang, Lei Zhao, Yuxiao Dong, and Jie Tang. 2024. Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot.arXiv preprint arXiv:2412.02612(2024)
Pith/arXiv arXiv 2024
-
[71]
Wenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Guangzhi Sun, Lu Lu, Yuxuan Wang, and Chao Zhang. 2024. Salmonn-omni: A codec-free llm for full-duplex speech understanding and generation.arXiv preprint arXiv:2411.18138(2024)
Pith/arXiv arXiv 2024
-
[72]
Dong Zhang, Gang Wang, Jinlong Xue, Kai Fang, Liang Zhao, Rui Ma, Shuhuai Ren, Shuo Liu, Tao Guo, Weiji Zhuang, et al. 2025. MiMo-Audio: Audio Language Models are Few-Shot Learners.arXiv preprint arXiv:2512.23808(2025)
arXiv 2025
-
[73]
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. InFindings of the Association for Computational Linguistics: EMNLP 2023. 15757–15773
2023
-
[74]
Zhen Zhang, Xuehai He, Weixiang Yan, Ao Shen, Chenyang Zhao, Shuohang Wang, Yelong Shen, and Xin Eric Wang. 2025. Soft thinking: Unlocking the reason- ing potential of llms in continuous concept space.arXiv preprint arXiv:2505.15778 (2025)
Pith/arXiv arXiv 2025
-
[75]
Dong Zhang, Xin Zhang, Jun Zhan, Shimin Li, Yaqian Zhou, and Xipeng Qiu
-
[76]
Yufan Zhuang, Liyuan Liu, Chandan Singh, Jingbo Shang, and Jianfeng Gao. 2025. Text generation beyond discrete token sampling.arXiv preprint arXiv:2505.14827 (2025)
arXiv 2025
-
[77]
Wenhao Zou, Yuwei Miao, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, and Jingwen Xu. 2026. LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning.arXiv preprint arXiv:2601.19952(2026). Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoke...
arXiv 2026
-
[78]
Xingjian Zhao, Zhe Xu, Qinyuan Cheng, Zhaoye Fei, Luozhijie Jin, Yang Wang, Hanfu Chen, Yaozhou Jiang, Qinghui Gao, Ke Chen, et al. 2025. MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance.arXiv preprint arXiv:2510.00499(2025)
arXiv 2025
-
[2024]
Speechgpt-gen: Scaling chain-of-information speech generation.arXiv preprint arXiv:2401.13527(2024)
Pith/arXiv arXiv 2024
-
[2025]
Think silently, think fast: Dynamic latent compression of llm reasoning chains.arXiv preprint arXiv:2505.16552(2025)
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.