Pith. sign in

REVIEW 4 major objections 4 minor 80 references

Compressing the intermediate reasoning text in a spoken language model improves spoken math accuracy while cutting text-token cost to roughly 40% of the full-reasoning baseline.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 11:14 UTC pith:33KBDUXW

load-bearing objection Genuinely new combination—compressed reasoning as joint speech-guidance and reasoning carrier in an SLM—but the headline 3% gain is confounded by an extra training stage and test-set-selected hyperparameters. the 4 major comments →

arxiv 2607.19932 v1 pith:33KBDUXW submitted 2026-07-22 cs.CL cs.SD

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

classification cs.CL cs.SD
keywords spoken language modelschain-of-modality reasoningcompact reasoningcurriculum learningspoken math QAtoken efficiencyinference cost
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that a spoken language model can solve spoken math problems more accurately and at lower token cost by compressing the textual reasoning it generates before speech, so that the same text both guides speech and carries the reasoning. The proposed ECoM Reasoning framework, built on the Chain-of-Modality architecture, drops the full user transcript and keeps only the most important reasoning tokens, selected by an importance-scoring model. To make this learnable, the authors train in three stages—plain spoken dialogue, full reasoning traces, then compressed reasoning—so the model internalizes full reasoning before being forced to be concise. On four spoken arithmetic benchmarks, the compressed model beats the full-reasoning baseline by about 3 points in accuracy while using roughly 40 percent of the text tokens, and beats the no-reasoning baseline by about 21 points. The result matters because it suggests stronger spoken reasoning does not require a larger inference budget.

Core claim

The central discovery is that the text modality in a spoken language model can be repurposed from a full transcript-plus-reasoning scaffold into a single compact reasoning carrier that simultaneously conditions speech generation and encodes the solution path. Concretely, ECoM Reasoning removes the user-text segment from the output sequence and compresses the reasoning span to 40% of its tokens while leaving the assistant's answer text intact. Trained with Progressive Compression—standard CoM, then CoM with explicit full reasoning, then compressed reasoning—the model reaches 60.66% average accuracy on four spoken arithmetic test sets, versus 57.74% for full-reasoning CoM and 39.58% for standa

What carries the argument

The load-bearing mechanism is a compressed reasoning representation, denoted ~Rt, built from two complementary compressions: sentence-level removal of the user's transcribed speech (the model learns to understand speech without emitting a transcript) and token-level pruning of the reasoning trace down to a chosen ratio (40% is best) using a trained token-importance scorer. Training then proceeds through a three-stage curriculum—spoken dialogue, full explicit reasoning, compressed reasoning—so that the model first masters the full chain and later learns to pack the same information into fewer tokens. The assistant's final text is deliberately not compressed, because the ablations show it is t

Load-bearing premise

The 40% compression ratio and the three-stage curriculum were chosen by ranking accuracy on the same four benchmarks used to report the final gains, so the claimed advantage presupposes those choices are not overfits to those test sets.

What would settle it

Choose the compression ratio and training schedule on a held-out development set, then run the same comparison on spoken-math benchmarks not used for model selection; the central claim fails if the ~3-point accuracy advantage over full-reasoning CoM does not survive.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A spoken math model trained with compressed reasoning beats both no-reasoning and full-reasoning baselines, so explicit reasoning and token savings are not in conflict.
  • Cutting text tokens from ~96.6 to ~30.2 per response reduces first-speech-token latency from 4.90 s to 1.60 s, bringing compressed reasoning closer to real-time spoken interaction.
  • A compression ratio around 40% is a sweet spot: 0% (no reasoning) and 20% (aggressive) both lose accuracy, while 100% wastes tokens; the reasoning traces contain removable redundancy.
  • Progressive training is necessary: training ECoM directly from scratch drops average accuracy to 27.36%, and hidden-state analysis shows progressive models retain full-chain information in their representations.
  • Leaving the assistant text uncompressed is essential; compressing Yt causes severe accuracy loss in both trained and inference-only settings.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the hidden-state similarity evidence generalizes, the compressed text acts as a sparse 'pointer' to a latent full reasoning chain; a testable extension is probing whether the model can reconstruct the full reasoning from the compressed token when prompted.
  • The fixed 40% ratio is a global choice; an adaptive per-problem compression budget—allocating more tokens to harder questions—might push accuracy further and is a natural next experiment.
  • The same 'compress the intermediate carrier' recipe may transfer to other modalities, e.g., compressing chain-of-thought in image or audio understanding tasks where intermediate representations are expensive, not just in speech output.
  • The authors' sentence-level removal of the user transcript suggests that explicit transcription is not needed for speech understanding once trained; if true, SLMs could omit user-text generation entirely in many settings, not just math.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes ECoM Reasoning, a framework for spoken language models that compresses the intermediate textual component so that it serves both as speech guidance and as a compact reasoning carrier. The method removes the user transcript X_t and applies token-level compression to the reasoning trace R_t using LLMLingua-2 importance scores. Training uses a three-stage progressive curriculum: standard CoM, CoM with full reasoning, then ECoM with compressed reasoning. On four spoken math benchmarks (AddSub, MultiArith, SingleEq, SVAMP), the authors report 60.66% average accuracy versus 57.74% for CoM Reasoning, with average generated text tokens reduced from 96.56 to 30.21, and an improved Acc/#Tok of 2.05 versus 0.61. Appendices include comparisons with other efficient-reasoning methods and speech-quality metrics.

Significance. If the result holds, this is a useful contribution: it extends efficient-reasoning methods to spoken language models and shows that a compressed textual representation can carry both reasoning and speech-guidance information. The paper is well-situated in the literature, and the authors provide a fair number of ablations, a latency measurement, speech-quality evaluation, and extended baseline comparisons. The core idea is interesting and timely. However, the central empirical claim is presently not isolated from training-budget effects or from test-set-based hyperparameter selection, so the reported 3-point accuracy gain over CoM Reasoning is not yet established as a property of compression.

major comments (4)
  1. [Section 3.2 and Section 4.2, Table 2] There is a training-compute confound. ECoM Reasoning is the Stage 3 model, while the CoM Reasoning baseline is the Stage 2 checkpoint. The paper never trains a compute-matched CoM Reasoning model for the same additional number of steps on full-form reasoning data with X_t retained. Table 3 shows that the 100% compression row (which removes X_t but keeps full R_t) reaches only 54.69%, 3.05 points below CoM Reasoning, so removing X_t actually hurts without the additional Stage 3 training. A control that gives CoM Reasoning the same extra training budget is needed to attribute the 60.66 vs. 57.74 gain to compression rather than to additional optimization steps.
  2. [Section 4.3.1, Table 3 and Section 4.3.2, Table 5] The 40% compression ratio and the three-stage curriculum are selected by ranking average accuracy on the same four test sets used to report the headline numbers. There is no held-out model-selection split, and all numbers come from a single seed. The final 60.66% accuracy and the 2.05 Acc/#Tok efficiency are therefore at risk of being tuned-configuration artifacts. The authors should either report results on a held-out split of the evaluation data, or provide multi-seed variance estimates so the 3-point gain over CoM Reasoning can be assessed statistically.
  3. [Section 4.2, Table 2 and Section 4.3.1, Table 3] The paper reports no variance or significance testing for any accuracy or #Tok number. The main comparison, 60.66 vs. 57.74, is a 2.92-point difference, and with a single training seed it is impossible to tell whether this exceeds run-to-run noise. Additionally, the #Tok metric uses an IQR-based outlier removal with threshold coefficient 10, which is an arbitrary choice not reported with sensitivity analysis. Given that token efficiency is a central claim, the authors should provide standard deviations across seeds and, ideally, report raw or median token counts.
  4. [Section 3.1.2 and Table 3, 100% row] The paper states that progressive training makes full removal of X_t harmless, but the evidence does not support this. In Table 3, the 100% compression row (R_t uncompressed, X_t removed) is 54.69%, which is substantially below CoM Reasoning (57.74%). Without an ablation that keeps X_t present under the same Stage 3 training budget and compression ratio, the claim that X_t removal is harmless is not established. This matters because the efficiency gain partly comes from removing the user transcript.
minor comments (4)
  1. [Abstract] "Improves accuracy by 21%" and "by 3%" should be stated as percentage points: 60.66 vs. 39.58 is +21.08 points absolute, and 60.66 vs. 57.74 is +2.92 points absolute. The current wording overstates the relative improvements.
  2. [Abstract and Section 4.2] "Using only 40% of the text tokens" is ambiguous. The retained ratio of R_t is 40%, but the measured #Tok is 30.21 vs. 96.56 for CoM Reasoning, i.e., about 31% of the token count. Please clarify which comparison the 40% refers to.
  3. [Appendix A, Table 7] The training configuration is inconsistent: the text says 2 epochs with each epoch approximately 63,900 steps (about 127,800 total steps), while Table 7 lists Max Steps = 300,000. Clarify whether 300,000 steps is the total across all stages, per stage, or a maximum not reached.
  4. [Figure 3a] The caption reports non-overlapping 95% confidence intervals for hidden-state similarity, but the paper does not describe how these intervals were computed. Please specify the bootstrap procedure or number of samples.

Circularity Check

2 steps flagged

Formal probabilistic derivation is self-contained, but the headline empirical results are partly circular because the 40% compression ratio and the three-stage training schedule are selected on the same evaluation benchmarks used to report the final accuracy/token-efficiency claims.

specific steps
  1. fitted input called prediction [Section 4.2 'Main Results' (Table 2) and Section 4.3.1 'Effect of Compression Ratio' (Table 3)]
    "For ECoM Reasoning, we report the variant with reasoning compressed to 40% of the original length. ... We find that a compression ratio of 40% achieves the best overall performance, yielding the highest response accuracy and high token efficiency."

    The 40% ratio is not an independent, pre-specified design choice; it is the highest-accuracy point selected from a sweep over 100/80/60/40/20/0% on the same four benchmarks (AddSub, MultiArith, SingleEq, SVAMP) that are then used in Table 2 to report the headline 60.66 Acc, 30.21 #Tok, 2.05 Acc/#Tok and the abstract's claim of '3% over CoM with full reasoning traces while using only 40% of the text tokens.' The reported gain is therefore the best point from a model-selection sweep, not an out-of-sample prediction, so the central efficiency/accuracy claim is partly forced by that selection.

  2. fitted input called prediction [Section 4.3.2 'Impact of Training Strategy' (Table 5)]
    "Our default three-stage strategy consistently achieves the best response accuracy, suggesting that it provides the most effective curriculum for transferring the benefits of textual guidance and explicit reasoning into the compressed format."

    The Progressive Compression training strategy is validated by comparing one-, two-, three-, and four-stage schedules and selecting the 'default three-stage' schedule that performs best on the same evaluation benchmarks used for the final reported numbers. Thus the claim that progressive compression is responsible for the accuracy gain is supported by a schedule selected on the test benchmarks, not by a held-out model-selection split or a pre-registered comparison.

full rationale

The probabilistic formulation in Eqs. (1)-(11) is a modeling expansion, not a derivation whose output equals its input: no equation is defined in terms of the headline result, and no fitted constant is disguised as a predicted variable. The self-citations in the paper (DrVoice, Fun-Audio-Chat, SLAM-LLM, URO-Bench) are tooling, data-construction, and evaluation dependencies, not load-bearing premises from which the compression claim is derived, so they do not constitute circularity. The main circularity-adjacent issue is empirical: the 40% compression ratio and the three-stage curriculum are selected by ranking accuracy on the exact four benchmarks that subsequently produce the abstract's '21%' and '3%' improvements and the 2.05 Acc/#Tok efficiency number. This makes the headline outcome partly a test-set-selected configuration rather than an independent validation. Separately, the absence of a compute-matched CoM Reasoning control is a training-budget confound, which I treat as a correctness risk rather than circularity. Overall, the derivation is self-contained, but the central empirical claim is partially selection-driven, giving a moderate circularity score.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The central empirical claim rests on externally fitted scoring and several domain assumptions rather than on many paper-specific free parameters. The main fitted/ad hoc numbers are the compression ratio and curriculum schedule, both selected on the evaluation benchmarks, plus the IQR trimming coefficient that shapes reported token counts. No new physical or architectural entity is introduced.

free parameters (3)
  • Reasoning-token retention ratio for R_t = 0.40
    Selected as the best average accuracy among {100,80,60,40,20,0}% on the four test benchmarks (Table 3).
  • Curriculum schedule = 3-stage: CoM -> CoM Reasoning -> ECoM
    Chosen over one-, two-, four-stage, and reversed-order schedules by accuracy on the same test sets (Table 5).
  • IQR outlier threshold coefficient for #Tok = 10
    Used to trim extreme outliers before reporting average generated text tokens; the choice is not justified or swept, and it directly shapes the headline token-efficiency numbers.
axioms (5)
  • domain assumption Viterbi-style approximation (Eqs. 2-3): the single most probable user transcript, reasoning text, and assistant text suffice for speech response generation.
    Underlies the CoM/ECoM generation chain. If exact marginalization matters, the deterministic MAP approximation could misrepresent the posterior.
  • domain assumption LLMLingua-2 token-importance scores trained on text data transfer to spoken math reasoning traces.
    Section 3.1.2/Eq. 8 uses LLMLingua-2's P(x_i | x_<=n; theta_MB) as the importance oracle; no SLM-specific re-training or quantitative validation is provided beyond manual inspection.
  • ad hoc to paper Full removal of the user transcript X_t is harmless after progressive training.
    Stated as an empirical finding in Section 3.1.2, this removal is central to the token savings and is not isolated in an ablation against a variant that keeps X_t.
  • domain assumption Synthesized speech (GPT-4o-mini-TTS and CosyVoice3) adequately represents spoken math questions for training and evaluation.
    All training and test speech is synthetic per Section 4.1; generalization to natural spontaneous speech is unverified.
  • domain assumption ASR transcription plus LLM judge correctly measures reasoning accuracy of generated speech.
    The URO-Bench protocol in Section 4.2 transcribes generated speech with Whisper-large-v3 and evaluates with GPT-mini, coupling reasoning accuracy to speech intelligibility and ASR behavior.

pith-pipeline@v1.3.0-alltime-deepseek · 19732 in / 12514 out tokens · 122550 ms · 2026-08-01T11:14:42.850161+00:00 · methodology

0 comments
read the original abstract

Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One important reason is that SLMs reason over purely verbalized mathematical expressions, which are harder to interpret than symbolic text. However, directly transferring text-based reasoning to SLMs is nontrivial due to architectural constraints and the additional computational requirements. To address this challenge, we propose Efficient Chain-of-Modality Reasoning (ECoM Reasoning), the first framework to introduce compressed reasoning into SLMs. By compressing the textual component so that it jointly serves as speech guidance and reasoning representation, ECoM Reasoning improves reasoning accuracy while using a smaller token budget than the standard Chain-of-Modality (CoM) architecture, which generates intermediate text before speech. To train this capability, we further propose Progressive Compression, a curriculum-based strategy that gradually trains the model from full-form reasoning to compressed reasoning. Experiments on spoken mathematical question answering benchmarks show that ECoM Reasoning improves accuracy by 21% over standard CoM without explicit reasoning, and by 3% over CoM with full reasoning traces while using only 40% of the text tokens, demonstrating that it enhances SLM reasoning while remaining inference-efficient.

Figures

Figures reproduced from arXiv: 2607.19932 by Chao-Hong Tan, Pengchao Feng, Qian Chen, Wen Wang, Xiangang Li, Xie Chen.

Figure 1
Figure 1. Figure 1: Overview of (a) the standard CoM framework and (b) the ECoM Reasoning framework, both built upon the Chain-of-Modality (CoM) architecture. By compressing the textual component, ECoM Reasoning enables the intermediate text to simultaneously guide speech generation and carry the core reasoning process, while maintaining low inference cost. more compact rationales or by adaptively constraining reasoning durin… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the proposed progressive training pipeline for [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Case studies on progressive training and assistant [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

80 extracted references · 35 linked inside Pith

  1. [1]

    Liquid AI. 2025. LFM2 Technical Report.arXiv preprint arXiv:2511.23404(2025)

  2. [2]

    Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Lon...

  3. [3]

    Siddhant Arora, Jinchuan Tian, Hayato Futami, Jee-weon Jung, Jiatong Shi, Yosuke Kashiwagi, Emiru Tsunoo, and Shinji Watanabe. 2025. Chain-of-thought training for open e2e spoken dialogue systems.arXiv preprint arXiv:2506.00722 (2025)

  4. [4]

    Siddhant Arora, Jinchuan Tian, Hayato Futami, Jiatong Shi, Yosuke Kashiwagi, Emiru Tsunoo, and Shinji Watanabe. 2025. Chain-of-thought reasoning in streaming full-duplex end-to-end spoken dialogue systems.arXiv preprint arXiv:2510.02066(2025)

  5. [5]

    Vidhisha Balachandran, Jingya Chen, Lingjiao Chen, Shivam Garg, Neel Joshi, Yash Lara, John Langford, Besmira Nushi, Vibhav Vineet, Yue Wu, et al. 2025. Inference-time scaling for complex tasks: Where we stand and what lies ahead. arXiv preprint arXiv:2504.00294(2025)

  6. [6]

    Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. InProceedings of the 26th annual international conference on machine learning. 41–48

  7. [7]

    Qiguang Chen, Dengyun Peng, Jinhao Liu, HuiKang Su, Jiannan Guan, Libo Qin, and Wanxiang Che. 2025. Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models. arXiv preprint arXiv:2508.11582(2025)

  8. [8]

    Wenxi Chen, Ziyang Ma, Ruiqi Yan, Yuzhe Liang, Xiquan Li, Ruiyang Xu, Zhikang Niu, Yanqiao Zhu, Yifan Yang, Zhanxun Liu, et al . 2025. Slam-omni: Timbre- controllable voice interaction system with single-stage training. InFindings of the Association for Computational Linguistics: ACL 2025. 2262–2282

  9. [9]

    Yushen Chen et al. 2025. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

  10. [10]

    Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, and Lijuan Wang. 2025. SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models. arXiv preprint arXiv:2510.06917(2025)

  11. [11]

    Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin, Kevin Lin, Shujie Liu, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, and Lijuan Wang. 2025. Stitch: Simultaneous thinking and talking with chunked reasoning for spoken language models.arXiv preprint arXiv:2507.15375(2025)

  12. [12]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168(2021)

  13. [13]

    Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024. Moshi: a speech-text foundation model for real-time dialogue.arXiv preprint arXiv:2410.00037(2024)

  14. [14]

    Yuntian Deng, Yejin Choi, and Stuart Shieber. 2024. From explicit cot to implicit cot: Learning to internalize cot step by step.arXiv preprint arXiv:2405.14838 (2024)

  15. [15]

    Yuntian Deng, Kiran Prasad, Roland Fernandez, Paul Smolensky, Vishrav Chaud- hary, and Stuart Shieber. 2023. Implicit chain of thought reasoning via knowledge distillation.arXiv preprint arXiv:2311.01460(2023)

  16. [16]

    Ding Ding, Zeqian Ju, Yichong Leng, Songxiang Liu, Tong Liu, Zeyu Shang, Kai Shen, Wei Song, Xu Tan, Heyi Tang, et al . 2025. Kimi-audio technical report. arXiv preprint arXiv:2504.18425(2025)

  17. [17]

    Yexing Du, Ziyang Ma, Yifan Yang, Keqi Deng, Xie Chen, Bo Yang, Yang Xiang, Ming Liu, and Bing Qin. 2024. Cot-st: Enhancing llm-based speech translation with multimodal chain-of-thought.arXiv preprint arXiv:2409.19510(2024)

  18. [18]

    Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu, Tianyu Zhao, Hao Wang, Xiang Lv, Hui Wang, Chongjia Ni, Xian Shi, et al. 2025. Cosyvoice 3: Towards in-the-wild speech generation via scaling-up and post-training.arXiv preprint arXiv:2505.17589(2025)

  19. [19]

    Qingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang, and Yang Feng. 2025. LLaMA-omni 2: LLM-based real-time spoken chatbot with autoregressive stream- ing speech synthesis. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 18617–18629

  20. [20]

    Hasan Abed Al Kader Hammoud, Kumail Alhamoud, Abed Hammoud, Elie Bou- Zeid, Marzyeh Ghassemi, and Bernard Ghanem. 2025. Train long, think short: Curriculum learning for efficient reasoning.arXiv preprint arXiv:2508.08940 (2025)

  21. [21]

    Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kush- man. 2014. Learning to solve arithmetic word problems with verb categorization. InProceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 523–533

  22. [22]

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card.arXiv preprint arXiv:2410.21276(2024)

  23. [23]

    Sieun Hyeon, Kyudan Jung, Jaehee Won, Nam-Joon Kim, Hyun Gon Ryu, Hyuk- Jae Lee, and Jaeyoung Do. 2025. Mathspeech: Leveraging small lms for accurate conversion in mathematical speech-to-formula. InProceedings of the AAAI Con- ference on Artificial Intelligence, Vol. 39. 24194–24202

  24. [24]

    Huiqiang Jiang et al. 2023. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

  25. [25]

    Huiqiang Jiang et al. 2024. LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

  26. [26]

    Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics3 (2015), 585–597

  27. [27]

    Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, et al. 2024. Numi- namath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository13, 9 (2024), 9

  28. [28]

    Jijie Li, Li Du, Hanyu Zhao, Bo-wen Zhang, Liangdong Wang, Boyan Gao, Guang Liu, and Yonghua Lin. 2025. Infinity instruct: Scaling instruction selection and synthesis to enhance language models.arXiv preprint arXiv:2506.11116(2025)

  29. [29]

    Jindong Li, Yali Fu, Li Fan, Jiahong Liu, Yao Shu, Chengwei Qin, Menglin Yang, Irwin King, and Rex Ying. 2025. Implicit reasoning in large language models: A comprehensive survey.arXiv preprint arXiv:2509.02350(2025)

  30. [30]

    Shufan Li and Aditya Grover. 2025. Predgen: Accelerated inference of large language models through input-time speculation for real-time speech interaction. arXiv preprint arXiv:2506.15556(2025)

  31. [31]

    Tianpeng Li, Jun Liu, Tao Zhang, Yuanbo Fang, Da Pan, Mingrui Wang, Zheng Liang, Zehuan Li, Mingan Lin, Guosheng Dong, et al. 2025. Baichuan-audio: A uni- fied framework for end-to-end speech interaction.arXiv preprint arXiv:2502.17239 (2025)

  32. [32]

    Zeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen, Zhijian Xu, Yingying Cheng, Fan Zhang, and Qiang Xu. 2026. Making Slow Thinking Faster: Com- pressing LLM Chain-of-Thought via Step Entropy. InThe Fourteenth International Conference on Learning Representations

  33. [33]

    Junhong Lin, Xinyue Zeng, Jie Zhu, Song Wang, Julian Shun, Jun Wu, and Dawei Zhou. 2025. Plan and budget: Effective and efficient test-time scaling on large language model reasoning.arXiv preprint arXiv:2505.16122(2025)

  34. [34]

    Yueqian Lin, Zhengmian Hu, Qinsi Wang, Yudong Liu, Hengfan Zhang, Jayaku- mar Subramanian, Nikos Vlassis, Hai Helen Li, and Yiran Chen. 2025. Voice evaluation of reasoning ability: Diagnosing the modality-induced performance gap.arXiv preprint arXiv:2509.26542(2025)

  35. [35]

    Siyuan Liu, Jiahui Xu, Feng Jiang, Kuang Wang, Zefeng Zhao, Chu-Ren Huang, Jinghang Gu, Changqing Yin, and Haizhou Li. 2026. Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue Systems.arXiv preprint arXiv:2602.23266(2026)

  36. [36]

    Ziyang Ma, Zhuo Chen, Yuping Wang, Eng Siong Chng, and Xie Chen. 2025. Audio-cot: Exploring chain-of-thought reasoning in large audio language model. arXiv preprint arXiv:2501.07246(2025)

  37. [37]

    Ziyang Ma, Guanrou Yang, Wenxi Chen, Zhifu Gao, Yexing Du, Xiquan Li, Zhisheng Zheng, Haina Zhu, Jianheng Zhuo, Zheshu Song, et al. 2026. SLAM- LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing.IEEE Journal of Selected Topics in Signal Processing(2026)

  38. [38]

    Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaro- engchai, Soroosh Mariooryad, Ehud Rivlin, RJ Skerry-Ryan, and Michelle Tadmor Ramanovich. 2023. Spoken question answering and speech continuation using spectrogram-powered llm.arXiv preprint arXiv:2305.15255(2023)

  39. [39]

    Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, et al. 2024. Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression. In Findings of the Association for Computational Linguistics: ACL 2024. 963–981. MM ’26, November 10–14, 2026, Rio de Jan...

  40. [40]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems?. InProceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies. 2080–2094

  41. [41]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. InInternational conference on machine learning. PMLR, 28492–28518

  42. [42]

    Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. InProceedings of the 2015 conference on empirical methods in natural language processing. 1743–1752

  43. [43]

    Hejian Sang, Yuanda Xu, Zhengze Zhou, Ran He, Zhipeng Wang, and Jiachen Sun. 2026. On-Policy Self-Distillation for Reasoning Compression.arXiv preprint arXiv:2603.05433(2026)

  44. [44]

    Xuan Shen et al . 2025. Efficient Reasoning with Hidden Thinking.CoRR abs/2501.19201 (2025). arXiv:2501.19201 doi:10.48550/ARXIV.2501.19201

  45. [45]

    Qundong Shi, Jie Zhou, Biyuan Lin, Junbo Cui, Guoyang Zeng, Yixuan Zhou, Ziyang Wang, Xin Liu, Zhen Luo, Yudong Wang, et al. 2026. UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models. arXiv preprint arXiv:2601.01373(2026)

  46. [46]

    Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou, SK Bong, Yashesh Gaur, Jay Mahadeokar, Ozlem Kalinli, and Mike Seltzer. 2025. Can Speech LLMs Think while Listening?arXiv preprint arXiv:2510.07497(2025)

  47. [47]

    Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Na Zou, et al . 2025. Stop overthinking: A survey on efficient reasoning for large language models.arXiv preprint arXiv:2503.16419(2025)

  48. [48]

    Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye, Huadai Liu, Honggang Zhang, Wei Xue, and Yike Guo. 2024. Both ears wide open: Towards language-driven spatial audio generation.arXiv preprint arXiv:2410.10676(2024)

  49. [49]

    Chao-Hong Tan, Qian Chen, Wen Wang, Chong Deng, Qinglin Zhang, Luyao Cheng, Hai Yu, Xin Zhang, Xiang Lv, Tianyu Zhao, et al. 2025. DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representa- tions.arXiv preprint arXiv:2506.09349(2025)

  50. [50]

    Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Ruihua Song, and Jian Luan

  51. [51]

    Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. https://qwenlm. github.io/blog/qwen2.5/

  52. [52]

    Tongyi Fun Team, Qian Chen, Luyao Cheng, Chong Deng, Xiangang Li, Jiaqing Liu, Chao-Hong Tan, Wen Wang, Junhao Xu, Jieping Ye, et al. 2025. Fun-Audio- Chat Technical Report.arXiv preprint arXiv:2512.20156(2025)

  53. [53]

    Chen Wang, Minpeng Liao, Zhongqiang Huang, Junhong Wu, Chengqing Zong, and Jiajun Zhang. 2024. Blsp-emo: Towards empathetic large speech-language models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 19186–19199

  54. [54]

    Chaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu, Yan Lu, Jinyu Li, and Zhizheng Wu. 2026. Closing the Modality Reasoning Gap for Speech Large Language Models.arXiv preprint arXiv:2601.05543(2026)

  55. [55]

    Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, and Hao Fei. 2025. Multimodal chain-of-thought reasoning: A comprehensive survey.arXiv preprint arXiv:2503.12605(2025)

  56. [56]

    Chengwei Wei, Bin Wang, Jung-jae Kim, and Nancy F Chen. 2025. Towards spoken mathematical reasoning: Benchmarking speech-based models over multi- faceted math problems.arXiv preprint arXiv:2505.15000(2025)

  57. [57]

    Donghang Wu, Haoyang Zhang, Jun Chen, Hexin Liu, Eng Siong Chng, Fei Tian, Xuerui Yang, Xiangyu Zhang, Daxin Jiang, Gang Yu, et al . 2025. Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models.arXiv preprint arXiv:2510.09592(2025)

  58. [58]

    Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, and Wenjie Li. 2025. Tokenskip: Controllable chain-of-thought compression in llms. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 3351–3363

  59. [59]

    Jingran Xie, Shun Lei, Yue Yu, Yang Xiang, Hui Wang, Xixin Wu, and Zhiyong Wu

  60. [60]

    Zhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu, Shuicheng Yan, and Chun- yan Miao. 2025. Audio-reasoner: Improving reasoning capability in large audio language models.arXiv preprint arXiv:2503.02318(2025)

  61. [61]

    InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Leveraging chain of thought towards empathetic spoken dialogue without corresponding question-answering data. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5

  62. [62]

    Zhifei Xie and Changqiao Wu. 2024. Mini-omni: Language models can hear, talk while thinking in streaming.arXiv preprint arXiv:2408.16725(2024)

  63. [63]

    Zhifei Xie, Ziyang Ma, Zihang Liu, Kaiyu Pang, Hongyu Li, Jialin Zhang, Yue Liao, Deheng Ye, Chunyan Miao, and Shuicheng Yan. 2025. Mini-omni-reasoner: Token- level thinking-in-speaking in large speech models.arXiv preprint arXiv:2508.15827 (2025)

  64. [64]

    Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. 2025. Chain of draft: Thinking faster by writing less.arXiv preprint arXiv:2502.18600(2025)

  65. [65]

    Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, et al. 2025. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765(2025)

  66. [66]

    Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Mengzhe Chen, Qian Chen, and Lei Xie. 2024. E-chat: Emotion-sensitive spoken dialogue system with large language models. In2024 IEEE 14th International Symposium on Chinese Spoken Language Processing (ISCSLP). IEEE, 586–590

  67. [67]

    Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. 2024. Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing.arXiv preprint arXiv:2406.08464 (2024)

  68. [68]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Cheng- peng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfen...

  69. [69]

    Ruiqi Yan, Xiquan Li, Wenxi Chen, Zhikang Niu, Chen Yang, Ziyang Ma, Kai Yu, and Xie Chen. 2025. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models.arXiv preprint arXiv:2502.17810(2025)

  70. [70]

    Aohan Zeng, Zhengxiao Du, Mingdao Liu, Kedong Wang, Shengmin Jiang, Lei Zhao, Yuxiao Dong, and Jie Tang. 2024. Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot.arXiv preprint arXiv:2412.02612(2024)

  71. [71]

    Wenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Guangzhi Sun, Lu Lu, Yuxuan Wang, and Chao Zhang. 2024. Salmonn-omni: A codec-free llm for full-duplex speech understanding and generation.arXiv preprint arXiv:2411.18138(2024)

  72. [72]

    Dong Zhang, Gang Wang, Jinlong Xue, Kai Fang, Liang Zhao, Rui Ma, Shuhuai Ren, Shuo Liu, Tao Guo, Weiji Zhuang, et al. 2025. MiMo-Audio: Audio Language Models are Few-Shot Learners.arXiv preprint arXiv:2512.23808(2025)

  73. [73]

    Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. InFindings of the Association for Computational Linguistics: EMNLP 2023. 15757–15773

  74. [74]

    Zhen Zhang, Xuehai He, Weixiang Yan, Ao Shen, Chenyang Zhao, Shuohang Wang, Yelong Shen, and Xin Eric Wang. 2025. Soft thinking: Unlocking the reason- ing potential of llms in continuous concept space.arXiv preprint arXiv:2505.15778 (2025)

  75. [75]

    Dong Zhang, Xin Zhang, Jun Zhan, Shimin Li, Yaqian Zhou, and Xipeng Qiu

  76. [76]

    Yufan Zhuang, Liyuan Liu, Chandan Singh, Jingbo Shang, and Jianfeng Gao. 2025. Text generation beyond discrete token sampling.arXiv preprint arXiv:2505.14827 (2025)

  77. [77]

    Wenhao Zou, Yuwei Miao, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, and Jingwen Xu. 2026. LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning.arXiv preprint arXiv:2601.19952(2026). Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoke...

  78. [78]

    Xingjian Zhao, Zhe Xu, Qinyuan Cheng, Zhaoye Fei, Luozhijie Jin, Yang Wang, Hanfu Chen, Yaozhou Jiang, Qinghui Gao, Ke Chen, et al. 2025. MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance.arXiv preprint arXiv:2510.00499(2025)

  79. [2024]

    Speechgpt-gen: Scaling chain-of-information speech generation.arXiv preprint arXiv:2401.13527(2024)

  80. [2025]

    Think silently, think fast: Dynamic latent compression of llm reasoning chains.arXiv preprint arXiv:2505.16552(2025)