Pith. sign in

REVIEW 2 minor 130 references

A Survey of Audio Reasoning in Multimodal Foundation Models

T0 review · 0 major / 2 minor · reviewed 2026-05-21 · grok-4.3

Pith's one-line read Audio reasoning in multimodal foundation models requires a dedicated survey and unified formulation because of its unique continuous and multi-scale characteristics.

desk verdict This is a standard survey that organizes existing audio reasoning work under a new taxonomy but introduces no original methods, data, or results. read the letter →

arxiv 2605.21008 v1 pith:S7FX2JXG submitted 2026-05-20 eess.AS

classification eess.AS
keywords audioreasoningmultimodalfoundationmodelsreasoning-augmentedgenerationaudio-to-textaudio-visualchain-of-thoughtreinforcementlearningspokeninteraction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper seeks to establish a coherent roadmap for audio reasoning by being the first to survey the field specifically. It distinguishes direct predictive modeling from reasoning-augmented generation to better organize how models align audio signals with language semantics. A reader would care if this leads to more reliable systems that can infer from speech, environmental sounds, and combined audio-visual inputs without losing fine details. The work reviews foundations, organizes advances in four categories, and covers methods like prompting and training techniques. It also points out obstacles such as data scarcity and the need to balance reasoning with speed.

What carries the argument

A unified formulation that distinguishes direct predictive modeling from reasoning-augmented generation to handle the alignment of continuous acoustic signals with discrete language model semantics while preserving fine-grained information.

What would settle it

An experiment showing that general multimodal reasoning techniques without audio-specific adaptations achieve equivalent performance on audio tasks would challenge the premise for a dedicated survey.

Watch

Extended reading notes

Core claim

The authors present the first dedicated survey of audio reasoning in multimodal foundation models. They introduce a unified formulation to separate direct predictive modeling from reasoning-augmented generation, review the architectural and training foundations, and systematically organize recent advances across Audio-to-Text, Audio-to-Speech, Audio-Visual Reasoning, and Agentic Audio Reasoning. The survey further examines emerging paradigms including Chain-of-Thought prompting, supervised fine-tuning, reinforcement learning, and latency-aware spoken interaction, along with evaluation practices and open challenges.

Load-bearing premise

The challenges in audio reasoning are fundamentally distinct from those in text and vision, necessitating a separate survey and a new unified formulation.

Editorial extensions

If this is right

  • Advances in Audio-to-Text and Audio-to-Speech can be more systematically compared and improved.
  • Agentic Audio Reasoning can support interactive spoken agents that perform step-by-step inference.
  • Methods like reinforcement learning can help overcome shortcut learning in audio tasks.
  • Latency-aware designs enable practical real-time audio reasoning applications.
  • Evaluation practices can evolve to better test for modality hallucination and grounding.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • This categorization could help in designing experiments that test reasoning depth versus prediction accuracy in audio models.
  • Connections to visual reasoning suggest potential for unified multi-modal reasoning frameworks beyond audio alone.
  • Addressing the listed obstacles might lead to foundation models that handle real-world audio interactions more robustly.
  • One could test the formulation by applying it to emerging audio datasets to see if it reveals new patterns in progress.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The paper claims to deliver the first dedicated survey of audio reasoning in multimodal foundation models. It provides a unified formulation distinguishing direct predictive modeling from reasoning-augmented generation, reviews architectural and training foundations, and organizes advances in Audio-to-Text, Audio-to-Speech, Audio-Visual Reasoning, and Agentic Audio Reasoning. The survey also covers paradigms like Chain-of-Thought prompting, supervised fine-tuning, reinforcement learning, latency-aware interaction, evaluation practices, challenges, and future directions.

Significance. If the claims hold, this survey would be significant for the field by establishing a coherent framework and roadmap for audio reasoning, which is currently limited compared to text and vision. The explicit identification of distinct audio challenges and obstacles like data scarcity and modality hallucination provides a useful structure for future work. As a survey without new quantitative claims, its value lies in synthesis and organization of existing literature.

minor comments (2)
  1. [Abstract] Abstract: The premise that audio poses fundamentally distinct challenges from text and vision is stated to motivate the scope; a brief explicit contrast with vision-language reasoning surveys would strengthen the justification for a dedicated audio survey.
  2. The unified formulation is introduced in the abstract but its concrete mathematical or conceptual details are not visible in the provided high-level description; ensuring the formulation is presented with clear notation and examples in the main text would improve accessibility.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive assessment and recommendation for minor revision. The summary accurately reflects the paper's contributions in providing a unified formulation of audio reasoning and organizing advances across Audio-to-Text, Audio-to-Speech, Audio-Visual, and Agentic paradigms, while highlighting key challenges such as data scarcity and modality hallucination.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

This is a survey paper whose central contribution is a review and taxonomy of existing literature on audio reasoning. It states a motivation based on modality differences and offers a unified formulation to organize prior work, but introduces no new quantitative predictions, fitted parameters, or formal derivations that could reduce to its own inputs. All load-bearing content consists of citations to external studies and internal consistency of the proposed categories, with no self-referential loops or self-citation chains that substitute for independent evidence. The paper is therefore self-contained against external benchmarks.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

This is a literature survey with no new mathematical derivations, fitted parameters, or postulated entities; it relies on standard domain assumptions from multimodal AI research.

assumptions (1)
  • domain assumption Audio poses challenges distinct from text and vision because it is continuous, temporally dense, and contains linguistic, paralinguistic, and environmental information at multiple time scales.
    Invoked in the abstract to motivate the need for specialized audio reasoning models.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey of Audio Reasoning in Multimodal Foundation Models." pith.science (2026). https://pith.science/paper/S7FX2JXG

@misc{pith2026260521008,
  author       = {Pith},
  title        = {Pith review of: A Survey of Audio Reasoning in Multimodal Foundation Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S7FX2JXG}},
  note         = {Machine review of arXiv:2605.21008}
}
read the original abstract

Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from those of text and vision. It is continuous, temporally dense, and contains linguistic, paralinguistic, and environmental information at multiple time scales. As a result, audio reasoning models must align acoustic signals with the discrete semantic space of large language models, while still preserving fine-grained information needed for reliable inference. Progress is also limited by three major obstacles: the scarcity of genuinely audio-grounded reasoning data, shortcut learning and modality hallucination, and the tension between reasoning depth and real-time latency in spoken interaction. In this paper, we present the first dedicated survey of audio reasoning. We provide a unified formulation that distinguishes direct predictive modeling from reasoning-augmented generation, review the architectural and training foundations of audio reasoning models, and systematically organize recent advances in Audio-to-Text, Audio-to-Speech, Audio-Visual Reasoning and Agentic Audio Reasoning. We further examine emerging paradigms such as Chain-of-Thought prompting, supervised fine-tuning, reinforcement learning, and latency-aware spoken interaction, and discuss evaluation practices, open challenges, and future directions. Our goal is to offer a coherent roadmap for developing robust, efficient, and natively grounded audio reasoning systems.

Figures

Figures reproduced from arXiv: 2605.21008 by the authors.

Figure 1
Figure 1. Timeline of representative audio reasoning models. Models are organized chronologically and grouped by major paradigms, including Audio-to-Text, Audio-to-Speech, Audio-Visual, and agentic audio reasoning. from direct generation to structured problem solving. This tax￾onomy clarifies the scope of audio reasoning and highlights the field’s current fragmentation across formulation, architecture, training, interaction, … view at source ↗
Figure 2
Figure 2. A compact taxonomy of audio reasoning. We organize the literature into four paradigms: Audio-to-Text reasoning, Audio-to-Speech reasoning, Audio-Visual reasoning, and Agentic Audio Reasoning. Representative meth￾ods and design patterns are discussed in the corresponding sections. under a common probabilistic view. For clarity, Table I sum￾marizes the main symbols used throughout the paper. A. General Formulation of … view at source ↗
Figure 3
Figure 3. Overview of major audio reasoning paradigms. The figure summarizes four paradigms covered in this survey: Audio-to-Text, Audio-to-Speech, Audio-Visual, and Agentic Audio Reasoning. It contrasts text-output reasoning, cross-modal audio-visual grounding, sequential and real-time speech-output reasoning, and agentic workflows based on predefined pipelines or dynamic tool calling. sufficiently complete to trigger early … view at source ↗

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Reference graph

Works this paper leans on

130 extracted references · 130 canonical work pages

  1. [1]

    Chain-of-thought prompting elicits reasoning in large 17 language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhouet al., “Chain-of-thought prompting elicits reasoning in large 17 language models,”Advances in neural information processing systems, vol. 35, pp. 24 824–24 837, 2022

  2. [2]

    From System 1 to System 2: A Survey of Reasoning Large Language Models

    Z.-Z. e. a. Li, “From system 1 to system 2: A survey of reasoning large language models,”arXiv preprint arXiv:2502.17419, 2025

  3. [3]

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

    D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Biet al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,”arXiv preprint arXiv:2501.12948, 2025

  4. [4]

    Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

    J. D. et al., “Beyond benchmarks: Matharena as an evaluation platform for mathematics with llms,” 2026. [Online]. Available: https://arxiv.org/abs/2605.00674

  5. [5]

    Let’s verify step by step,

    H. e. a. Lightman, “Let’s verify step by step,” inInternational Confer- ence on Learning Representations, vol. 2024, 2024, pp. 39 578–39 601

  6. [6]

    Scaling llm test-time compute optimally can be more effective than scaling parameters for reasoning,

    C. V . Snell, J. Lee, K. Xu, and A. Kumar, “Scaling llm test-time compute optimally can be more effective than scaling parameters for reasoning,” inThe Thirteenth International Conference on Learning Representations, 2025

  7. [7]

    Inference scaling laws: An empirical analysis of compute-optimal inference for llm problem- solving,

    Y . Wu, Z. Sun, S. Li, S. Welleck, and Y . Yang, “Inference scaling laws: An empirical analysis of compute-optimal inference for llm problem- solving,” inThe Thirteenth International Conference on Learning Representations, 2025

  8. [8]

    Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought rea- soning,

    H. e. a. Shao, “Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought rea- soning,”Advances in Neural Information Processing Systems, vol. 37, pp. 8612–8642, 2024

Show all 130 references
  1. [9]

    Compositional chain- of-thought prompting for large multimodal models,

    C. Mitra, B. Huang, T. Darrell, and R. Herzig, “Compositional chain- of-thought prompting for large multimodal models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2024, pp. 14 420–14 431

  2. [10]

    On the landscape of spoken language models: A comprehensive survey,

    S. Arora, K.-W. Chang, C.-M. Chien, Y . Peng, H. Wu, Y . Adi, E. Dupoux, H.-Y . Lee, K. Livescu, and S. Watanabe, “On the landscape of spoken language models: A comprehensive survey,”arXiv preprint arXiv:2504.08528, 2025

  3. [11]

    Mmau: A massive multi-task audio understanding and reasoning benchmark,

    S. e. a. Sakshi, “Mmau: A massive multi-task audio understanding and reasoning benchmark,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 84 929–84 964

  4. [12]

    Sd-eval: A benchmark dataset for spoken dialogue under- standing beyond words,

    J. e. a. Ao, “Sd-eval: A benchmark dataset for spoken dialogue under- standing beyond words,”Advances in Neural Information Processing Systems, vol. 37, pp. 56 898–56 918, 2024

  5. [13]

    Recent advances in discrete speech tokens: A review,

    Y . e. a. Guo, “Recent advances in discrete speech tokens: A review,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  6. [14]

    What enables human language? a biocultural frame- work,

    I. e. a. Arnon, “What enables human language? a biocultural frame- work,”Science, vol. 390, no. 6775, p. eadq8303, 2025

  7. [15]

    Representation of internal speech by single neurons in human supramarginal gyrus,

    S. K. e. a. Wandelt, “Representation of internal speech by single neurons in human supramarginal gyrus,”Nature human behaviour, vol. 8, no. 6, pp. 1136–1149, 2024

  8. [17]

    OmniFlatten: An end-to-end GPT model for seamless voice conversation,

    Q. e. a. Zhang, “OmniFlatten: An end-to-end GPT model for seamless voice conversation,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: ...

  9. [18]

    To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning,

    Z. R. S. et al., “To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning,” inThe Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=w6nlcS8Kkn

  10. [19]

    Benchmarking open-ended audio dialogue understanding for large audio-language models,

    K. e. a. Gao, “Benchmarking open-ended audio dialogue understanding for large audio-language models,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vie...

  11. [20]

    Recent advances in speech language models: A survey,

    W. Cui, D. Yu, X. Jiao, Z. Meng, G. Zhang, Q. Wang, S. Y . Guo, and I. King, “Recent advances in speech language models: A survey,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 13 943– 13 970

  12. [21]

    Sparks of large au- dio models: A survey and outlook,

    S. Latif, M. Shoukat, F. Shamshad, M. Usama, Y . Ren, H. Cuayáhuitl, W. Wang, X. Zhang, R. Togneri, E. Cambriaet al., “Sparks of large au- dio models: A survey and outlook,”arXiv preprint arXiv:2308.12792, 2023

  13. [22]

    Audio-language models for audio-centric tasks: A survey,

    Y . Su, J. Bai, Q. Xu, K. Xu, and Y . Dou, “Audio-language models for audio-centric tasks: A survey,”arXiv preprint arXiv:2501.15177, 2025

  14. [23]

    A survey on speech large language models for understanding,

    J. Peng, Y . Wang, B. Li, Y . Guo, H. Wang, Y . Fang, Y . Xi, H. Li, X. Li, K. Zhanget al., “A survey on speech large language models for understanding,”IEEE Journal of Selected Topics in Signal Processing, 2025

  15. [24]

    Towards general auditory intelligence: Large multimodal models for machine listening and speaking,

    S. Wang, Z. Jin, C. Tang, Q. Li, B. Li, C. Chen, Y . Hu, W. Yu, Y . Li, J. Zhuanget al., “Towards general auditory intelligence: Large multimodal models for machine listening and speaking,”arXiv preprint arXiv:2511.01299, 2025

  16. [25]

    Towards holistic evaluation of large audio-language models: A comprehensive survey,

    C.-K. Yang, N. S. Ho, and H.-y. Lee, “Towards holistic evaluation of large audio-language models: A comprehensive survey,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 10 155–10 181

  17. [26]

    Multimodal chain-of-thought reasoning: A comprehensive survey,

    Y . Wang, S. Wu, Y . Zhang, S. Yan, Z. Liu, J. Luo, and H. Fei, “Multimodal chain-of-thought reasoning: A comprehensive survey,” arXiv preprint arXiv:2503.12605, 2025

  18. [27]

    Robust speech recognition via large-scale weak super- vision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” inInternational Conference on Machine Learning (ICML). PMLR, 2023, pp. 28 492–28 518

  19. [28]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021

  20. [29]

    Beats: Audio pre-training with acoustic tokenizers,

    S. Chen, Y . Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “Beats: Audio pre-training with acoustic tokenizers,” inInternational Conference on Machine Learning (ICML). PMLR, 2023, pp. 5178–5193

  21. [30]

    Ast: Audio spectrogram trans- former,

    Y . Gong, Y .-A. Chung, and J. Glass, “Ast: Audio spectrogram trans- former,” inProc. Interspeech 2021, 2021, pp. 571–575

  22. [31]

    Robust speech recognition via large-scale weak supervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” 2022. [Online]. Available: https://arxiv.org/abs/2212. 04356

  23. [32]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosaleet al., “Llama 2: Open foundation and fine-tuned chat models,”arXiv preprint arXiv:2307.09288, 2023

  24. [33]

    The llama 3 herd of models,

    A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fanet al., “The llama 3 herd of models,”arXiv preprint arXiv:2407.21783, 2024

  25. [34]

    Qwen2 technical report,

    A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huanget al., “Qwen2 technical report,”arXiv preprint arXiv:2407.10671, 2024

  26. [35]

    Vi- cuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,

    W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vi- cuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,” https://lmsys.org/blog/2023-03-30-vicuna/, 2023, accessed: 2023-03-30

  27. [36]

    Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot,

    A. Zeng, Z. Du, M. Liu, K. Wang, S. Jiang, L. Zhao, Y . Dong, and J. Tang, “Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot,”arXiv preprint arXiv:2412.02612, 2024

  28. [37]

    LLaMA- Omni: Seamless speech interaction with large language models,

    Q. Fang, S. Niu, R. Zhou, Z. Lin, M. Chen, and Y . Feng, “LLaMA- Omni: Seamless speech interaction with large language models,”arXiv preprint arXiv:2409.06666, 2024

  29. [38]

    Speech gpt: Empowering large language models with intrinsic cross- modal conversational abilities,

    D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y . Zhou, and X. Qiu, “Speech gpt: Empowering large language models with intrinsic cross- modal conversational abilities,” inFindings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp. 15 757–15 773

  30. [39]

    Moshi: a speech-text foundation model for real- time dialogue,

    A. Défossezet al., “Moshi: a speech-text foundation model for real- time dialogue,”arXiv preprint arXiv:2410.00080, 2024

  31. [40]

    Blsp: Bootstrapping language-speech pre-training via behavior alignment of continuation writing,

    C. Wang, M. Liao, Z. Huang, J. Lu, J. Wu, Y . Liu, C. Zong, and J. Zhang, “Blsp: Bootstrapping language-speech pre-training via behavior alignment of continuation writing,” 2024. [Online]. Available: https://arxiv.org/abs/2309.00916

  32. [41]

    Salmonn: Towards generic hearing abilities for large language models,

    C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “Salmonn: Towards generic hearing abilities for large language models,” inThe Twelfth International Conference on Learning Representations (ICLR), 2024

  33. [42]

    Proximal policy optimization algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017. [Online]. Available: https://arxiv.org/abs/1707.06347

  34. [43]

    Deepseekmath: Pushing 18 the limits of mathematical reasoning in open language models,

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . K. Li, Y . Wu, and D. Guo, “Deepseekmath: Pushing 18 the limits of mathematical reasoning in open language models,” 2024. [Online]. Available: https://arxiv.org/abs/2402.03300

  35. [44]

    Audio-cot: Exploring chain-of-thought reasoning in large audio language model,

    Z. Ma, Z. Chen, Y . Wang, E. S. Chng, and X. Chen, “Audio-cot: Exploring chain-of-thought reasoning in large audio language model,” arXiv preprint arXiv:2501.07246, 2025

  36. [45]

    Sar-lm: Symbolic audio reasoning with large language models,

    T. Taheri, Y . Ma, and E. Benetos, “Sar-lm: Symbolic audio reasoning with large language models,”arXiv preprint arXiv:2511.06483, 2025

  37. [46]

    Audio-reasoner: Improving reasoning capability in large audio language models,

    Z. Xie, M. Lin, Z. Liu, P. Wu, S. Yan, and C. Miao, “Audio-reasoner: Improving reasoning capability in large audio language models,”arXiv preprint arXiv:2503.02318, 2025

  38. [47]

    Audio flamingo sound-cot technical report: Improving chain-of-thought reasoning in sound understanding,

    Z. Kong, A. Goel, J. F. Santos, S. Ghosh, R. Valle, W. Ping, and B. Catanzaro, “Audio flamingo sound-cot technical report: Improving chain-of-thought reasoning in sound understanding,”arXiv preprint arXiv:2508.11818, 2025

  39. [48]

    Audio- cogito: Towards deep audio reasoning in large audio language models,

    L. Li, H. Chen, Z. Li, Q. Hu, J. Kang, J. Li, L. Xie, and Y . Li, “Audio- cogito: Towards deep audio reasoning in large audio language models,” arXiv preprint arXiv:2604.12527, 2026

  40. [49]

    Reinforcement learning outperforms supervised fine-tuning: A case study on audio question answering,

    G. Li, J. Liu, H. Dinkel, Y . Niu, J. Zhang, and J. Luan, “Reinforcement learning outperforms supervised fine-tuning: A case study on audio question answering,”arXiv preprint arXiv:2503.11197, 2025

  41. [50]

    Omni-r1: Do you really need audio to fine-tune your audio llm?

    A. Rouditchenko, S. Bhati, E. Araujo, S. Thomas, H. Kuehne, R. Feris, and J. Glass, “Omni-r1: Do you really need audio to fine-tune your audio llm?”arXiv preprint arXiv:2505.09439, 2025

  42. [52]

    Data- balanced curriculum learning for audio question answering,

    G. Wijngaard, E. Formisano, M. Esposito, and M. Dumontier, “Data- balanced curriculum learning for audio question answering,”arXiv preprint arXiv:2507.06815, 2025

  43. [53]

    Phi-4 technical report,

    M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kauffmannet al., “Phi-4 technical report,”arXiv preprint arXiv:2412.08905, 2024

  44. [54]

    Sari: Structured audio reasoning via curriculum-guided reinforcement learning,

    C. Wen, T. Guo, S. Zhao, W. Zou, and X. Li, “Sari: Structured audio reasoning via curriculum-guided reinforcement learning,”arXiv preprint arXiv:2504.15900, 2025

  45. [55]

    Omni- autothink: Adaptive multimodal reasoning via reinforcement learning,

    D. Yang, S. Liu, D. Wang, Y . Wang, G. Wan, and H. Meng, “Omni- autothink: Adaptive multimodal reasoning via reinforcement learning,” arXiv preprint arXiv:2512.03783, 2025

  46. [56]

    Omni-clst: Error-aware curriculum learning with guided selec- tive chain-of-thought for audio question answering,

    J. Zhao, H. Su, L. Fan, Z. Luo, H. Wang, H. Sun, and Y . Qin, “Omni-clst: Error-aware curriculum learning with guided selec- tive chain-of-thought for audio question answering,”arXiv preprint arXiv:2509.12275, 2025

  47. [57]

    Think smart, not hard: Difficulty adaptive reasoning for large audio language models,

    Z. Sheng, S. Zhou, C. Gong, and Z. Li, “Think smart, not hard: Difficulty adaptive reasoning for large audio language models,”arXiv preprint arXiv:2509.21960, 2025

  48. [58]

    Aud- semthinker: Enhancing audio-language models through reasoning over semantics of sound,

    G. Wijngaard, E. Formisano, M. Esposito, and M. Dumontier, “Aud- semthinker: Enhancing audio-language models through reasoning over semantics of sound,”arXiv preprint arXiv:2505.14142, 2025

  49. [59]

    Measuring audio’s impact on correctness: Audio- contribution-aware post-training of large audio language models,

    H. He, X. Du, R. Sun, Z. Dai, Y . Xiao, M. Yang, J. Zhou, X. Li, Z. Liu, Z. Lianget al., “Measuring audio’s impact on correctness: Audio- contribution-aware post-training of large audio language models,”arXiv preprint arXiv:2509.21060, 2025

  50. [60]

    Step-audio-r1 technical report,

    F. Tian, X. T. Zhang, Y . Zhang, H. Zhang, Y . Li, D. Liu, Y . Deng, D. Wu, J. Chen, L. Zhaoet al., “Step-audio-r1 technical report,”arXiv preprint arXiv:2511.15848, 2025

  51. [61]

    Step-audio 2 technical report,

    B. Wu, C. Yan, C. Hu, C. Yi, C. Feng, F. Tian, F. Shen, G. Yu, H. Zhang, J. Liet al., “Step-audio 2 technical report,”arXiv preprint arXiv:2507.16632, 2025

  52. [62]

    Audio-thinker: Guiding audio language model when and how to think via reinforcement learning,

    S. Wu, C. Li, W. Wang, H. Zhang, H. Wang, M. Yu, and D. Yu, “Audio-thinker: Guiding audio language model when and how to think via reinforcement learning,”arXiv preprint arXiv:2508.08039, 2025

  53. [63]

    Audio-deepthinker: Progressive reasoning-aware reinforcement learning for high-quality chain-of-thought emergence in audio language models,

    X. He, C. Li, J. Wang, Y . Rong, T. Xie, W. Wang, L. Liu, and D. Yu, “Audio-deepthinker: Progressive reasoning-aware reinforcement learning for high-quality chain-of-thought emergence in audio language models,”arXiv preprint arXiv:2604.18187, 2026

  54. [64]

    Incentivizing consistent, effective and scalable reasoning capability in audio llms via reasoning process rewards,

    J. Fan, R. Ren, J. Li, R. Pandey, P. G. Shivakumar, I. Bulyko, A. Gandhe, G. Liu, and Y . Gu, “Incentivizing consistent, effective and scalable reasoning capability in audio llms via reasoning process rewards,”arXiv preprint arXiv:2510.20867, 2025

  55. [65]

    Soundmind: Rl-incentivized logic reasoning for audio-language models,

    X. Diao, C. Zhang, K. Kong, W. Wu, C. Ma, Z. Ouyang, P. Qing, S. V osoughi, and J. Gui, “Soundmind: Rl-incentivized logic reasoning for audio-language models,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 528– 540

  56. [66]

    Beyond single-audio: Advancing multi-audio processing in audio large language models,

    Y . Chen, X. Yue, X. Gao, C. Zhang, L. F. D’Haro, R. T. Tan, and H. Li, “Beyond single-audio: Advancing multi-audio processing in audio large language models,” inFindings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 10 917–10 930

  57. [67]

    Polyaudio: Advancing multi-audio analysis & reasoning in large audio language models,

    S. Kumar, S. Ghosh, Y . Lin, Y . Chen, R. Duraiswami, and D. Manocha, “Polyaudio: Advancing multi-audio analysis & reasoning in large audio language models,” 2025

  58. [68]

    Emotion- thinker: Prosody-aware reinforcement learning for explainable speech emotion reasoning,

    D. Wang, S. Liu, T. Zhang, Y . Chen, J. Li, and H. Meng, “Emotion- thinker: Prosody-aware reinforcement learning for explainable speech emotion reasoning,”arXiv preprint arXiv:2601.15668, 2026

  59. [69]

    Qwen2.5-omni technical report,

    J. X. et al., “Qwen2.5-omni technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2503.20215

  60. [70]

    Kimi-audio technical report,

    D. Ding, Z. Ju, Y . Leng, S. Liu, T. Liu, Z. Shang, K. Shen, W. Song, X. Tan, H. Tanget al., “Kimi-audio technical report,”arXiv preprint arXiv:2504.18425, 2025

  61. [71]

    Mini-omni: Language models can hear, talk while thinking in streaming,

    Z. Xie and C. Wu, “Mini-omni: Language models can hear, talk while thinking in streaming,”arXiv preprint arXiv:2408.16725, 2024

  62. [72]

    Mini-omni2: Towards open-source gpt-4o model with vision, speech and duplex,

    ——, “Mini-omni2: Towards open-source gpt-4o model with vision, speech and duplex,”arXiv preprint arXiv:2410.11190, 2024

  63. [73]

    SLAM-omni: Timbre-controllable voice interaction system with single-stage training,

    W. e. a. Chen, “SLAM-omni: Timbre-controllable voice interaction system with single-stage training,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: Association for Computational L...

  64. [74]

    Minimizing modality gap from the input side: Your speech llm can be a prosody-aware text llm,

    W. Cui, X.-H. Li, D. Tan, Q. Zheng, and I. King, “Minimizing modality gap from the input side: Your speech llm can be a prosody-aware text llm,”arXiv preprint arXiv:2605.05927, 2026. [Online]. Available: https://arxiv.org/abs/2605.05927

  65. [76]

    Qwen3. 5-omni technical report,

    Q. Team, “Qwen3. 5-omni technical report,”arXiv preprint arXiv:2604.15804, 2026

  66. [77]

    Mimo-audio: Audio language models are few-shot learners,

    L.-C.-T. Xiaomi, “Mimo-audio: Audio language models are few-shot learners,” 2025. [Online]. Available: https://github.com/XiaomiMiMo/ MiMo-Audio

  67. [78]

    Opens2s: Advancing fully open-source end-to-end empathetic large speech language model,

    C. Wang, T. Peng, W. Yang, Y . Bai, G. Wang, J. Lin, L. Jia, L. Wu, J. Wang, C. Zonget al., “Opens2s: Advancing fully open-source end-to-end empathetic large speech language model,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Syste...

  68. [79]

    Shanks: Simultaneous hear- ing and thinking for spoken language models,

    C.-H. Chiang, X. Wang, L. Li, C.-C. Lin, K. Lin, S. Liu, Z. Wang, Z. Yang, H.-y. Lee, and L. Wang, “Shanks: Simultaneous hear- ing and thinking for spoken language models,”arXiv preprint arXiv:2510.06917, 2025

  69. [80]

    Can speech LLMs think while listening?

    Y .-J. Shih, D. Raj, C. Wu, W. Zhou, S. Bong, Y . Gaur, J. Mahadeokar, O. Kalinli, and M. Seltzer, “Can speech LLMs think while listening?” inThe Fourteenth International Conference on Learning Representations, 2026. [Online]. Available: https: //openreview.net/forum?id=dFVenZdVbX

  70. [81]

    Chronological thinking in full-duplex spoken dialogue language models,

    D. Wu, H. Zhang, C. Chen, T. Zhang, F. Tian, X. Yang, G. Yu, H. Liu, N. Hou, Y . Huet al., “Chronological thinking in full-duplex spoken dialogue language models,”arXiv preprint arXiv:2510.05150, 2025

  71. [82]

    The silent thought: Modeling internal cognition in full- duplex spoken dialogue models via latent reasoning,

    D. Wu, T. Zhang, Y . Li, H. Liu, C. Chen, E. S. Chng, and Y . Bengio, “The silent thought: Modeling internal cognition in full- duplex spoken dialogue models via latent reasoning,”arXiv preprint arXiv:2603.17837, 2026

  72. [83]

    STITCH: Simultaneous thinking and talking with chunked reasoning for spoken language models,

    C.-H. C. et al., “STITCH: Simultaneous thinking and talking with chunked reasoning for spoken language models,” inThe Fourteenth International Conference on Learning Representations, 2026. [Online]. Available: https://openreview.net/forum?id=5Z1eMhCeTb

  73. [84]

    Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models,

    Z. Xie, Z. Ma, Z. Liu, K. Pang, H. Li, J. Zhang, Y . Liao, D. Ye, C. Miao, and S. Yan, “Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models,”arXiv preprint arXiv:2508.15827, 2025

  74. [85]

    Mind-paced speaking: A dual-brain approach to real-time reasoning in spoken language models,

    D. Wu, H. Zhang, J. Chen, H. Liu, E. S. Chng, F. Tian, X. Yang, X. Zhang, D. Jiang, G. Yuet al., “Mind-paced speaking: A dual-brain approach to real-time reasoning in spoken language models,”arXiv preprint arXiv:2510.09592, 2025

  75. [86]

    Reflecting twice before speaking with empathy: Self-reflective alternating inference for empathy-aware end-to-end spoken dialogue,

    Y . Jia, P. Liu, H. Sun, J. Zhou, X. Cheng, C. Liu, K. Zeng, X. Cai, and Y . Qin, “Reflecting twice before speaking with empathy: Self-reflective alternating inference for empathy-aware end-to-end spoken dialogue,” arXiv preprint arXiv:2601.18281, 2026

  76. [87]

    The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning,

    S. Kim, S. J. Joo, D. Kim, J. Jang, S. Ye, J. Shin, and M. Seo, “The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning,” inThe 2023 Conference on Empirical Methods in Natural Language Processing, 2023. [Online]. Avail...

  77. [88]

    Kame: Tandem architec- ture for enhancing knowledge in real-time speech-to-speech conversa- tional ai,

    S. Kuroki, Y . Kubo, T. Akiba, and Y . Tang, “Kame: Tandem architec- ture for enhancing knowledge in real-time speech-to-speech conversa- tional ai,”arXiv preprint arXiv:2510.02327, 2025

  78. [89]

    Moshirag: Asynchronous knowledge retrieval for full-duplex speech language models,

    C.-M. Chien, M. Orsini, E. Kharitonov, N. Zeghidour, K. Livescu, and A. Défossez, “Moshirag: Asynchronous knowledge retrieval for full-duplex speech language models,” 2026. [Online]. Available: https://arxiv.org/abs/2604.12928

  79. [90]

    Qwen3-omni technical report,

    J. X. et al., “Qwen3-omni technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2509.17765

  80. [91]

    Vita: Towards open-source interactive omni multimodal llm,

    C. Fu, H. Lin, Z. Long, Y . Shen, M. Zhao, Y . Zhang, X. Wang, D. Yin, L. Ma, X. Zhenget al., “Vita: Towards open-source interactive omni multimodal llm,”arXiv preprint arXiv:2408.05211, 2024

  81. [92]

    Veomni: Scaling any modality model training with model-centric distributed recipe zoo,

    Q. Ma, Y . Zheng, Z. Shi, Z. Zhao, B. Jia, Z. Huang, Z. Lin, Y . Li, J. Yang, Y . Peng, Z. Zhang, and X. Liu, “Veomni: Scaling any modality model training with model-centric distributed recipe zoo,”

  82. [93]

    Available: https://arxiv.org/abs/2508.02317

    [Online]. Available: https://arxiv.org/abs/2508.02317

  83. [94]

    Baichuan-omni technical report,

    Y . L. et al., “Baichuan-omni technical report,” 2024. [Online]. Available: https://arxiv.org/abs/2410.08565

  84. [95]

    Lyra: An efficient and speech-centric framework for omni-cognition,

    Z. Zhong, C. Wang, Y . Liu, S. Yang, L. Tang, Y . Zhang, J. Li, T. Qu, Y . Li, Y . Chen, S. Yu, S. Wu, E. Lo, S. Liu, and J. Jia, “Lyra: An efficient and speech-centric framework for omni-cognition,” 2024. [Online]. Available: https://arxiv.org/abs/2412.09501

  85. [96]

    Direct preference optimization: Your language model is secretly a reward model,

    R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” 2024. [Online]. Available: https://arxiv.org/abs/2305.18290

  86. [97]

    Baichuan-omni-1.5 technical report,

    Y . L. et al., “Baichuan-omni-1.5 technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2501.15368

  87. [98]

    Longcat-flash-omni technical report,

    M. L. T. et al., “Longcat-flash-omni technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2511.00279

  88. [99]

    Group sequence policy optimization,

    C. Zheng, S. Liu, M. Li, X.-H. Chen, B. Yu, C. Gao, K. Dang, Y . Liu, R. Men, A. Yang, J. Zhou, and J. Lin, “Group sequence policy optimization,” 2025. [Online]. Available: https: //arxiv.org/abs/2507.18071

  89. [101]

    Worldsense: Evaluating real-world omnimodal understanding for multimodal llms,

    J. Hong, S. Yan, J. Caiet al., “Worldsense: Evaluating real-world omnimodal understanding for multimodal llms,” inICLR 2026, 2026

  90. [102]

    Audio-centric video understanding benchmark without text shortcut,

    Y . Yang, J. Zhuang, G. Sunet al., “Audio-centric video understanding benchmark without text shortcut,” inEMNLP 2025, 2025

  91. [103]

    See, hear, and understand: Benchmarking audiovisual human speech understanding in multimodal large language models,

    L. T. P. Nguyen and e. a. Yu, “See, hear, and understand: Benchmarking audiovisual human speech understanding in multimodal large language models,”arXiv preprint arXiv:2512.02231, 2025

  92. [104]

    Omni-captioner: Data pipeline, mod- els, and benchmark for omni detailed perception,

    Z. Ma, R. Xu, Z. Xinget al., “Omni-captioner: Data pipeline, mod- els, and benchmark for omni detailed perception,”arXiv preprint arXiv:2510.12720, 2025

  93. [105]

    Speech-hands: A self-reflection voice agentic approach to speech recognition and audio reasoning with omni perception,

    Z. e. a. Wan, “Speech-hands: A self-reflection voice agentic approach to speech recognition and audio reasoning with omni perception,”arXiv preprint arXiv:2601.09413, 2026

  94. [106]

    Audiogenie-reasoner: A training- free multi-agent framework for coarse-to-fine audio deep reasoning,

    Y . Rong, C. Li, D. Yu, and L. Liu, “Audiogenie-reasoner: A training- free multi-agent framework for coarse-to-fine audio deep reasoning,” arXiv preprint arXiv:2509.16971, 2025

  95. [107]

    Lts-voiceagent: A listen-think-speak framework for efficient streaming voice interaction via semantic triggering and incremental reasoning,

    W. Zou, Y . Miao, Z. Ma, J. Xu, J. Gao, J. Hao, R. He, and J. Xu, “Lts-voiceagent: A listen-think-speak framework for efficient streaming voice interaction via semantic triggering and incremental reasoning,” arXiv preprint arXiv:2601.19952, 2026

  96. [108]

    Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities,

    Z. Zhou, R. Wang, and Z. Wu, “Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities,” 2025. [Online]. Available: https://arxiv.org/abs/2505.17862

  97. [109]

    Aura: Agent for understanding, reasoning, and auto- mated tool use in voice-driven tasks,

    L. M. Maben, G. G. Lakshmy, S. Radhakrishnan, S. Arora, and S. Watanabe, “Aura: Agent for understanding, reasoning, and auto- mated tool use in voice-driven tasks,”arXiv preprint arXiv:2506.23049, 2025

  98. [110]

    Audiotoola- gent: An agentic framework for audio-language models,

    G. Wijngaard, E. Formisano, M. Dumontier, and J. Jitsev, “Audiotoola- gent: An agentic framework for audio-language models,”arXiv preprint arXiv:2510.02995, 2025

  99. [111]

    Autagent: A reinforcement learning framework for tool-augmented audio reasoning,

    S. Tong, X. Li, Y . Wang, B. Bi, Y . Cai, S. Liu, Y . He, and C. Hao, “Autagent: A reinforcement learning framework for tool-augmented audio reasoning,”arXiv preprint arXiv:2602.13685, 2026

  100. [112]

    Stream rag: Instant and accurate spoken dialogue systems with streaming tool usage,

    S. Arora, H. Khan, K. Sun, X. L. Dong, S. Choudhary, S. Moon, X. Zhang, A. Sagar, S. T. Appini, K. Patnaiket al., “Stream rag: Instant and accurate spoken dialogue systems with streaming tool usage,”arXiv preprint arXiv:2510.02044, 2025

  101. [113]

    V oxmind: An end-to-end agentic spoken dialogue system,

    T. e. a. Liang, “V oxmind: An end-to-end agentic spoken dialogue system,”arXiv preprint arXiv:2604.15710, 2026

  102. [114]

    V oicebench: Benchmarking llm-based voice assistants,

    Y . Chen, X. Yue, C. Zhang, X. Gao, R. T. Tan, and H. Li, “V oicebench: Benchmarking llm-based voice assistants,”arXiv preprint arXiv:2410.17196, 2024

  103. [115]

    Audio multichallenge: A multi-turn evaluation of spoken dialogue systems on natural human interaction,

    A. Gosai, T. Vuong, U. Tyagi, S. Li, W. You, M. Bavare, A. Uçar, Z. Fang, B. Jang, B. Liuet al., “Audio multichallenge: A multi-turn evaluation of spoken dialogue systems on natural human interaction,” arXiv preprint arXiv:2512.14865, 2025

  104. [116]

    Mtr-duplexbench: Towards a comprehensive evaluation of multi-round conversations for full-duplex speech language models,

    H. Zhang, W. Cui, H. Xu, X. Li, L. Zhu, H. Bai, S. Ma, and I. King, “Mtr-duplexbench: Towards a comprehensive evaluation of multi-round conversations for full-duplex speech language models,”arXiv preprint arXiv:2511.10262, 2025

  105. [117]

    Speech-ifeval: Evaluating instruction-following and quantifying catastrophic forgetting in speech- aware language models,

    K.-H. Lu, C.-Y . Kuan, and H.-y. Lee, “Speech-ifeval: Evaluating instruction-following and quantifying catastrophic forgetting in speech- aware language models,”arXiv preprint arXiv:2505.19037, 2025

  106. [118]

    Inserter: Speech instruction following with unsupervised interleaved pre-training,

    D. Wang, J. Xu, R. Chu, Z. Guo, X. Wang, J. Wu, D. Yang, S. Ji, and J. Lin, “Inserter: Speech instruction following with unsupervised interleaved pre-training,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2...

  107. [119]

    Mmau: A massive multi-task audio understanding and reasoning benchmark,

    S. e. a. Sakshi, “Mmau: A massive multi-task audio understanding and reasoning benchmark,”arXiv preprint arXiv:2410.19168, 2024

  108. [120]

    MMSU: A massive multi-task spoken language understanding and reasoning benchmark,

    D. Wang, J. Wu, J. Li, D. Yang, X. Chen, T. Zhang, and H. Meng, “MMSU: A massive multi-task spoken language understanding and reasoning benchmark,”CoRR, vol. abs/2506.04779, 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2506.04779

  109. [121]

    Mmar: A challenging benchmark for deep reasoning in speech, audio, music, and their mix,

    Z. Ma, Y . Ma, Y . Zhu, C. Yang, Y .-W. Chao, R. Xu, W. Chen, Y . Chen, Z. Chen, J. Conget al., “Mmar: A challenging benchmark for deep reasoning in speech, audio, music, and their mix,”arXiv preprint arXiv:2505.13032, 2025

  110. [122]

    Air-bench: Benchmarking large audio- language models via generative comprehension,

    Q. Yang, J. Xu, W. Liu, Y . Chu, Z. Jiang, X. Zhou, Y . Leng, Y . Lv, Z. Zhao, C. Zhouet al., “Air-bench: Benchmarking large audio- language models via generative comprehension,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume ...

  111. [123]

    Wavbench: Benchmarking reasoning, colloquialism, and paralinguistics for end-to-end spoken dialogue models,

    Y . Li, S. Ji, Y . Chen, T. Liang, H. Ying, Y . Wang, J. Li, J. Fang, and Z. Zhao, “Wavbench: Benchmarking reasoning, colloquialism, and paralinguistics for end-to-end spoken dialogue models,”arXiv preprint arXiv:2602.12135, 2026

  112. [124]

    Mmau-pro: A challenging and comprehensive benchmark for holistic evaluation of audio general intelligence,

    S. Kumar, Š. Sedlá ˇcek, V . Lokegaonkar, F. López, W. Yu, N. Anand, H. Ryu, L. Chen, M. Pli ˇcka, M. Hlavá ˇceket al., “Mmau-pro: A challenging and comprehensive benchmark for holistic evaluation of audio general intelligence,”arXiv preprint arXiv:2508.13992, 2025

  113. [125]

    V oxeval: Benchmarking the knowledge understanding capabilities of end-to-end spoken language models,

    W. Cui, X. Jiao, Z. Meng, and I. King, “V oxeval: Benchmarking the knowledge understanding capabilities of end-to-end spoken language models,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 16 735–16 753

  114. [126]

    Mdar: A multi-scene dynamic audio reasoning benchmark,

    H. Li, C. Jiang, H. Wang, M. Zhang, J. Sun, Z. Yang, Y . Cao, S. Dou, X. Fan, B. Fanet al., “Mdar: A multi-scene dynamic audio reasoning benchmark,”arXiv preprint arXiv:2509.22461, 2025

  115. [127]

    Towards spoken math- ematical reasoning: Benchmarking speech-based models over multi- faceted math problems,

    C. Wei, B. Wang, J.-j. Kim, and N. F. Chen, “Towards spoken math- ematical reasoning: Benchmarking speech-based models over multi- faceted math problems,”arXiv preprint arXiv:2505.15000, 2025

  116. [128]

    Wildspeech-bench: Benchmarking end-to-end speechllms in the wild,

    L. Zhang, J. Zhang, B. Lei, C. Wu, A. Liu, W. Jia, and X. Zhou, “Wildspeech-bench: Benchmarking end-to-end speechllms in the wild,” arXiv preprint arXiv:2506.21875, 2025

  117. [129]

    Uro-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models,

    R. Yan, X. Li, W. Chen, Z. Niu, C. Yang, Z. Ma, K. Yu, and X. Chen, “Uro-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models,”arXiv preprint arXiv:2502.17810, 2025

  118. [130]

    V oiceagentbench: Are voice assistants ready for agentic tasks?

    D. Jain, H. Shukla, G. Rajeev, A. Kulkarni, C. Khatri, and S. Agarwal, “V oiceagentbench: Are voice assistants ready for agentic tasks?”arXiv preprint arXiv:2510.07978, 2025

  119. [131]

    Sakura: On the multi- hop reasoning of large audio-language models based on speech and audio information,

    C.-K. Yang, N. Ho, Y .-T. Piao, and H.-y. Lee, “Sakura: On the multi- hop reasoning of large audio-language models based on speech and audio information,”arXiv preprint arXiv:2505.13237, 2025

  120. [132]

    τ-voice: Benchmarking full-duplex voice agents on real-world domains,

    S. Ray, K. Dhandhania, V . Barres, and K. Narasimhan, “τ-voice: Benchmarking full-duplex voice agents on real-world domains,” 2026. [Online]. Available: https://arxiv.org/abs/2603.13686

  121. [133]

    Audiorag: A challenging bench- mark for audio reasoning and information retrieval,

    J. Lin, C. Zhang, T. Wang, and H. Li, “Audiorag: A challenging bench- mark for audio reasoning and information retrieval,”arXiv preprint arXiv:2602.10656, 2026

  122. [134]

    Full-duplex-bench-v3: Benchmarking tool use for full-duplex voice agents under real-world disfluency,

    G.-T. Lin, C. Chen, Z. Chen, and H.-y. Lee, “Full-duplex-bench-v3: Benchmarking tool use for full-duplex voice agents under real-world disfluency,”arXiv preprint arXiv:2604.04847, 2026. 20 Zhihan Guoreceived the B.Eng. degree from Bei- jing Institute of Technology. She is curr...

Pith tools

Reviewed May 21, 2026 · model on record in the stance chip above.