REVIEW 3 major objections 7 minor 6 cited by
A Preliminary Exploration with GPT-4o Voice Mode
T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This report evaluates GPT-4o's audio understanding across 180 tasks and three benchmarks, arguing that it leads current audio-language models in speech and music reasoning and hallucination resistance, while its built-in safety refusals…
desk verdict First broad public capability map of GPT-4o voice mode, with a genuinely interesting refusal analysis, but the comparative scores hinge on an unstated refusal-scoring rule and an undisclosed benchmark overlap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by three measurement instruments and one refusal detector. Dynamic-SUPERB Phase 2 contributes 180 community-built tasks spanning speech, audio, and music, with relative scores computed against a Whisper-LLaMA cascade baseline and, for classification tasks, against a repeated random-guess baseline. MMAU contributes expert-written multiple-choice questions across audio, music, and speech domains, giving a direct comparison of reasoning and knowledge. CMM contributes the perception accuracy and hallucination resistance metrics that support the central hallucination claim. The refusal detector, combining template string matching with an LLM judge, separates cases where GPT-4o will not answer from cases where it answers incorrectly, which is essential because many of the model's lowest scores are refusals rather than wrong predictions.
What would settle it
Re-run every Dynamic-SUPERB and MMAU item multiple times with varied prompting and sampling, then measure the variance in accuracy and refusal rate; if scores swing widely across runs or reworded instructions, then the claimed strengths and weaknesses are properties of the protocol rather than of the models.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a proprietary end-to-end audio-language model can combine speech, audio, and music understanding with strong instruction following and markedly lower hallucination rates than open large audio-language models, but that its built-in safeguards systematically refuse a class of tasks, including speaker identification, age classification, MOS prediction, and audio deepfake detection. The supporting numbers are an MMAU Test average of 60.46% versus 52.97% for the next proprietary baseline and 52.50% for the best open model, and a CMM hallucination resistance of 83.75% versus 59% for the next-best measured model. On Dynamic-SUPERB, GPT-4o leads on intent classification, stress detection, stuttering detection, multilingual speech recognition, and singing analysis, yet posts negative relative scores on phonological feature classification, audio duration prediction, and several music tasks. The paper treats these scores as a preliminary map of what current LALMs can and cannot do, not as a final verdict on GPT-4o.
Load-bearing premise
The whole comparison stands on the assumption that running the official benchmark pipelines once per sample produces scores that are stable and representative of each model's true ability.
Editorial extensions
If this is right
- If the MMAU numbers hold, GPT-4o voice mode is the strongest audio-language model in this comparison, ahead of the other proprietary model and all open baselines.
- If the CMM numbers hold, GPT-4o's hallucination resistance is substantially higher than every other measured model, even though its perception accuracy is lower than the best open model.
- GPT-4o's refusal pattern means current benchmark leaderboards for proprietary LALMs conflate capability with safety willingness, so future comparisons should report refusal rates alongside accuracy.
- Several Dynamic-SUPERB tasks remain at or below random-guess accuracy for all LALMs, showing that universal instruction-following speech models are still far from solved.
- Tasks that depend mostly on text-transcribable content can still be solved well by cascaded systems, so end-to-end LALMs earn their advantage mainly on acoustic, prosodic, and musical information.
Reading between the lines
- Editorial inference: The near-total refusals on tasks where even strong models perform near chance suggest refusal rate may function as a model-internal signal of low confidence, meaning refusal patterns could be mined as a cheap difficulty label for benchmark items.
- Editorial inference: The finding that random guessing beats many LALMs on hard classification tasks raises the possibility that multiple-choice formats understate generative audio understanding; free-form responses with rubric-based scoring could show a different capability profile.
- Editorial inference: The dataset-dependent refusal rates for fundamentally the same task, such as speaker verification, imply that safety post-training is prompt-sensitive; systematically varying instruction wording would map each model's refusal boundary more precisely before comparing models against each other.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This report evaluates GPT-4o voice mode (gpt-4o-audio-preview-2024-10-01) on three benchmarks: Dynamic-SUPERB Phase 2, MMAU, and CMM. For Dynamic-SUPERB, the authors compute relative scores against Whisper-LLaMA and random baselines and report refusal rates for each task domain. The central qualitative claims are that GPT-4o shows strong audio, speech, and music understanding on several task families (intent classification, multilingual ASR, singing analysis, and hallucination resistance on CMM), while it struggles with duration prediction and instrument classification and frequently refuses tasks related to speaker traits, health inference, deepfake detection, and music analysis. The paper repeatedly frames itself as a preliminary exploration and notes sensitivity to evaluation protocol.
Significance. If the core comparisons are sound, this is a useful early snapshot of a proprietary end-to-end audio-language model, with the additional value of quantifying refusal behavior rather than discarding it. The authors should be credited for reporting refusal rates with two detectors, for introducing a random baseline to contextualize relative scores, and for explicitly acknowledging protocol sensitivity. However, the headline claims are only as reliable as the scoring protocol; the current manuscript leaves the treatment of refusals unspecified, does not define the relative-score metric, and provides no uncertainty quantification.
major comments (3)
- [§4 intro, Fig. 3, Fig. 25] The relative score is never defined. The text says the computation 'follow[s] the setting in Dynamic-SUPERB' but gives no formula or normalization. This matters because the metric produces extreme values when the Whisper-LLaMA baseline is near floor: LTU-AS receives -649.46 on Phoneme Segment Counting (abs diff) in Fig. 3, and SALMONN-13B receives -1771.85 on Audio Duration Prediction in Fig. 25. Define the metric, report the raw scores, and flag or exclude tasks where the baseline denominator is near zero so that the relative numbers are not read as meaningful effect sizes.
- [§3.1, §4–§6] All GPT-4o results appear to come from a single run with no confidence intervals or repeated sampling. Given the observed variability in refusal behavior and API stochasticity, the comparisons—especially small relative-score differences and refusal-rate differences across datasets—are difficult to interpret without uncertainty quantification. The 100-run random baseline does not address model variance; at minimum, the authors should state whether the API calls were repeated and report standard errors or a sensitivity analysis.
- [§3.1, ref [8], author list] The authors are also contributors to Dynamic-SUPERB (ref [8] shares multiple co-authors with this paper) and evaluate GPT-4o on that benchmark using the official scripts, but no conflict-of-interest or contribution disclosure is made. Add an explicit statement identifying the authors' role in constructing the benchmark and the specific tasks or scripts to which they contributed, so that readers can assess the risk of benchmark-construction bias.
minor comments (7)
- [Table 2, §4.2] The task name 'L2 English Accuracy/Fluency/Prodosy Ranking' misspells 'Prosody'; the same typo appears in the text in §4.2.
- [§4 intro] The sentence 'we simply repeats the experiments for 100 times' should read 'we simply repeat the experiments 100 times.'
- [§4.5] Figure 9 and Figure 10 are referenced in reverse order in the text ('Table 5, Figure 10 and Figure 9 respectively'); please correct the ordering or the figure numbering.
- [§4.5, §4.8] 'Whipser-LLaMA' is a typo for 'Whisper-LLaMA' in both sections.
- [§4.6] The phrase 'reflecting the the wealth of textual knowledge' contains a duplicated 'the.'
- [Fig. 25] The GPT-4o entry for 'Audio Duration Prediction' appears blank in the relative-score figure; clarify whether this is a formatting artifact, a refusal, or an N/A case.
- [§4.1] The claim that Random baseline outperforming LALMs 'supports our conjecture that GPT-4o lacks confidence' is not directly supported by that observation; the refusal-confidence link is speculative and should be flagged as such.
Circularity Check
No significant circularity: GPT-4o's scores are empirical measurements against benchmarks whose ground truth is external to the evaluated model, so the central claims do not reduce to the paper's own inputs.
full rationale
This paper is an empirical evaluation report rather than a derivation, so there is no derivation chain whose outputs are defined in terms of its inputs. GPT-4o's performance on Dynamic-SUPERB, MMAU, and CMM is obtained by running a fixed proprietary model against benchmark items with independently fixed labels and official scoring scripts. Several authors are also contributors to Dynamic-SUPERB [8], which is a self-citation, but it is not load-bearing in a circular sense: the benchmark's ground truth is external to the evaluated model, and no parameter of GPT-4o is fitted to those labels here. The refusal-rate analysis in Section 3.2 and the relative-score tables (e.g., Section 4.2) are arithmetic or rule-based transformations of collected responses, not predictions derived from fitted quantities. The extreme relative scores for near-zero baselines, such as LTU-AS's -649.46 on phoneme segment counting, indicate metric instability rather than circularity. The paper repeatedly acknowledges that model performance varies with evaluation protocols (Abstract, Section 7) and notes that some tasks lack official protocols (Section 4.1 footnote) or yield 100% N/A rates (Sections 4.11 and 4.13). It does not explicitly state how refusals are scored in Dynamic-SUPERB, but that is a reproducibility and validity concern, not a circularity concern. No equation or definition in the paper makes the claimed conclusion true by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption Dynamic-SUPERB, MMAU, and CMM provide valid ground truth and fair evaluation pipelines for comparing LALMs.
- domain assumption Relative score to Whisper-LLaMA is a meaningful metric even when the baseline's performance is near zero.
- domain assumption The refusal detection methods (string matching and LLaMA-3.1-8B-Instruct) correctly identify refusals across tasks.
- domain assumption Single evaluation runs of GPT-4o are representative, i.e., the model's responses are stable across API calls.
Cite this review
Pith. "Pith review of A Preliminary Exploration with GPT-4o Voice Mode." pith.science (2026). https://pith.science/paper/NR56754H
@misc{pith2026250209940,
author = {Pith},
title = {Pith review of: A Preliminary Exploration with GPT-4o Voice Mode},
year = {2026},
howpublished = {\url{https://pith.science/paper/NR56754H}},
note = {Machine review of arXiv:2502.09940}
}
read the original abstract
With the rise of multimodal large language models, GPT-4o stands out as a pioneering model, driving us to evaluate its capabilities. This report assesses GPT-4o across various tasks to analyze its audio processing and reasoning abilities. We find that GPT-4o exhibits strong knowledge in audio, speech, and music understanding, performing well in tasks like intent classification, spoken command classification, semantic and grammatical reasoning., multilingual speech recognition, and singing analysis. It also shows greater robustness against hallucinations than other large audio-language models (LALMs). However, it struggles with tasks such as audio duration prediction and instrument classification. Additionally, GPT-4o's safety mechanisms cause it to decline tasks like speaker identification, age classification, MOS prediction, and audio deepfake detection. Notably, the model exhibits a significantly different refusal rate when responding to speaker verification tasks on different datasets. This is likely due to variations in the accompanying instructions or the quality of the input audio, suggesting the sensitivity of its built-in safeguards. Finally, we acknowledge that model performance varies with evaluation protocols. This report only serves as a preliminary exploration of the current state of LALMs.
Figures
Figures from the paper (27 more)
Forward citations
Cited by 6 Pith papers
-
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
Label-free real-vs-noise scoring of audio-encoder neurons, followed by sparse amplification, substantially improves LALM perception of non-semantic speech attributes without retraining.
-
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
Lychee-FD resolves modality interference in full-duplex spoken language models by separating acoustic and semantic parameters in deep layers and adding a dense semantic alignment channel, achieving state-of-the-art pe...
-
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
By projecting hidden states to the vocabulary at every layer, the paper shows that failed attribute recognition in three LALMs is marked by mid-network information peaks followed by degradation, and that models rely o...
-
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
SASTNet, which fuses Whisper semantic features with Wav2Vec2 and AudioMAE acoustic features, improves source tracing for codec-based deepfake speech on CodecFake+, while exposing that prior models overfit to silence.
-
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.
-
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
A proposed audio-native translation framework called MAST with fractional diffusion is described, but no evidence is given that it produces working translations.
Reference graph
Works this paper leans on
-
[8]
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang, Andy T Liu, Chen-An Li, Yu-Xiang Lin, Wei-Cheng Tseng, Anuj Diwan, Yi-Jen Shih, Jiatong Shi, et al. Dynamic-superb phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks. arXiv preprint arXiv:2411.05361, 2024
arXiv 2024
-
[1]
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuan- jun Lv, Jinzheng He, Junyang Lin, et al. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759, 2024
arXiv 2024
-
[2]
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919, 2023
arXiv 2023
-
[3]
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. Moshi: a speech-text foundation model for real-time dialogue. arXiv preprint arXiv:2410.00037, 2024. 34
-
[4]
Audio entailment: Assessing deductive reasoning for audio understanding
Soham Deshmukh, Shuo Han, Hazim Bukhari, Benjamin Elizalde, Hannes Gamper, Rita Singh, and Bhiksha Raj. Audio entailment: Assessing deductive reasoning for audio understanding. arXiv preprint arXiv:2407.18062, 2024
arXiv 2024
-
[5]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024
arXiv 2024
-
[6]
Llama-omni: Seamless speech interaction with large language models
Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, and Yang Feng. Llama-omni: Seamless speech interaction with large language models. arXiv preprint arXiv:2409.06666, 2024
arXiv 2024
-
[7]
Wavllm: Towards robust and adaptive speech large language model
Shujie Hu, Long Zhou, Shujie Liu, Sanyuan Chen, Hongkun Hao, Jing Pan, Xunying Liu, Jinyu Li, Sunit Sivasankaran, Linquan Liu, et al. Wavllm: Towards robust and adaptive speech large language model. arXiv preprint arXiv:2404.00656, 2024
arXiv 2024
Show all 31 references
-
[9]
Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech
Chien-yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Haibin Wu, Siddhant Arora, Kai-Wei Chang, Jiatong Shi, Yifan Peng, et al. Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech. In ICASSP 2024-2024 I...
2024
-
[10]
Audiogpt: Understanding and generating speech, music, sound, and talking head
Rongjie Huang et al. Audiogpt: Understanding and generating speech, music, sound, and talking head. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 38, pages 23802–23804, 2024
2024
-
[11]
Mistral 7b
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023
-
[12]
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in neural information processing systems, 33:17022–17033, 2020
2020
-
[13]
Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation
Chun-Yi Kuan, Chih-Kai Yang, Wei-Ping Huang, Ke-Han Lu, and Hung-yi Lee. Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation. In 2024 IEEE Spoken Language Technology Workshop (SLT), pages 1060–10...
2024
-
[14]
Rlaif vs
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash. Rlaif vs. rlhf: scaling reinforcement learning from human feedback with ai feedback. In Proceedings of the 41...
2024
-
[15]
The curse of multi-modalities: Evaluating hallucinations of large multimodal models across language, visual, and audio
Sicong Leng, Yun Xing, Zesen Cheng, Yang Zhou, Hang Zhang, Xin Li, Deli Zhao, Shijian Lu, Chunyan Miao, and Lidong Bing. The curse of multi-modalities: Evaluating hallucinations of large multimodal models across language, visual, and audio. arXiv preprint arXiv:2410.12787, 2024
-
[16]
Align-slm: Textless spoken language models with reinforcement learning from ai feedback
Guan-Ting Lin, Prashanth Gurunath Shivakumar, Aditya Gourav, Yile Gu, Ankur Gandhe, Hung- yi Lee, and Ivan Bulyko. Align-slm: Textless spoken language models with reinforcement learning from ai feedback. arXiv preprint arXiv:2411.01834, 2024
2024 arXiv
-
[17]
Music understand- ing llama: Advancing text-to-music generation with question answering and captioning
Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan. Music understand- ing llama: Advancing text-to-music generation with question answering and captioning. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages...
2024
-
[18]
Generative spoken dialogue language modeling
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al. Generative spoken dialogue language modeling. Transactions of the Association for Computational Linguistics, 11:250–266, 2023
2023
- [19]
-
[20]
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:2773...
2022
-
[21]
Mmau: A massive multi-task audio understanding and reasoning benchmark
S Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, Ramaneswaran Selvakumar, Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, and Dinesh Manocha. Mmau: A massive multi-task audio understanding and reasoning benchmark. arXiv preprint arXiv:2410.19168, 2024
-
[22]
Salmonn: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, MA Zejun, and Chao Zhang. Salmonn: Towards generic hearing abilities for large language models. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[23]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024
2024 arXiv
-
[24]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[25]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[26]
Hear: Holistic evaluation of audio representations
Joseph Turian, Jordie Shier, Humair Raj Khan, Bhiksha Raj, Björn W Schuller, Christian J Steinmetz, Colin Malloy, George Tzanetakis, Gissel Velarde, Kirk McNally, et al. Hear: Holistic evaluation of audio representations. In NeurIPS 2021 Competitions and Demonstrations Track, ...
2021
-
[27]
Lin, Andy T
Shu wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y . Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman M...
2021
-
[28]
Mini-omni: Language models can hear, talk while thinking in streaming
Zhifei Xie and Changqiao Wu. Mini-omni: Language models can hear, talk while thinking in streaming. arXiv preprint arXiv:2408.16725, 2024
2024 arXiv
-
[29]
Marble: Music audio representation benchmark for universal evaluation
Ruibin Yuan, Yinghao Ma, Yizhi Li, Ge Zhang, Xingran Chen, Hanzhi Yin, Yiqi Liu, Jiawen Huang, Zeyue Tian, Binyue Deng, et al. Marble: Music audio representation benchmark for universal evaluation. Advances in Neural Information Processing Systems, 36:39626–39647, 2023
2023
-
[30]
Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. arXiv preprint arXiv:2305.11000, 2023
2023 arXiv
-
[31]
Speechalign: Aligning speech generation to human preferences
Dong Zhang, Zhaowei Li, Shimin Li, Xin Zhang, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. Speechalign: Aligning speech generation to human preferences. arXiv preprint arXiv:2404.05600, 2024. 36
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.