REVIEW 5 major objections 6 minor 1 cited by
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A new benchmark of 1,680 adversarial audio samples tests six voice assistants and reports GPT-4o as the most resilient.
desk verdict Useful benchmark, but the 'GPT-4o clearly best overall' claim is not backed by the paper's own tables—on implicit noise Llama-Omni wins on multiple metrics in all three evaluations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the CAA quadruplet $(a_i, t_i, a_i^{\text{no\_attack}}, A_i)$: an original utterance, its transcript, a clean agent-read audio recording, and a set of attack variants. Content attacks are generated by GPT-4-guided synonym substitution, token rearrangement, or minimal token variation; emotional attacks re-synthesize the transcript with an opposite emotion or overlay opposite-mood background music; explicit noise attacks overlay natural, industrial, or human noise; implicit noise attacks overlay a 15 Hz infrasound or a 22 kHz ultrasound signal. Evaluation machinery then compares no-attack versus attacked responses with word error rate, ROUGE-L, and cosine similarity, plus GPT-4o-rated no-attack coherence, attack coherence, attack correlation, and linguistic robustness, and human ratings of coherence.
What would settle it
Run the 1,680 CAA samples through the six models, then have blind human raters score the responses without knowing which model produced them and without letting GPT-4o score its own outputs, and check whether GPT-4o still leads on the coherence, correlation, and linguistic-robustness metrics. If its rank drops, the paper's central claim fails; a quicker check is whether GPT-4o's self-scores on its own outputs exceed human scores on the same outputs.
Extended reading notes
Core claim
The central claim is that conversational robustness of large audio-language models can be measured by a universal, content-preserving attack benchmark, and that under that benchmark GPT-4o clearly outperforms the other five evaluated models. The paper asserts that GPT-4o consistently delivers coherent, contextually relevant, and linguistically solid responses even under severe adversarial conditions, and attributes this to its extensive pre-training on large-scale datasets. The benchmark itself is the other part of the discovery: four attack families (content, emotional, explicit noise, and implicit noise) applied to re-synthesized clear speech provide a standardized way to compare vulnerabilities that previous model-specific targeted attacks did not offer.
Load-bearing premise
The ranking depends on GPT-4o being an impartial judge of all six models, including itself; if its self-scores are inflated, the headline result weakens.
Editorial extensions
If this is right
- Models that transcribe audio before responding, such as SpeechGPT, degrade more under all four attack families than models that process audio more directly.
- Natural noise is the most damaging explicit noise category overall, and infrasound is more damaging than ultrasound for most of the evaluated models.
- Resilience to emotional mismatch is not necessarily good news, because it coincides with low emotional awareness, which is a weakness for natural conversation.
- Training on noisy audio and large-scale pre-training are the main observed correlates of resilience, pointing to data diversity as a defense strategy.
- The released generation scripts let the benchmark grow beyond the current 1,680 samples, so the same attack families can be applied to new utterances and models.
Reading between the lines
- An implication the authors leave implicit is that CAA only tests attacks that preserve human intelligibility and overall meaning, so it likely underestimates gradient-based targeted attacks tuned to a specific model; a high CAA rank does not guarantee resistance to those attacks.
- A testable extension would be to add audio jailbreak samples once open-source audio jailbreak methods exist, since the paper notes it could not generate such samples and this remains an underexplored threat.
- The finding that inaudible 15 Hz infrasound degrades several models more than 22 kHz ultrasound suggests a concrete mechanism check: comparing model input-embedding sensitivity to sub-20 Hz spectral energy would connect benchmark results to underlying audio encoders.
- A practical deployment consequence is that voice assistants relying on speech-to-text front ends should expect adversarial audio to corrupt the transcription step, so robustness claims should be reported with per-attack transcription accuracy alongside final response quality.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Chat-Audio Attacks (CAA), a benchmark of 1,680 adversarial audio samples built from 360 speech utterances drawn from MELD, TVQA, and Common Voice. Each utterance is turned into no-attack, content-attack, emotional-attack, explicit-noise-attack, and implicit-noise-attack audio variants. The authors evaluate six large audio-language models (SpeechGPT, SALMONN, Qwen2-Audio, LLama-Omni, Gemini-1.5-Pro, GPT-4o) with three evaluation strategies: standard metrics (WER, ROUGE-L, cosine similarity), GPT-4o-based scores (NC, ACoh, ACor, LR), and human ratings (NC, ACoh). The paper concludes that GPT-4o is the most resilient LALM overall. Data and generation scripts are released publicly.
Significance. The construction of a public, multi-attack audio benchmark with release scripts is a useful contribution to a young area, and the three evaluation channels are a thoughtful attempt to go beyond ASR-style metrics. The sample size is reasonable for a first benchmark, and the tables appear internally consistent. The main value of the paper currently depends on the reliability of its comparative ranking, which is not yet established: the headline claim that GPT-4o is clearly best overall is underdetermined by the paper's own tables and by the absence of a predefined aggregation rule, uncertainty estimates, and inter-rater agreement measures. If these issues are repaired, the benchmark and its public artifacts would be a solid resource for evaluating LALM robustness.
major comments (5)
- [Section 4, Tables 2-4] The headline claim that 'GPT-4o clearly emerges as the best-performing model overall' is not derived from any stated aggregation rule across the four attack families and three evaluation channels, and no significance or variance information is reported. The tables themselves contradict a uniform ranking: for ultrasound attacks, Table 2 gives LLama-Omni WER 0.37 vs GPT-4o 1.13, ROUGE-L 0.75 vs 0.17, and COS 0.79 vs 0.28; for infrasound, WER 0.67 vs 1.25, ROUGE-L 0.56 vs 0.22, and COS 0.63 vs 0.35. Table 3 shows LLama-Omni ACoh 3.31 vs 2.70 and ACor 3.53 vs 2.26 on ultrasound, and Table 4 shows LLama-Omni human ACoh 3.15 vs 3.08 on implicit noise. The Discussion itself states that Llama-Omni shows greater robustness on both types of implicit noise. Under a different weighting of attack families, the 'overall best' verdict would change; the conclusion therefore needs a transparent aggregation rule plus uncertainty estimates.
- [Section 3.3] The GPT-4o-Based Evaluation uses GPT-4o as the scoring model for outputs generated by GPT-4o, so the top ranking on this channel may be inflated by self-scoring bias. No comparison of GPT-4o scores against human scores is reported, and no alternative judge or score-distribution analysis is provided. The paper should either calibrate this channel against human ratings, use a different judge model, or restrict the conclusion to the standard and human channels.
- [Section 3.4] The human evaluation relies on five raters and reports only averaged scores. Inter-rater agreement (e.g., Krippendorff's alpha or Fleiss' kappa) and per-cell variance or confidence intervals are missing, which makes differences such as the implicit-noise ACoh of 3.15 vs 3.08 in Table 4 uninterpretable. Without these, the human channel cannot substantiate the ranking.
- [Section 3.1] The sentence 'In these evaluation methods, all audio content is presented in the form of transcribed text' implies that the GPT-4o-based and human judges never listen to the attacked audio; they rate text transcripts of model outputs against text transcripts of the input. This means the three evaluation channels mostly measure robustness of the semantic/response layer to the content of the attacks, not the model's acoustic processing. The benchmark's claim to evaluate audio attacks would be strengthened by at least one listening-based evaluation or by a clear argument that transcript-based evaluation is sufficient.
- [Sections 2.2 and 2.3] The no-attack and the content/emotion attack audio are re-synthesized with AzureSpeechSDK rather than being natural speech, and explicit/implicit noise is overlaid programmatically without verification of the frequency content received by each model's audio encoder. This limits ecological validity: the baseline is TTS speech, and the implicit-noise attacks assume, rather than check, that the inaudible tones are actually present in the model input rather than filtered out by preprocessing. Please report spectra or input-level checks, or qualify the conclusions accordingly.
minor comments (6)
- [Throughout] Typos and inconsistent capitalization: 'discusse' in Section 1, 'Common V oice' used repeatedly, and 'LLama-Omni' vs 'Llama-Omni' are inconsistent.
- [References] The reference list contains a placeholder '?' after Köpf et al., a malformed inline citation '(aud, 2023)', and an incomplete entry for 'Szegedy, 2013'; the GPT-4 citations should point to the GPT-4 technical report rather than the ChatGPT homepage.
- [Table 5] The GPT-4o row has empty cells for Parameters, Language Model, and Audio Model; these should read 'not disclosed' to avoid implying that the information is absent.
- [Table 6] The SpeechGPT row for Opp-Emo Music shows an empty response; the paper should state how empty outputs were handled in WER, ROUGE-L, COS, and the LLM/human ratings.
- [Introduction] The text says each of the 360 sets 'encompassing four distinct types of audio attacks,' but Table 1 shows emotional attacks exist only for MELD and implicit/explicit noise counts differ by source; clarify the set structure.
- [Section 4] The attribution of GPT-4o's robustness to 'extensive pre-training on large-scale datasets' is speculative and not supported by evidence in this paper; consider removing or softening the claim.
Circularity Check
Partial circularity: GPT-4o is the judge in one evaluation channel and also the model ranked first, so the 'best overall' claim is partly self-scored.
-
other
[Section 3.3 (GPT-4o-Based Evaluation) and Section 4 (Discussion), with Tables 3 and 5]
"The GPT-4o-Based Evaluation utilizes a more sophisticated set of criteria, leveraging the advanced capabilities of GPT-4o to simulate real-world interaction scenarios and evaluate the effects of adversarial attacks on model behavior. ... GPT-4o clearly emerges as the best-performing model overall."
Section 3.1 lists GPT-4o among the six evaluated models, so the Table 3 scores that rank GPT-4o first on NC, ACoh, ACor, and LR are GPT-4o's own ratings of its own outputs, with no reported calibration against human judgments for this channel. The Discussion's 'best overall' verdict draws on this self-scored channel together with Tables 2 and 4, with no pre-specified aggregation rule. On the GPT-4o-Based channel, GPT-4o's rank is set by GPT-4o's own judgment by construction, making one of the three evaluation pillars a self-assessment rather than an independent measurement. The CAA benchmark itself and the Standard and Human evaluations are independent, so the circularity is partial.
full rationale
No load-bearing self-citations or imported uniqueness theorems appear: the benchmark construction, attack generation, and standard metric comparisons are self-contained and do not reduce to the conclusions. The main circular feature is the GPT-4o-Based Evaluation, where the scoring model is identical to the top-ranked system; this affects the strength of the headline robustness claim but does not by itself determine all three evaluation channels. The paper's other notable weakness, the missing aggregation rule behind 'GPT-4o clearly emerges as the best-performing model overall,' is an underdetermination or correctness issue rather than a circularity, and the Discussion even concedes that Llama-Omni is more robust on implicit noise. The unlabeled '(Köpf et al., 2024;?)' citation in the Introduction is an incomplete reference, not a circularity. Overall, one evaluation channel is self-referential, so the score is moderate.
Assumptions & free parameters
assumptions (4)
- domain assumption GPT-4o is a valid and impartial automated evaluator for response coherence, correlation, and linguistic robustness.
- domain assumption Azure TTS re-synthesis of transcripts preserves the conversational properties of the original human speech well enough to serve as a baseline.
- domain assumption Inaudible signals below 20 Hz or above 20 kHz are a meaningful robustness test and do not distort the speech for listeners.
- domain assumption Transcribed text of model responses captures the full effect of audio attacks.
Cite this review
Pith. "Pith review of Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models." pith.science (2026). https://pith.science/paper/DKS4J4SV
@misc{pith2026241114842,
author = {Pith},
title = {Pith review of: Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/DKS4J4SV}},
note = {Machine review of arXiv:2411.14842}
}
read the original abstract
Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research focused on model-specific adversarial methods, real-world applications demand a more generalizable and universal approach to audio adversarial attacks. In this paper, we introduce the Chat-Audio Attacks (CAA) benchmark including four distinct types of audio attacks, which aims to explore the vulnerabilities of LALMs to these audio attacks in conversational scenarios. To evaluate the robustness of LALMs, we propose three evaluation strategies: Standard Evaluation, utilizing traditional metrics to quantify model performance under attacks; GPT-4o-Based Evaluation, which simulates real-world conversational complexities; and Human Evaluation, offering insights into user perception and trust. We evaluate six state-of-the-art LALMs with voice interaction capabilities, including Gemini-1.5-Pro, GPT-4o, and others, using three distinct evaluation methods on the CAA benchmark. Our comprehensive analysis reveals the impact of four types of audio attacks on the performance of these models, demonstrating that GPT-4o exhibits the highest level of resilience. Our data can be accessed via the following link: \href{https://github.com/crystraldo/CAA}{CAA}.
Figures
Forward citations
Cited by 1 Pith paper
-
Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models
A survey that organizes audio and video AI security research into adversarial, backdoor, and jailbreak attacks, with extra attention to multimodal large language models.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
2023. https://mindgard.ai/resources/audio-based-jailbreak-attacks-on-multi-modal-llms?hs_amp=true Audio-based jailbreak attacks on multi-modal llms . Accessed: 2023-10-15
work page 2023
-
[4]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[5]
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. 2019. Common voice: A massively-multilingual speech corpus. arXiv preprint arXiv:1912.06670
arXiv 2019
-
[6]
Zal \'a n Borsos, Rapha \"e l Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2022. Audiolm: a language modeling approach to audio generation.(2022). arXiv preprint arXiv:2209.03143
arXiv 2022
-
[7]
Nicholas Carlini and David Wagner. 2018. Audio adversarial examples: Targeted attacks on speech-to-text. In 2018 IEEE security and privacy workshops (SPW), pages 1--7. IEEE
work page 2018
-
[8]
Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, and Bo Xu. 2023. X-llm: Bootstrapping advanced large language models by treating multi-modalities as foreign languages. arXiv preprint arXiv:2305.04160
arXiv 2023
Show all 45 references
-
[9]
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. 2024. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759
2024 arXiv
-
[10]
Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, and Yang Feng. 2024. Llama-omni: Seamless speech interaction with large language models. arXiv preprint arXiv:2409.06666
2024 arXiv
-
[11]
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, et al. 2024. Prompting large language models with speech recognition abilities. In ICASSP 2024-2024 IEEE International Conference on Acoust...
2024
-
[12]
Yuan Gong and Christian Poellabauer. 2017. Crafting adversarial examples for speech paralinguistics applications. arXiv preprint arXiv:1711.03280
2017 arXiv
-
[13]
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572
2014 arXiv
-
[14]
Alex Graves, Santiago Fern \'a ndez, Faustino Gomez, and J \"u rgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning, pages 369--376
2006
-
[15]
Muhammad Usman Hadi, Qasem Al Tashi, Abbas Shah, Rizwan Qureshi, Amgad Muneer, Muhammad Irfan, Anas Zafar, Muhammad Bilal Shaikh, Naveed Akhtar, Jia Wu, et al. 2024. Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospect...
2024
-
[16]
Zhankui He, Zhouhang Xie, Rahul Jha, Harald Steck, Dawen Liang, Yesu Feng, Bodhisattwa Prasad Majumder, Nathan Kallus, and Julian McAuley. 2023. Large language models as zero-shot conversational recommenders. In Proceedings of the 32nd ACM international conference on informati...
2023
-
[17]
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is bert really robust? a strong baseline for natural language attack on text classification and entailment. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 8018--8025
2020
-
[18]
Andre Kassis and Urs Hengartner. 2021. Practical attacks on voice spoofing countermeasures. arXiv preprint arXiv:2107.14642
2021 arXiv
-
[19]
Corey Kereliuk, Bob L Sturm, and Jan Larsen. 2015. Deep learning and music adversaries. IEEE Transactions on Multimedia, 17(11):2059--2071
2015
-
[20]
Saydulu Kolasani. 2023. Optimizing natural language processing, large language models (llms) for efficient customer service, and hyper-personalization to enable sustainable growth and revenue. Transactions on Latest Trends in Artificial Intelligence, 4(4)
2023
-
[21]
o pf, Yannic Kilcher, Dimitri von R \
Andreas K \"o pf, Yannic Kilcher, Dimitri von R \"u tte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Rich \'a rd Nagyfi, et al. 2024. Openassistant conversations-democratizing large language model alignment. Advances in Neura...
2024
-
[22]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25
2012
-
[23]
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al. 2021. On generative spoken language modeling from raw audio. Transactions of the Association for Computational Lingu...
2021
-
[24]
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018. Tvqa: Localized, compositional video question answering. arXiv preprint arXiv:1809.01696
2018 arXiv
-
[25]
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020. Bert-attack: Adversarial attack against bert using bert. arXiv preprint arXiv:2004.09984
2020 arXiv
-
[26]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74--81
2004
-
[27]
Microsoft. 2023. Azure Cognitive Services Speech SDK . https://github.com/Azure-Samples/cognitive-services-speech-sdk
2023
-
[28]
Paarth Neekhara, Shehzeen Hussain, Prakhar Pandey, Shlomo Dubnov, Julian McAuley, and Farinaz Koushanfar. 2019. Universal adversarial perturbations for speech recognition systems. arXiv preprint arXiv:1905.03828
2019 arXiv
-
[29]
OpenAI . 2023. Chatgpt. https://openai.com/blog/chatgpt/. 1, 2
2023
-
[30]
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2018. Meld: A multimodal multi-party dataset for emotion recognition in conversations. arXiv preprint arXiv:1810.02508
2018 arXiv
-
[31]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492--28518. PMLR
2023
-
[32]
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arX...
2024 arXiv
-
[33]
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Semantically equivalent adversarial rules for debugging nlp models. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (volume 1: long papers), pages 856--865
2018
-
[34]
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu-Ghazaleh. 2023. Survey of vulnerabilities in large language models revealed by adversarial attacks. arXiv preprint arXiv:2310.10844
2023 arXiv
-
[35]
Shyam Sudhakaran, Miguel Gonz \'a lez-Duque, Matthias Freiberger, Claire Glanois, Elias Najarro, and Sebastian Risi. 2024. Mariogpt: Open-ended text2level generation through large language models. Advances in Neural Information Processing Systems, 36
2024
-
[36]
C Szegedy. 2013. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199
2013 arXiv
-
[37]
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023. Salmonn: Towards generic hearing abilities for large language models. arXiv preprint arXiv:2310.13289
2023 arXiv
-
[38]
Graham Todd, Sam Earle, Muhammad Umair Nasir, Michael Cerny Green, and Julian Togelius. 2023. Level generation through large language models. In Proceedings of the 18th International Conference on the Foundations of Digital Games, pages 1--8
2023
-
[39]
Jason Wei and Kai Zou. 2019. Eda: Easy data augmentation techniques for boosting performance on text classification tasks. arXiv preprint arXiv:1901.11196
2019 arXiv
-
[40]
Jian Wu, Yashesh Gaur, Zhuo Chen, Long Zhou, Yimeng Zhu, Tianrui Wang, Jinyu Li, Shujie Liu, Bo Ren, Linquan Liu, et al. 2023. On decoder-only architecture for speech-to-text and large language model integration. In 2023 IEEE Automatic Speech Recognition and Understanding Work...
2023
-
[41]
Yi Xie, Zhuohang Li, Cong Shi, Jian Liu, Yingying Chen, and Bo Yuan. 2021. Enabling fast and universal audio adversarial attack using generative model. In Proceedings of the AAAI conference on artificial intelligence, volume 35, pages 14129--14137
2021
-
[42]
Yi Xie, Cong Shi, Zhuohang Li, Jian Liu, Yingying Chen, and Bo Yuan. 2020. Real-time, universal, and robust adversarial attacks against speaker recognition systems. In ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 173...
2020
-
[43]
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. arXiv preprint arXiv:2305.11000
2023 arXiv
-
[44]
Xingyu Zhang, Xiongwei Zhang, Wei Liu, Xia Zou, Meng Sun, and Jian Zhao. 2022. Waveform level adversarial example generation for joint attacks against both automatic speaker verification and spoofing countermeasures. Engineering Applications of Artificial Intelligence, 116:105469
2022
-
[45]
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Man Cheung, and Min Lin. 2024. On evaluating adversarial robustness of large vision-language models. Advances in Neural Information Processing Systems, 36
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.