{"id":"34b44db9-e6e9-4267-93b6-ad0ce9c20ef7","arxiv_id":"2411.14842","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The CAA benchmark applies content, emotional, explicit noise, and implicit noise attacks to six audio-language models and finds GPT-4o the most robust.","lead":"This paper introduces Chat-Audio Attacks (CAA), a benchmark with 1,680 audio samples that stress-test voice-enabled AI models with four kinds of acoustic attack. It ranks six large audio-language models and reports that GPT-4o resists these attacks best. A smart generalist might read it to see how fragile current voice AI is against noisy, emotional, or altered audio.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GPT-4o ranking rests on an unstated aggregate; on implicit-noise attacks LLama-Omni beats GPT-4o across several metrics in all three evaluations, so 'clearly best overall' is not established by the paper's own tables.","rationale":"The reader identified GPT-4o self-scoring as the weakest assumption. That concern is legitimate but partially mitigated: Table 4 (human evaluation) also ranks GPT-4o first on NC and on most ACoh entries, so the top ranking is not solely an artifact of the judge being GPT-4o. The more load-bearing issue is that the headline conclusion requires an aggregation over attack types, and no such aggregation is defined or justified. On implicit noise, the paper's own results show LLama-Omni ahead of GPT-4o on standard metrics and human ACoh, and the Discussion explicitly acknowledges this. Thus 'clearly best overall' is an overstatement without a stated weighting or statistical test. This does not invalidate the benchmark contribution; it means the central claim must be qualified. The reader's CONDITIONAL verdict remains appropriate, so the verdict should be UNCHANGED, with the condition sharpened to require either an explicit aggregation rule or a per-family reporting of the ranking.","tokens_in":15525,"tokens_out":8839,"duration_ms":76179,"concrete_test":"Download the released CAA data and compute a pre-registered aggregate robustness score: for each attack family (content, emotion, explicit noise, implicit noise), average the per-model normalized scores across WER, ROUGE-L, COS, ACoh, ACor, LR, and human NC/ACoh, giving equal weight to each family. Then bootstrap over the 360 samples to obtain confidence intervals for the aggregate. If LLama-Omni or Gemini-1.5-Pro is within one bootstrap standard error of GPT-4o (or above it) on the equal-weighted aggregate, the claim 'clearly emerges as the best-performing model overall' must be replaced by a qualified, per-family statement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'GPT-4o clearly emerges as the best-performing model overall' (Section 4) requires an aggregation across four attack families, but no aggregation rule or significance test is provided. The paper's own tables contradict a uniform ranking: for implicit noise (infrasound/ultrasound), LLama-Omni outperforms GPT-4o on standard metrics (Table 2: ultrasound WER 0.37 vs 1.13, ROUGE-L 0.75 vs 0.17, COS 0.79 vs 0.28; infrasound WER 0.67 vs 1.25, ROUGE-L 0.56 vs 0.22, COS 0.63 vs 0.35), on the GPT-4o-based evaluation (Table 3: ultrasound ACoh 3.31 vs 2.70, ACor 3.53 vs 2.26), and on human ACoh (Table 4: 3.15 vs 3.08). The Discussion itself concedes that Llama-Omni demonstrates greater robustness on both types of implicit noise. Since the 'overall best' verdict is not derived from a predefined weighted score, the conclusion is underdetermined: a model can be best under one weighting and second-best under another. The GPT-4o self-scoring concern raised by the reader is real, but the human evaluation partially corroborates GPT-4o's top ranking on content and emotion; the more decisive problem is the missing aggregation and the explicit contradiction on implicit noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Chat-Audio Attacks (CAA), a benchmark of 1,680 adversarial audio samples built from 360 speech utterances drawn from MELD, TVQA, and Common Voice. Each utterance is turned into no-attack, content-attack, emotional-attack, explicit-noise-attack, and implicit-noise-attack audio variants. The authors evaluate six large audio-language models (SpeechGPT, SALMONN, Qwen2-Audio, LLama-Omni, Gemini-1.5-Pro, GPT-4o) with three evaluation strategies: standard metrics (WER, ROUGE-L, cosine similarity), GPT-4o-based scores (NC, ACoh, ACor, LR), and human ratings (NC, ACoh). The paper concludes that GPT-4o is the most resilient LALM overall. Data and generation scripts are released publicly.","tokens_in":15850,"tokens_out":6762,"duration_ms":60094,"significance":"The construction of a public, multi-attack audio benchmark with release scripts is a useful contribution to a young area, and the three evaluation channels are a thoughtful attempt to go beyond ASR-style metrics. The sample size is reasonable for a first benchmark, and the tables appear internally consistent. The main value of the paper currently depends on the reliability of its comparative ranking, which is not yet established: the headline claim that GPT-4o is clearly best overall is underdetermined by the paper's own tables and by the absence of a predefined aggregation rule, uncertainty estimates, and inter-rater agreement measures. If these issues are repaired, the benchmark and its public artifacts would be a solid resource for evaluating LALM robustness.","major_comments":[{"comment":"The headline claim that 'GPT-4o clearly emerges as the best-performing model overall' is not derived from any stated aggregation rule across the four attack families and three evaluation channels, and no significance or variance information is reported. The tables themselves contradict a uniform ranking: for ultrasound attacks, Table 2 gives LLama-Omni WER 0.37 vs GPT-4o 1.13, ROUGE-L 0.75 vs 0.17, and COS 0.79 vs 0.28; for infrasound, WER 0.67 vs 1.25, ROUGE-L 0.56 vs 0.22, and COS 0.63 vs 0.35. Table 3 shows LLama-Omni ACoh 3.31 vs 2.70 and ACor 3.53 vs 2.26 on ultrasound, and Table 4 shows LLama-Omni human ACoh 3.15 vs 3.08 on implicit noise. The Discussion itself states that Llama-Omni shows greater robustness on both types of implicit noise. Under a different weighting of attack families, the 'overall best' verdict would change; the conclusion therefore needs a transparent aggregation rule plus uncertainty estimates.","section":"Section 4, Tables 2-4"},{"comment":"The GPT-4o-Based Evaluation uses GPT-4o as the scoring model for outputs generated by GPT-4o, so the top ranking on this channel may be inflated by self-scoring bias. No comparison of GPT-4o scores against human scores is reported, and no alternative judge or score-distribution analysis is provided. The paper should either calibrate this channel against human ratings, use a different judge model, or restrict the conclusion to the standard and human channels.","section":"Section 3.3"},{"comment":"The human evaluation relies on five raters and reports only averaged scores. Inter-rater agreement (e.g., Krippendorff's alpha or Fleiss' kappa) and per-cell variance or confidence intervals are missing, which makes differences such as the implicit-noise ACoh of 3.15 vs 3.08 in Table 4 uninterpretable. Without these, the human channel cannot substantiate the ranking.","section":"Section 3.4"},{"comment":"The sentence 'In these evaluation methods, all audio content is presented in the form of transcribed text' implies that the GPT-4o-based and human judges never listen to the attacked audio; they rate text transcripts of model outputs against text transcripts of the input. This means the three evaluation channels mostly measure robustness of the semantic/response layer to the content of the attacks, not the model's acoustic processing. The benchmark's claim to evaluate audio attacks would be strengthened by at least one listening-based evaluation or by a clear argument that transcript-based evaluation is sufficient.","section":"Section 3.1"},{"comment":"The no-attack and the content/emotion attack audio are re-synthesized with AzureSpeechSDK rather than being natural speech, and explicit/implicit noise is overlaid programmatically without verification of the frequency content received by each model's audio encoder. This limits ecological validity: the baseline is TTS speech, and the implicit-noise attacks assume, rather than check, that the inaudible tones are actually present in the model input rather than filtered out by preprocessing. Please report spectra or input-level checks, or qualify the conclusions accordingly.","section":"Sections 2.2 and 2.3"}],"minor_comments":[{"comment":"Typos and inconsistent capitalization: 'discusse' in Section 1, 'Common V oice' used repeatedly, and 'LLama-Omni' vs 'Llama-Omni' are inconsistent.","section":"Throughout"},{"comment":"The reference list contains a placeholder '?' after Köpf et al., a malformed inline citation '(aud, 2023)', and an incomplete entry for 'Szegedy, 2013'; the GPT-4 citations should point to the GPT-4 technical report rather than the ChatGPT homepage.","section":"References"},{"comment":"The GPT-4o row has empty cells for Parameters, Language Model, and Audio Model; these should read 'not disclosed' to avoid implying that the information is absent.","section":"Table 5"},{"comment":"The SpeechGPT row for Opp-Emo Music shows an empty response; the paper should state how empty outputs were handled in WER, ROUGE-L, COS, and the LLM/human ratings.","section":"Table 6"},{"comment":"The text says each of the 360 sets 'encompassing four distinct types of audio attacks,' but Table 1 shows emotional attacks exist only for MELD and implicit/explicit noise counts differ by source; clarify the set structure.","section":"Introduction"},{"comment":"The attribution of GPT-4o's robustness to 'extensive pre-training on large-scale datasets' is speculative and not supported by evidence in this paper; consider removing or softening the claim.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The benchmark and data release are valuable, and the taxonomy is clear. However, the current framing of the overall ranking overstates what the evidence supports, and the paper's own tables undermine the 'clearly best overall' verdict. A revision that reframes the results as multidimensional, adds uncertainty and agreement measures, and addresses the self-scoring issue would make the contribution sound. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the CAA benchmark is a genuine contribution and worth taking seriously, but the headline result does not survive contact with the paper's own tables. GPT-4o is strong on content and emotion attacks, but on implicit noise Llama-Omni beats it on WER, ROUGE-L, and COS in Table 2, on ACoh and ACor in Table 3, and on human ACoh in Table 4. The Discussion even concedes Llama-Omni handles implicit noise better. So 'clearly emerges as the best-performing model overall' (Section 4) is underdetermined: there is no stated aggregation rule or significance test across the four attack families.\n\nWhat is genuinely new: this is the first benchmark I know of that combines content, emotional, explicit-noise, and implicit-noise attacks on LALMs under conversational conditions, and the three-evaluation design (standard metrics, GPT-4o judge, human raters) is a reasonable way to triangulate robustness. The dataset of 1,680 samples and the release link are useful. The attack construction is not deep—TTS re-reading plus signal overlays—but that is fine for a benchmark whose value is in the comparison.\n\nSoft spots, in order of seriousness. First, the missing aggregation and the implicit-noise contradiction. Second, GPT-4o serves as both the model being ranked and the judge in one of the three evaluation channels; this is a genuine partial circularity, though human scores independently corroborate GPT-4o's top position on content and emotion. Third, no error bars or statistical tests, and the human evaluation uses five raters without agreement metrics. Fourth, the audio is TTS-synthesized, which limits ecological validity for real speech. The limitations section is honest about the controlled setting and the absence of audio jailbreak attacks.\n\nBottom line: the paper is a solid empirical contribution to an under-served subfield, but the central robustness claim needs to be reined in or properly aggregated. I would send it to peer review, with the expectation that the authors add uncertainty quantification, define the aggregation, and recalibrate the 'best overall' claim. I'd also want the data/code checked before publication.","headline":"Useful benchmark, but the 'GPT-4o clearly best overall' claim is not backed by the paper's own tables—on implicit noise Llama-Omni wins on multiple metrics in all three evaluations.","tokens_in":16331,"tokens_out":2269,"would_cite":true,"duration_ms":20643,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new benchmark of 1,680 adversarial audio samples tests six voice assistants and reports GPT-4o as the most resilient.","keywords":["audio adversarial attacks","large audio-language models","Chat-Audio Attacks benchmark","voice assistant robustness","GPT-4o","adversarial audio evaluation","content attack","implicit noise attack"],"falsifier":"Run the 1,680 CAA samples through the six models, then have blind human raters score the responses without knowing which model produced them and without letting GPT-4o score its own outputs, and check whether GPT-4o still leads on the coherence, correlation, and linguistic-robustness metrics. If its rank drops, the paper's central claim fails; a quicker check is whether GPT-4o's self-scores on its own outputs exceed human scores on the same outputs.","tokens_in":15353,"feed_emoji":"🔊","tokens_out":8955,"duration_ms":78828,"temperature":0.7,"pith_summary":"The paper sets out to test whether large audio-language models can keep holding a conversation when their audio input is attacked. It builds a benchmark, CAA, of 1,680 audio samples: for each of 360 source utterances it creates no-attack and attacked variants spanning content edits, emotional mismatches, explicit background noise, and inaudible infrasound or ultrasound. Six voice-capable models are scored under three evaluation routes: standard text metrics, a GPT-4o-based conversational scorecard, and human ratings. The central conclusion is that GPT-4o is the most resilient of the six across all three routes, while SpeechGPT and SALMONN are the most fragile. If correct, that gives researchers a reusable stress test for voice assistants and points to training data scale rather than a specific audio architecture as the main source of resilience.","feed_headline":"Four audio attack types test voice AI and GPT-4o resists best","feed_subtitle":"The 1,680-sample benchmark ranks six voice models under content, emotion, and noise attacks.","key_machinery":"The carrying object is the CAA quadruplet $(a_i, t_i, a_i^{\\text{no\\_attack}}, A_i)$: an original utterance, its transcript, a clean agent-read audio recording, and a set of attack variants. Content attacks are generated by GPT-4-guided synonym substitution, token rearrangement, or minimal token variation; emotional attacks re-synthesize the transcript with an opposite emotion or overlay opposite-mood background music; explicit noise attacks overlay natural, industrial, or human noise; implicit noise attacks overlay a 15 Hz infrasound or a 22 kHz ultrasound signal. Evaluation machinery then compares no-attack versus attacked responses with word error rate, ROUGE-L, and cosine similarity, plus GPT-4o-rated no-attack coherence, attack coherence, attack correlation, and linguistic robustness, and human ratings of coherence.","core_discovery":"The central claim is that conversational robustness of large audio-language models can be measured by a universal, content-preserving attack benchmark, and that under that benchmark GPT-4o clearly outperforms the other five evaluated models. The paper asserts that GPT-4o consistently delivers coherent, contextually relevant, and linguistically solid responses even under severe adversarial conditions, and attributes this to its extensive pre-training on large-scale datasets. The benchmark itself is the other part of the discovery: four attack families (content, emotional, explicit noise, and implicit noise) applied to re-synthesized clear speech provide a standardized way to compare vulnerabilities that previous model-specific targeted attacks did not offer.","pith_inferences":["An implication the authors leave implicit is that CAA only tests attacks that preserve human intelligibility and overall meaning, so it likely underestimates gradient-based targeted attacks tuned to a specific model; a high CAA rank does not guarantee resistance to those attacks.","A testable extension would be to add audio jailbreak samples once open-source audio jailbreak methods exist, since the paper notes it could not generate such samples and this remains an underexplored threat.","The finding that inaudible 15 Hz infrasound degrades several models more than 22 kHz ultrasound suggests a concrete mechanism check: comparing model input-embedding sensitivity to sub-20 Hz spectral energy would connect benchmark results to underlying audio encoders.","A practical deployment consequence is that voice assistants relying on speech-to-text front ends should expect adversarial audio to corrupt the transcription step, so robustness claims should be reported with per-attack transcription accuracy alongside final response quality."],"forward_implications":["Models that transcribe audio before responding, such as SpeechGPT, degrade more under all four attack families than models that process audio more directly.","Natural noise is the most damaging explicit noise category overall, and infrasound is more damaging than ultrasound for most of the evaluated models.","Resilience to emotional mismatch is not necessarily good news, because it coincides with low emotional awareness, which is a weakness for natural conversation.","Training on noisy audio and large-scale pre-training are the main observed correlates of resilience, pointing to data diversity as a defense strategy.","The released generation scripts let the benchmark grow beyond the current 1,680 samples, so the same attack families can be applied to new utterances and models."],"supporting_citations":[{"why":"Supplies the MELD conversational audio with emotion labels used for 120 of the 360 source utterances and for generating emotional attacks.","marker":"Poria et al., 2018"},{"why":"Supplies TVQA dialogues, source of 120 English conversational samples in the benchmark.","marker":"Lei et al., 2018"},{"why":"Supplies Common Voice utterances, source of the remaining 120 audio samples.","marker":"Ardila et al., 2019"},{"why":"The Azure Speech SDK re-synthesizes the no-attack and attacked audio, establishing the clean baseline and the emotional tone variants.","marker":"Microsoft, 2023"},{"why":"Establishes the gradient-based audio adversarial attack method that CAA's universal, conversation-level attack design is positioned against.","marker":"Carlini and Wagner, 2018"},{"why":"Provides the universal audio adversarial attack framework that motivates the benchmark's attack generation approach.","marker":"Xie et al., 2021"},{"why":"GPT-4 is used for transcript refinement, quality filtering, content-attack generation, and the GPT-4o-based scoring, so the benchmark pipeline and the headline evaluation both depend on it.","marker":"OpenAI, 2023"},{"why":"Identifies GPT-4o, the model the paper finds most resilient, and supplies the background claim that it handles audio-text interaction in noisy environments.","marker":"Achiam et al., 2023"}],"fun_headline_variants":["Chat-audio attacks: benchmark reveals GPT-4o is most robust","Audio attack benchmark: four threats, six models, GPT-4o wins","GPT-4o withstands audio attacks best in new benchmark","Benchmark tests voice AI under four audio attack types","Who survives chat-audio attacks? GPT-4o leads, says benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking depends on GPT-4o being an impartial judge of all six models, including itself; if its self-scores are inflated, the headline result weakens.","fun_headline_variants_meta":{"raw":{"variants":["Chat-audio attacks: benchmark reveals GPT-4o is most robust","Audio attack benchmark: four threats, six models, GPT-4o wins","GPT-4o withstands audio attacks best in new benchmark","Benchmark tests voice AI under four audio attack types","Who survives chat-audio attacks? GPT-4o leads, says benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000473,"raw_usage":{"total_tokens":2331,"prompt_tokens":905,"completion_tokens":1426,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":1348}},"tokens_in":521,"tokens_out":1426,"duration_ms":11264,"temperature":1.0,"reasoning_tokens":1348,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:48:23.219387+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the 1,680 CAA samples through the six models, then have blind human raters score the responses without knowing which model produced them and without letting GPT-4o score its own outputs, and check whether GPT-4o still leads on the coherence, correlation, and linguistic-robustness metrics. If its rank drops, the paper's central claim fails; a quicker check is whether GPT-4o's self-scores on its own outputs exceed human scores on the same outputs.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Azure Speech SDK re-synthesizes the no-attack and attacked audio, establishing the clean baseline and the emotional tone variants."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the gradient-based audio adversarial attack method that CAA's universal, conversation-level attack design is positioned against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the universal audio adversarial attack framework that motivates the benchmark's attack generation approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"GPT-4 is used for transcript refinement, quality filtering, content-attack generation, and the GPT-4o-based scoring, so the benchmark pipeline and the headline evaluation both depend on it."}],"review_version":1}