{"id":"3eab6b15-792d-48be-8ea5-1ebe88ca5def","arxiv_id":"2608.09158","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An inaudible 5-20 Hz intermittent waveform degrades performance across six audio-language models, and a spectral-clustering requery guard partially recovers it.","lead":"This paper shows that a fixed, low-frequency (5-20 Hz) audio waveform can cut the accuracy of six audio-language models by up to 67 percentage points while remaining nearly inaudible to human listeners. It also proposes a detection-and-requery defense that partially restores accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Physical-chain absence undermines both halves of the central claim: no real capture chain is shown to deliver beta=4 to the model while a human in the same room finds the emission inaudible.","rationale":"The paper honestly flags the physical-chain limitation in its Limitations section and Appendix A, and the reader already made the verdict CONDITIONAL on exactly this point. I agree this is the weakest, load-bearing assumption. I considered whether the attack's statistical comparison to Gaussian and PNL baselines is a separate concern: Table 13 shows ILL is not significantly better than the audible baselines on the 18 model-task cells after Holm correction, but the contribution is low-band stealth rather than uniform dominance, so I do not treat that as fatal. The cross-model transfer with a fixed waveform and the ablation showing that structured state-sequence construction adds effect are real internal evidence that the digital phenomenon is not a fluke of one model. The missing physical realization is therefore the only concern that can invalidate the abstract's central claim as stated. Since the paper explicitly restricts its claims and the reader's CONDITIONAL verdict already reflects that restriction, no verdict adjustment is needed; the concrete test should be run before the claim is upgraded to a demonstrated practical threat.","tokens_in":18181,"tokens_out":4615,"duration_ms":55438,"concrete_test":"Perform a controlled over-the-air experiment: place a calibrated low-frequency source (e.g., a rotary subwoofer as in Park and Robertson 2009) and a consumer smartphone or laptop microphone in a room; emit ILL while speech plays; adjust the source level until the post-ADC and post-codec signal fed to the six LALMs has the same 5-20 Hz RMS relative to speech as beta=4 in Appendix D; then measure task accuracy and, with human listeners physically present or with binaural calibrated recordings at the listener position, collect audibility ratings. If the accuracy degradation vanishes or the ratings significantly exceed the clean-condition baseline, the abstract's combined claim would need to be qualified as digital-only or rejected as a practical threat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract is that ILL is an inaudible input that degrades LALMs. For that claim to hold, two physical conditions must hold simultaneously. First, a real source in a real room must produce a received 5-20 Hz component whose post-ADC/analog-front-end/codec amplitude equals the beta=4 RMS used in Equation 14 (Appendix D). Appendix A cites sources showing low-frequency generation and microphone sensitivity, but it also lists AC coupling, high-pass filters, AGC, and codecs as threats; it never measures them. Therefore the attack-effectiveness half rests on an unverified assumption that the model-side waveform equals the digitally mixed x + delta. Second, inaudibility: human listeners heard the same digital composite x + delta through conventional headphones or loudspeakers (Appendix F). Consumer transducers typically roll off below 20 Hz, so a listener's 1.33 rating may reflect the playback transducer's inability to reproduce 5-20 Hz, not human insensitivity to physically present low-frequency pressure. No test plays the actual ILL emission in the same room where a human and the target microphone coexist. Thus the combined claim 'degrades by 67 pp while rated 1.33' is not physically established: one half assumes the chain preserves 5-20 Hz to the model, while the other half's playback likely removes it from the listener. If either half fails, ILL is either ineffective or audible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Intermittent Low-Frequency Lockout (ILL), a fixed universal waveform in the 5-20 Hz band, constructed offline using attention-based active-interval estimation and corpus-derived frequency-state transitions, and evaluates it as a black-box attack on six large audio-language models (LALMs) across speech recognition, translation, audio question answering, and emotion classification. It also proposes Distributional Requery Guard (DRG), a clustering-based detector that flags low-frequency distribution shifts and requests a second recording for semantic recovery, and evaluates this defense on the same tasks. The digital experiments report task accuracy reductions of up to 67 percentage points, an audible noise ratio of 0.06-0.08%, a mean human audibility rating of 1.33 (close to 1.17 for clean audio), and defense gains from 28.5% to 46.1% attacked accuracy after clean reacquisition.","tokens_in":18415,"tokens_out":9191,"duration_ms":90055,"significance":"The paper targets a plausible and underexplored input-surface risk: signals below the conventional audibility cutoff that may still reach LALM frontends and influence behavior. The digital evaluation is broad and internally consistent, with a common RMS budget across baselines, ablation studies, and internal attention/confidence analyses. The method is black-box in the sense that the waveform is fixed and constructed from a reference model without querying the target. If the physical channel were validated, the work would constitute a meaningful contribution to LALM safety. However, the central 'inaudible' and 'practical' claims currently rest on a simulated receiver-side evaluation and on human ratings obtained through playback transducers that likely attenuate the very band under study. This gap is load-bearing and must be addressed by physical validation or by a substantial reframing of the claims to a digital simulation study.","major_comments":[{"comment":"The abstract's central claim that ILL is an inaudible input posing a practical hidden channel requires that a real acoustic emission deliver a 5-20 Hz component with the modeled amplitude beta=4 to the target microphone while a human in the same room perceives it as inaudible. The manuscript states in the Limitations that the experiments 'simulate microphone reception and do not reproduce the complete acoustic path from a loudspeaker through a real-world acoustic environment to a microphone,' and Appendix A provides only literature support for feasibility rather than end-to-end measurements. Meanwhile, the human audibility evaluation in Appendix F played the digitally mixed composite x+delta through conventional headphones or loudspeakers, which typically have strong roll-off below 20 Hz; a mean rating of 1.33 may therefore reflect the transducer's inability to reproduce 5-20 Hz rather than human insensitivity to a physically present low-frequency pressure field. Thus the two halves of the combined claim are not jointly established: the attack-effectiveness half assumes the chain preserves the low-frequency component to the model, while the inaudibility half likely removes it from the listener. This needs to be fixed by physical experiments or by explicitly scaling back the claims to a simulated receiver-side effect.","section":"Threat Model; Limitations; Appendix A; Appendix F; Eq. (14)"},{"comment":"The audible noise ratio (ANR) is computed on the digital waveform and therefore says nothing about whether the acoustic field at the listener's ear contains perceptible low-frequency energy. The statement that 'the agreement between the human ratings and ANR supports the low perceptibility of ILL' is only valid for the playback chain used in the rating study, not for a physical deployment. Moreover, the human rating results are reported only as aggregate means and medians; no variance, confidence intervals, or per-condition statistics are given, so the difference between 1.33 and 1.17 is not shown to be statistically meaningful. Please report full distributions, confidence intervals, and ideally disaggregate by playback mode (headphones versus loudspeaker).","section":"Eq. (13); Fig. 3; Appendix F"},{"comment":"The common RMS budget sets RMS(delta)=4 for every attack, but the manuscript does not specify the RMS or amplitude normalization of the clean source audio x. If x is in a conventional sample range, a perturbation with RMS 4 can dominate the composite signal, which would make the model-side degradation unsurprising and would make the 'stealthy' interpretation of the human ratings depend entirely on playback roll-off. Please report the SNR (or the ratio of perturbation RMS to speech RMS) and verify that the reported attack effectiveness is not an artifact of a single arbitrary amplitude scale.","section":"Appendix D; Eq. (14)"},{"comment":"The Introduction claims that 'the same requery mechanism also recovers useful semantic evidence under other noise perturbations and attack methods,' but the DRG detector is trained and evaluated (Table 3) only on ILL-style low-frequency interference. For the four additional attack types in Figure 5, it is not specified whether DRG's detector actually flags those attacks and with what accuracy, or whether requery was forced for the comparison. Without this information, the transfer claim is not substantiated. Please clarify the experimental setup or restrict the claim to the ILL case.","section":"Fig. 5; Introduction; Table 3"}],"minor_comments":[{"comment":"The phrases 'specified in Section' and 'defined in Section' appear without section numbers (in the DRG description and the metrics paragraph); please add numbered references.","section":"Method; Experiments"},{"comment":"The description '112 complete human response sets' is ambiguous: report whether these are 112 distinct participants each rating all conditions, or 112 ratings per condition.","section":"Appendix F"},{"comment":"The abbreviations 'M&alpha', 'S&beta', 'Q&gamma', etc. are not defined in the captions; add a legend identifying M (MiniCPM), S (StepAudio), Q (Qwen3), and the attack symbols.","section":"Figures 5 and 6"},{"comment":"The audible noise ratio uses an upper cutoff of 8 kHz without explanation; either justify the cutoff or align it with the stated 20 kHz audibility boundary.","section":"Eq. (13)"},{"comment":"The color legend (gray/red/green) cannot be conveyed in a monochrome print; the parenthesized subscripts are sufficient, but the caption should state that the colors are only an aid in the electronic version.","section":"Table 1"},{"comment":"The mean corpus duration T-bar is defined but it is not immediately clear how it is used; please explicitly connect it to the definition of K=round(T_cyc/T-bar)+1.","section":"Method; Frequency Confusion Transfer"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid digital red-teaming study with careful attention to fair comparison, ablations, and internal analyses. The stress-test concern about the missing physical acoustic path lands: the 'inaudible' and 'practical' claims are not supported by the current evidence. A major revision that either adds a physical validation (even a limited one) or reframes the contribution as 'simulated receiver-side evaluation' would be necessary for acceptance. The human-subject reporting also needs statistical strengthening."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. ILL is a genuinely new attack construction: a fixed 5-20 Hz intermittent waveform, built once from a reference corpus and reference model, transferred to six LALMs with no per-target optimization. The accuracy drops are real in the digital domain, up to 67 pp on RAVDESS/StepAudio2, and the paper is careful to compare against Gaussian and PNL baselines under a common RMS budget (Appendix D), so the comparison is fair. The audibility claim is supported in the digital realm by ANR under 0.1% and a mean human rating of 1.33, though I'll get to the caveat. The ablation showing the structured state sequence beats random, fixed-frequency, and sweep waveforms, plus the attention and confidence analysis, make the mechanism plausible. DRG is refreshingly simple: spectral descriptor, K-means, conditional requery; detection F1 of 89-99% and recovery from 28.5% to 46.1% is measured, not hand-waved. Appendix E's paired stats are honest enough to say ILL is comparable to Gaussian and PNL, not uniformly stronger. The limitations section is candid about what was not tested.\n\nThe soft spots are the physical chain. The central advertised claim is that ILL is an inaudible input that degrades LALMs. That requires two physical facts to hold simultaneously: a real source-room-mic chain must preserve 5-20 Hz at the modeled beta=4 received amplitude, and a human in that same room must not hear the emission. The paper tests neither jointly. The human ratings were collected on digitally mixed composites played through consumer headphones or loudspeakers, which typically roll off below 20 Hz; the 1.33 rating may say more about the transducer than about human sensitivity to the physically emitted signal. Conversely, the model-side waveform is assumed from digital mixing, not measured through a real capture chain. Appendix A cites sources showing low-frequency generation and mic sensitivity, but also lists AC coupling, high-pass filters, AGC, and codecs as threats without measuring them. The stress-test note is right that this undermines the combined claim. That said, the authors explicitly acknowledge not reproducing the complete acoustic path, so this is an incomplete threat evaluation rather than a misleading one. The DRG detector is trained on the same ILL interference it must detect, making the detection F1 partly circular, though the recovery result is less so. Sample sizes are 100 per dataset and several free parameters (beta, gamma, n, alpha, W) are chosen via the same ablations; the conclusions are stable but not the last word.\n\nWho is this for? People working on audio robustness, red teaming, and LALM safety will get a useful, clearly described baseline attack and a defense idea worth testing. It deserves a serious referee, with the physical validation as the main requested revision.","headline":"A genuinely new low-frequency attack construction with solid digital-domain results, but the physical 'inaudible input' claim is not yet established and deserves a serious referee.","tokens_in":19007,"tokens_out":2541,"would_cite":true,"duration_ms":25108,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fixed 5–20 Hz waveform, inaudible to humans, degrades large audio-language models by up to 67 percentage points in accuracy.","keywords":["large audio-language models","low-frequency attack","inaudible interference","universal adversarial waveform","red teaming","audio safety","distributional requery","availability attack"],"falsifier":"Run a physical test in a typical office: emit the ILL waveform from a low-frequency-capable loudspeaker at a distance that yields a measured received amplitude near β=4 at a smartphone microphone, record a speech query, and compare the LALM's task accuracy to the clean baseline; if the low-frequency band is attenuated below the noise floor by the microphone's high-pass response or AGC, or if the model's accuracy does not drop substantially, the real-world transfer claimed by the paper would not hold.","tokens_in":17940,"feed_emoji":"🤫","tokens_out":9936,"duration_ms":77923,"temperature":0.7,"pith_summary":"The paper establishes that a fixed, universal waveform confined to the 5–20 Hz band—far below normal human hearing—can silently degrade the task performance of large audio-language models (LALMs) across multiple models and tasks. The proposed attack, Intermittent Low-Frequency Lockout (ILL), uses a single frozen template, constructed once from a reference corpus and reference model, that transfers without per-target optimization; it reduces accuracy by up to 67 percentage points on emotion classification while keeping audible noise below 0.1% of the waveform's energy and earning a mean human audibility rating near that of clean audio. The authors argue this reveals a hidden input channel that models perceive but users cannot, creating an availability and reliability risk. They also propose a lightweight defense, Distributional Requery Guard (DRG), which detects the low-frequency distribution shift and conditionally requests a second recording, recovering mean attacked accuracy from 28.5% to 46.1%. The core message is that model perception extends below the human hearing threshold, and safety evaluation must account for this mismatch.","feed_headline":"Inaudible 5–20 Hz wave drops audio AI accuracy 67 points","feed_subtitle":"A fixed universal waveform hits six LALMs, almost inaudible to humans; a requery guard partially restores accuracy.","key_machinery":"The machinery is the universal intermittent low-frequency waveform δ*(t): a fixed 5–20 Hz signal, silent to humans, that repeats an 'on' interval derived from Sentence Attention Scale Estimation (which finds durations of continuous semantic attention via boundary scores from the reference model's audio attention) and an 'off' interval determined by a duty ratio, and whose active segment is synthesized by Frequency Confusion Transfer (which quantizes corpus spectral centroids into states and decodes a most-probable state sequence with continuous phase, so the waveform's frequency trajectory mimics corpus spectral variation). The template is constructed once, remains fixed across test recordings and targets, and is meant to be emitted acoustically so the microphone captures the physical superposition x(t)+δ(t). Detection is carried out by DRG, a spectral K-means cluster on ℓ1-normalized frequency descriptors that flags low-frequency mass and requests a second recording.","core_discovery":"The paper's central claim is that low-frequency signals in the 5–20 Hz range, although inaudible to humans, can be used as a universal black-box attack against large audio-language models. The authors construct ILL, an intermittent waveform derived from two components: Sentence Attention Scale Estimation, which sets the on/off timing from multi-scale changes in a reference model's attention over a speech corpus, and Frequency Confusion Transfer, which converts the corpus's spectral variation into a phase-continuous low-frequency state sequence. The fixed template is emitted as a standalone acoustic signal that superposes with the user's speech at the microphone; in simulation, it reduces accuracy by up to 67 percentage points across six LALMs and four task types, while its spectral leakage above 20 Hz is 0.06–0.08% and human raters find it as inaudible as clean audio. The authors further claim that the disruption is accompanied by reduced attention to acoustic evidence and lower confidence, and that their Distributional Requery Guard detects the shift with F1 up to 99% and recovers useful semantics by requesting a second recording. The central discovery, as stated, is that a fixed inaudible waveform, independent of any test utterance or target model, constitutes a transferable and practically stealthy availability threat to LALMs.","pith_inferences":["Editorial inference: If the physical chain can deliver the modeled received amplitude β=4 at the microphone, the same ILL template could be tested over the air; the authors did not reproduce the complete acoustic path, so physical realism remains open.","Editorial inference: The success of a universal template suggests that similar below-perception side channels (e.g., ultrasonic or other sub-audible bands) might be exploited to perturb models that process those bands, though this is not tested here.","Editorial inference: DRG's principle—when input distribution shifts, ask for independent evidence—is a general availability-defense idea that could be applied to other input corruptions beyond low frequency, supported by the paper's transfer results on four other attack types.","Editorial inference: The strong effect of a fixed 5–20 Hz waveform on emotion classification (RAVDESS) suggests that prosodic and paralinguistic features are particularly sensitive to low-frequency contamination, a hypothesis the paper does not test directly."],"forward_implications":["LALMs that ingest raw audio inherit a security surface below the human hearing threshold, so any safety argument based on human audibility is insufficient.","A single universal waveform can disrupt multiple models without per-target optimization, lowering the barrier to launching an attack.","Objective spectral leakage and subjective audibility ratings must be reported together, since a low ANR does not by itself prove human imperceptibility.","Conditional reacquisition, as used by DRG, can recover task semantics in the presence of low-frequency interference, and the same requery helps for other noise types, suggesting a general defense principle.","The observed drop in audio attention mass and correct-answer probability under attack indicates that interference acts on internal evidence use, implying that monitoring these quantities might offer a detection signal."],"supporting_citations":[{"why":"Shows small electret condenser microphones retain nonzero response down to 0.5 Hz, supporting that the 5–20 Hz band can reach a model's frontend.","marker":"(Jeng, Yang, and Lee 2011)"},{"why":"Characterizes digital smartphone microphone response over 0.01–20 Hz in non-isolated conditions, strengthening physical plausibility of low-frequency capture.","marker":"(Asmar et al. 2018)"},{"why":"Demonstrates a compact rotary source can generate and propagate coherent infrasound over distance, supporting that the ILL band can exist as a controlled airborne signal.","marker":"(Park and Robertson 2009)"},{"why":"Provides the white-box adversarial audio baseline (Audio-Adv.) that ILL is compared against.","marker":"(Carlini and Wagner 2018)"},{"why":"Provides the universal acoustic attack baseline (Whisper) that transfers across models, the comparison most relevant to ILL's transferability.","marker":"(Raina et al. 2024)"},{"why":"Supplies the LibriSpeech corpus used to evaluate speech recognition degradation.","marker":"(Panayotov et al. 2015)"},{"why":"Supplies RAVDESS, the emotion dataset where the largest 67-point accuracy drop is observed.","marker":"(Livingstone and Russo 2018)"}],"fun_headline_variants":["Inaudible 5–20 Hz wave slashes audio AI accuracy 67 points","Low-frequency attack cuts audio model accuracy by 67 points","Silent 5-20 Hz waveform disrupts audio models, nearly inaudible","Inaudible universal waveform drops LALM accuracy 67 points"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The attack's real-world effect depends on the 5–20 Hz component surviving the full acoustic-to-digital chain—source, room, microphone, high-pass filters, AGC, and codec—at a received amplitude near the simulated β=4, a chain the authors simulated only as digital mixing and did not reproduce in the physical world.","fun_headline_variants_meta":{"raw":{"variants":["Inaudible 5–20 Hz wave slashes audio AI accuracy 67 points","Low-frequency attack cuts audio model accuracy by 67 points","Silent 5-20 Hz waveform disrupts audio models, nearly inaudible","Inaudible universal waveform drops LALM accuracy 67 points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000936,"raw_usage":{"total_tokens":4059,"prompt_tokens":1056,"completion_tokens":3003,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":672,"completion_tokens_details":{"reasoning_tokens":2922}},"tokens_in":672,"tokens_out":3003,"duration_ms":22279,"temperature":1.0,"reasoning_tokens":2922,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:23:42.710370+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a physical test in a typical office: emit the ILL waveform from a low-frequency-capable loudspeaker at a distance that yields a measured received amplitude near β=4 at a smartphone microphone, record a speech query, and compare the LALM's task accuracy to the clean baseline; if the low-frequency band is attenuated below the noise floor by the microphone's high-pass response or AGC, or if the model's accuracy does not drop substantially, the real-world transfer claimed by the paper would not hold.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows small electret condenser microphones retain nonzero response down to 0.5 Hz, supporting that the 5–20 Hz band can reach a model's frontend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Characterizes digital smartphone microphone response over 0.01–20 Hz in non-isolated conditions, strengthening physical plausibility of low-frequency capture."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates a compact rotary source can generate and propagate coherent infrasound over distance, supporting that the ILL band can exist as a controlled airborne signal."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the white-box adversarial audio baseline (Audio-Adv.) that ILL is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the universal acoustic attack baseline (Whisper) that transfers across models, the comparison most relevant to ILL's transferability."}],"review_version":1}