{"id":"d3eebe88-572f-4f2e-9c12-a72efd49ba26","arxiv_id":"2606.27701","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A novel simulation framework for acoustic adversarial attacks shows up to 94.5% relative WER increase on Whisper and wav2vec and defines a Dual-Form SNR to separate stealth from efficacy.","lead":"The paper introduces a high-throughput simulation framework for over-the-air acoustic attacks on voice control systems and tests over 8 million scenarios. Smart generalists should read it to better understand the real-world risks to AI voice interfaces that current digital-only attack models ignore.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Simulation fidelity to real over-the-air propagation not quantitatively validated against physical measurements","rationale":"The reader's weakest_assumption directly identifies the same internal dependency; no stronger or orthogonal concern (e.g., statistical reporting, baseline definition, or metric circularity) is visible from the abstract-level description of the central claim.","tokens_in":1671,"tokens_out":291,"duration_ms":14738,"concrete_test":"Take the 50 highest-gain adversarial examples reported in the simulation results; re-record them over-the-air in the paper's physical test room using the same source/receiver geometry and hardware; compute Whisper WER on the physical recordings and measure rank correlation with the simulated WER values. Correlation below 0.6 falsifies the fidelity assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline WER increases (up to 94.5%) and Dual-Form SNR rest on the claim that the high-throughput simulator captures the acoustic factors (room impulse responses, geometry, microphone directivity, etc.) that actually govern attack transfer. The abstract states real-world testing was performed, yet if the 8 M evaluations are generated inside the simulator without a reported side-by-side comparison (e.g., Pearson correlation or rank agreement between simulated and measured WER on the same adversarial waveforms), the quantitative risk numbers remain unanchored. This is the precise condition the reader flagged.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a high-throughput simulation framework for over-the-air acoustic adversarial attacks on ASR systems (Whisper, wav2vec). It reports results from over 8 million evaluations showing that acoustic awareness produces relative WER increases of up to 94.5%, and introduces a Dual-Form SNR metric to separate source stealth from attack efficacy. The work combines real-world testing, conceptual discussion, and large-scale simulation to argue that prior digital-only workflows abstract away critical acoustic factors.","tokens_in":1773,"tokens_out":500,"duration_ms":17390,"significance":"If the simulator's fidelity holds, the scale of the evaluation and the new SNR formulation would provide a concrete advance for repeatable physical-world attack assessment, moving beyond the abstractions criticized in the introduction. The empirical WER numbers and the decoupling metric are the primary contributions.","major_comments":[{"comment":"Validation section (or equivalent methods/results subsection describing the 8 M evaluations): the headline relative WER increases (up to 94.5 %) and the Dual-Form SNR rest on the claim that the simulator faithfully reproduces the acoustic factors (room impulse responses, geometry, microphone directivity) that govern real transfer. The abstract states real-world testing occurred, yet no quantitative side-by-side comparison (Pearson correlation, rank agreement, or error bars on matched simulated vs. measured WER for the same waveforms) is reported. This is load-bearing for the central risk claims.","section":"Validation / Results"},{"comment":"Definition of Dual-Form SNR (section introducing the metric): the operationalization must be shown to be independent of the simulation parameters that already encode attack success; otherwise the claimed decoupling of stealth from efficacy is tautological by construction. A concrete formula or pseudocode and an ablation on its sensitivity to room/geometry parameters would be required.","section":"Dual-Form SNR definition"}],"minor_comments":[{"comment":"Abstract: the sentence claiming 'real-world testing' should quantify how many physical trials were run and whether they were used for simulator calibration or only for qualitative illustration.","section":"Abstract"},{"comment":"Figure captions and axis labels for any WER-vs-SNR plots should explicitly state whether the plotted points are simulated or measured.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed report. The two major comments identify important areas for strengthening the validation and metric presentation. We respond point-by-point below and are prepared to revise the manuscript accordingly.","responses":[{"response":"We agree that a quantitative side-by-side validation is essential to support the simulator's fidelity claims. While the manuscript references real-world testing to motivate the framework, it does not include the requested matched comparisons (Pearson correlation, rank agreement, or error bars). In revision we will add a dedicated validation subsection reporting these metrics on the real-world waveforms we collected, thereby directly addressing the load-bearing concern for the reported WER increases.","revision_made":"yes","referee_comment":"[Validation / Results] Validation section (or equivalent methods/results subsection describing the 8 M evaluations): the headline relative WER increases (up to 94.5 %) and the Dual-Form SNR rest on the claim that the simulator faithfully reproduces the acoustic factors (room impulse responses, geometry, microphone directivity) that govern real transfer. The abstract states real-world testing occurred, yet no quantitative side-by-side comparison (Pearson correlation, rank agreement, or error bars on matched simulated vs. measured WER for the same waveforms) is reported. This is load-bearing for the central risk claims."},{"response":"We accept the need for explicit operationalization. The Dual-Form SNR is constructed from two distinct signal formulations—one evaluated at the source location and one at the victim microphone—that are deliberately separated from the adversarial perturbation parameters used to compute WER. In the revised manuscript we will insert the full mathematical definition, accompanying pseudocode, and a sensitivity ablation across room and geometry parameters to demonstrate that the metric remains stable and non-tautological with respect to attack success.","revision_made":"yes","referee_comment":"[Dual-Form SNR definition] Definition of Dual-Form SNR (section introducing the metric): the operationalization must be shown to be independent of the simulation parameters that already encode attack success; otherwise the claimed decoupling of stealth from efficacy is tautological by construction. A concrete formula or pseudocode and an ablation on its sensitivity to room/geometry parameters would be required."}],"tokens_in":1359,"tokens_out":476,"duration_ms":15854,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper builds a high-throughput simulator for over-the-air acoustic attacks on speech models and introduces a Dual-Form SNR to separate source stealth from victim success. They report running 8 million evaluations and relative WER increases up to 94.5% for Whisper and wav2vec when acoustics are included.\n\nWhat stands out is the attempt to move past purely digital adversarial examples. Most prior work abstracts away room effects, geometry, and microphone response, and this work tries to put those back in at scale. The Dual-Form SNR looks like a reasonable way to fix how current papers measure attacks. They also mention real-world testing alongside the simulation.\n\nThe soft spot is validation. The headline numbers rest on the simulator capturing the acoustic factors that actually matter in physical settings. If the 8 million runs are generated inside the model without a reported quantitative check—such as correlation or rank agreement between simulated and measured WER on the same waveforms—then the risk estimates stay somewhat unanchored. The abstract notes real-world tests occurred, but the strength of the claims depends on how closely those match the simulation outputs.\n\nThis is for researchers working on audio adversarial ML or voice-interface security. Readers who care about physical deployment risks will get value from the framing and the new metric. The paper shows honest engagement with the limitations of existing methods, so it deserves a serious referee to examine the simulation details and any side-by-side validation data.","headline":"The simulation framework and Dual-Form SNR metric address a gap in acoustic adversarial work, but the results hinge on unverified simulator fidelity to real propagation.","tokens_in":2232,"tokens_out":372,"would_cite":false,"duration_ms":18954,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A simulation framework for over-the-air acoustic attacks shows that acoustic awareness raises word error rates by up to 94.5% in models like Whisper.","keywords":["over-the-air acoustic attacks","voice control systems","simulation framework","word error rate","signal to noise ratio","adversarial attacks","speech recognition"],"falsifier":"Running the identical set of adversarial examples through physical loudspeakers and microphones in controlled rooms and measuring whether the observed word error rates match the simulated values would settle whether the framework's predictions hold.","tokens_in":2591,"feed_emoji":"🔊","tokens_out":644,"duration_ms":22445,"temperature":0.7,"pith_summary":"The paper builds a high-throughput simulation framework to study acoustic attacks on voice-controlled AI systems that accounts for physical sound propagation and room geometry. Testing more than eight million adversarial examples reveals that including these acoustic factors produces relative word error rate increases of up to 94.5 percent compared with purely digital approaches. The same framework introduces a Dual-Form Signal to Noise Ratio that separates how well an attack hides from its source from how effectively it disrupts the victim system. This setup removes the need to abstract away acoustic variables that earlier studies had set aside.","feed_headline":"Acoustic simulation raises voice AI attack success by 94.5%","feed_subtitle":"Large-scale testing of over-the-air attacks shows physical acoustics sharply increase word error rates in speech recognition models.","key_machinery":"The high-throughput reality simulation framework that incorporates acoustic geometry and detectability factors, together with the Dual-Form Signal to Noise Ratio metric.","core_discovery":"By running over eight million evaluations inside a novel high-throughput reality simulation framework, the authors show that acoustic awareness in over-the-air attacks yields relative Word Error Rate increases of up to 94.5 percent under Whisper and wav2vec, while the framework operationalizes a Dual-Form Signal to Noise Ratio that decouples source stealth from victim attack efficacy.","pith_inferences":["Defenses for voice AI may need to incorporate countermeasures that account for how sound travels through specific room shapes and surfaces.","The simulation method could be applied to other audio interfaces such as smart-home devices or automotive voice systems to assess similar risks.","Validation experiments that compare simulated results against matched physical recordings would clarify how much the framework can be trusted for real deployments."],"forward_implications":["Voice recognition systems face substantially higher risk from physical attacks once room acoustics and geometry are included in evaluations.","The Dual-Form SNR metric allows separate measurement of attack stealth and attack success, removing a prior limitation in the field.","Researchers can now conduct repeatable, large-scale tests of acoustic attacks without the logistical barriers of physical experiments.","Prior risk assessments that ignored acoustic variables systematically underestimate the effectiveness of over-the-air attacks."],"fun_headline_variants":["Acoustic sim shows 94.5% WER rise in voice AI attacks","8M evals link acoustics to 94.5% voice AI WER increase","Reality sim exposes 94.5% relative WER gains from attacks","High-scale sim finds physical acoustics raise WER 94.5%"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The simulation framework accurately captures the key acoustic factors affecting detectability and geometry influence in real-world over-the-air attacks.","fun_headline_variants_meta":{"raw":{"variants":["Acoustic sim shows 94.5% WER rise in voice AI attacks","8M evals link acoustics to 94.5% voice AI WER increase","Reality sim exposes 94.5% relative WER gains from attacks","High-scale sim finds physical acoustics raise WER 94.5%"]},"model":"grok-4.3","cost_usd":0.006587,"raw_usage":{"total_tokens":3051,"prompt_tokens":617,"num_sources_used":0,"completion_tokens":81,"cost_in_usd_ticks":65874500,"prompt_tokens_details":{"text_tokens":617,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2353,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":617,"tokens_out":81,"duration_ms":17471,"temperature":1.0,"reasoning_tokens":2353,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-02T21:25:16.513140+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the identical set of adversarial examples through physical loudspeakers and microphones in controlled rooms and measuring whether the observed word error rates match the simulated values would settle whether the framework's predictions hold.","supporting_citations":[],"review_version":2}