{"id":"d0a618cd-85c2-4910-afd8-31f0d0f83d66","arxiv_id":"2412.12217","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review summarizing adversarial example attacks and defenses in cybersecurity, with duplicated references and no new experimental results.","lead":"This paper is a survey of adversarial example attacks and defenses across cybersecurity applications such as malware detection, intrusion detection, and speaker verification. It compiles attack success rates and defense results from many prior studies but adds no new experiments.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central claim is directionally safe, but its evidence chain is weaker than the 71-reference apparatus suggests: duplicate citations inflate the corpus and no verification of quoted success rates is reported, so the 'comprehensive' conclusion is conditional on an unproven…","rationale":"Reader's CONDITIONAL verdict is defensible. The most load-bearing issue is not whether adversarial examples can degrade DL security classifiers—that is independently established and not in dispute—but whether this survey supplies a reliable evidence chain for that conclusion. Its only synthetic contribution is the recounting of numbers from the literature, and that recounting has no verification trail and is inflated by extensive duplicate citations. I therefore isolate deduplication and numerical verification as the single check that would settle whether the survey's evidence base is adequate. If the check passes, the CONDITIONAL can be relaxed to ACCEPT for the narrow claim; if it fails, the conclusion needs re-scoping or the survey needs correction. I do not see an internal inconsistency that would justify REJECT, and the central claim is consensus-consistent. Since my concern does not move the verdict away from the reader's CONDITIONAL, I mark the recommendation as UNCHANGED.","tokens_in":22038,"tokens_out":8170,"duration_ms":73258,"concrete_test":"Perform a deduplication-and-verification pass: merge references by DOI/arXiv ID, then rebuild the per-domain evidence table of Sections II–VI. If any of the five application domains is supported by fewer than three unique independent studies, or if spot-checking two headline numbers per domain against the original papers fails to reproduce them (e.g., [43] 98.9%, [47] AUC 0.9710→0.9319, [55] 35.7% ASR, [57] 94.31%), the survey's evidence chain is materially weaker than the 71-reference count implies and the conclusion should be re-scoped accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's conclusion—'a literature review underscores the effectiveness of attacks utilizing adversarial examples against ML/DL-based security systems'—is supported only by the survey's own recounting of quantitative results (e.g., 85% misclassification in [41], 63–69% in [42], 98.9%/96.5% in [43], 35.7% ASR in [55], 94.31% ASR in [57]). Three conditions are load-bearing: the figures are quoted faithfully, each citation is a distinct study, and each application domain is supported by independent experiments rather than repeated mentions of the same work. The second condition demonstrably fails. The reference list contains many duplicates, including [23]=[41], [24]=[45], [25]=[43], [26]=[47], [27]=[48], [28]=[6], [29]=[7], [30]=[8], [31]=[54], [33]=[60], [34]=[59], [35]=[58], [36]=[61], [38]=[65], [39]=[64], and [40]=[63]. After deduplication, the 71-reference corpus is substantially smaller, and the 'comprehensive' scope (Sections II–VI) is not established. The first condition is also unverified: the survey provides no audit trail for any of its headline numbers, so a misquoted figure would propagate directly into the conclusion. This is a correctness/completeness risk for the survey argument, not a refutation of the underlying claim, which has independent support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a survey of adversarial example (AE) attacks on deep learning (DL) systems operating in five cybersecurity domains: malware detection, botnet/DGA detection, network intrusion detection, user authentication, and encrypted traffic analysis. For each domain it recounts a set of selected attack and defense studies, quoting quantitative results such as misclassification rates, AUC values, and attack success rates, and it closes with a short countermeasures section and a conclusion arguing that AEs can significantly degrade the performance of ML/DL-based security systems. The abstract promises a comprehensive review covering attack generation methods, domain-specific impacts, trade-offs, and defense mechanisms.","tokens_in":22353,"tokens_out":9198,"duration_ms":75107,"significance":"The topic is important and timely: the resilience of DL-based security tools under adversarial perturbation is a central question in applied machine learning security. The manuscript correctly identifies recurring themes, such as the functionality-preservation constraint on malware perturbation and the black-box nature of many network attacks, and it brings together studies from several application areas within one narrative. The broad qualitative conclusion that adversarial examples can degrade DL-based security classifiers has independent support in the literature. However, the paper's evidentiary value is currently limited by extensive duplicate citations, a missing survey methodology, and a lack of any verifiable audit trail for the many quantitative results it reports. These problems affect the central claim of comprehensiveness and must be corrected for the survey to be citable.","major_comments":[{"comment":"The survey presents the same studies as multiple independent references. For example, [41] and [23] are the same paper (Grosse et al., arXiv:1606.04435), [42] is the ESORICS 2017 version of the same work, [26]=[47] (DeepDGA), [27]=[48] (CharBot), [28]=[6], [29]=[7], [30]=[8], [31]=[54], [32]=[55], [33]=[60], [34]=[59], [35]=[58], [36]=[61], [38]=[65], [39]=[64], and [40]=[63]. Section III first mentions [26] and [27], then devotes individual paragraphs to [47] and [48], which are the same two papers, and Section IV repeats this pattern for [30]/[56] and [29]/[57]. This duplication inflates the apparent evidence base and makes the 'comprehensive' claim unverifiable. The reference list must be deduplicated, the text renumbered, and the number of distinct studies stated explicitly.","section":"Reference list; Sections II–VI"},{"comment":"The paper does not state a survey methodology. There is no search strategy, database list, inclusion or exclusion criteria, time window, or deduplication procedure, and the conclusion's claim of comprehensiveness is therefore unsupported. Section I only announces the paper structure, while Section VIII asserts that the review is comprehensive without defining the corpus. A survey claiming to be 'comprehensive' must either specify how its references were collected and selected or temper the claim. Please add a methodology subsection and a limitations paragraph.","section":"Section I (Introduction) and Section VIII (Conclusion)"},{"comment":"References [15] and [71], cited together at the end of the first paragraph of Section I to support the claim that deep learning 'enables and facilitates many security-based applications,' are not about cybersecurity applications of deep learning. Reference [15] is titled 'Mitigating Challenges in Ethereum's Proof-of-Stake Consensus: Evaluating the Impact of EigenLayer and Lido' and reference [71] is 'Strengthening DeFi Security: A Static Analysis Approach to Flash Loan Vulnerabilities.' Neither supports the cited sentence. Reference [15] should be removed or replaced with a relevant citation, and the citation chain for the introduction should be checked for relevance.","section":"Section I, refs. [15], [71]"},{"comment":"The survey reports many exact quantitative results (e.g., 85% misclassification in [41], 63–69% in [42], 98.9%/96.5% in [43], AUC 0.9710 to 0.9319 in [47], 35.7% ASR in [55], 98.86% in [56], 94.31% in [57], 95% in [59], 98.43%/96.63% in [60]) without any audit trail: there is no summary table, no page or table numbers from the cited papers, and no indication of how each number was extracted. Because the conclusion aggregates these figures, a single misquotation would propagate directly into the survey's main claim. Please add a verification appendix or a table that maps each reported metric to its original source location, and state explicitly whether the numbers were checked against the originals.","section":"Sections II–VI"}],"minor_comments":[{"comment":"Section III's opening says 'modifying the AGD names' but should read 'modifying the DGA names,' and the paper-structure paragraph in Section I uses 'zombie networks' where 'botnets' is meant.","section":"Section III; Section I"},{"comment":"Section VII is titled 'Countermeasures' but covers only gradient masking, distillation, and feature compression; the abstract promises a discussion of adversarial training, which is absent from this section. Either add a subsection on adversarial training or revise the abstract.","section":"Section VII"},{"comment":"The paper would benefit from a comparative summary table (attack, domain, dataset, target model, metric, result, source), both to improve readability and to make the audit trail for reported numbers explicit.","section":"Overall"},{"comment":"Reference formatting is inconsistent: entry [46] lists 'Dl-fhmc' instead of 'DL-FHMC', entries [53] and [71] use nonstandard quotation marks, and several entries have mixed title capitalization. A careful copyedit of the bibliography is needed.","section":"References"},{"comment":"The paragraph on [44] mentions eight well-known off-the-shelf adversarial learning methods without naming them; naming these methods is necessary for the reader to assess the comparison.","section":"Section II"},{"comment":"The conclusion repeats generalities about stability, resilience, and security but does not state the limitations of the survey or concrete open questions; adding these would strengthen the paper's usefulness.","section":"Section VIII"}],"recommendation":"major_revision","confidential_remarks":"The reference duplications are extensive enough that a quick bibliometric check would have caught them; I would encourage the editor to require the authors to provide the deduplicated corpus and a list of which entries correspond to which studies before sending out a revised version. The unrelated self-citation [15] is a particular concern. The survey topic is suitable for the journal, but the current manuscript is not yet a reliable reference."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on Li Li's survey. The paper presents no new results—it's a review—so the bar is faithful, organized summary of known work. It does clear the low bar of readability: the five domain sections walk through each study's design and headline numbers, and someone new to adversarial ML in security could get a quick orientation. That's the extent of the credit.\n\nThe problems are in the evidence chain. The reference list is padded. The same papers appear twice under different numbers: Grosse et al. is [23] and [41], DeepDGA is [26] and [47], CharBot is [27] and [48], FGMD is [6] and [28], and there are at least a dozen more duplicates. That means the 71 references are really maybe 50, and the 'comprehensive' claim isn't supported. There's also no search strategy or inclusion criteria, so I have no way to assess whether the selection is representative. Every quantitative result quoted (85% misclassification, 98.9% for COPYCAT, 35.7% ASR for Tiki-Taka) is taken on faith; there's no verification trail. The self-citation [15] to an unrelated Ethereum paper is a red flag. The countermeasures section is three paragraphs—skeletal for a survey that claims to cover mitigation.\n\nThe central claim, that adversarial examples can degrade DL-based security tools, is uncontroversial and independently supported. So the conclusion isn't wrong. But the execution is too sloppy for this to serve as a reliable reference. A serious editor would want the deduplication done, the numbers checked, and a methodology paragraph added before sending it to reviewers. As submitted, I'd desk reject it.","headline":"A readable but sloppy survey whose citation inflation and missing methodology make it unreliable as a reference; the core claim is safe, the execution is not.","tokens_in":22817,"tokens_out":3342,"would_cite":false,"duration_ms":28673,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 71 studies concludes that adversarial examples reliably degrade deep-learning security systems across every major cybersecurity application the field has tested.","keywords":["adversarial examples","deep learning","cybersecurity","malware detection","intrusion detection","DGA detection","speaker verification","defense mechanisms"],"falsifier":"A systematic replication study that reruns the key reported experiments—such as COPYCAT on Windows malware, CharBot against FANCI, Tiki-Taka against CSE-CIC-IDS2018, and VMask against VGGVox—under the original conditions and finds materially lower attack success rates, or that shows production-grade security models resist the described perturbations, would weaken the survey's central claim.","tokens_in":21843,"feed_emoji":"🛡️","tokens_out":2896,"duration_ms":27415,"temperature":0.7,"pith_summary":"This survey argues that adversarial examples—small, carefully crafted perturbations to input data—are an effective and cross-domain threat to deep-learning-based security tools. By reviewing published attacks on malware classifiers, botnet and DGA detectors, intrusion detection systems, user authentication, and encrypted traffic analysis, the paper finds that such examples can force high misclassification rates and evade detection in each domain. The paper also catalogs defenses such as gradient masking, adversarial training, feature squeezing, and distillation, noting that they reduce but do not eliminate vulnerability. The intended takeaway is that adversarial robustness must be a primary design criterion for deploying deep learning in cybersecurity.","feed_headline":"Adversarial examples break deep-learning security tools","feed_subtitle":"A survey finds crafted inputs defeat malware, DGA, intrusion, and authentication systems; defenses only reduce the damage.","key_machinery":"The core mechanism is a domain-by-domain survey that organizes reported attack and defense results by cybersecurity application. For each domain it identifies the attack generation technique (forward derivatives, GAN-based generators, saliency maps, universal perturbations, query-based black-box methods) and the defense evaluated (adversarial training, distillation, feature squeezing, ensemble voting), then uses the cited paper's quantitative outcomes—such as misclassification rates, attack success rates, and AUC changes—as evidence that adversarial examples degrade model performance.","core_discovery":"The paper establishes, through a literature review, that adversarial examples cause significant performance deterioration in machine-learning and deep-learning-based security systems. Across the surveyed domains, attackers can craft inputs that evade detection with minimal or functionality-preserving modifications: malware classifiers are fooled at rates from 63% to 100%, DGA classifiers see recall plummet with only two character changes, intrusion detection systems are bypassed with less than 0.005% byte modification, and practical speaker verification is tricked by universal perturbations. The survey concludes that the destructive power of adversarial attacks on security models is real and that defenses like adversarial retraining improve resilience without making systems impervious.","pith_inferences":["The survey implies that adversarial robustness should be reported alongside accuracy for any security-related deep learning model, since evasion is a realistic operational risk rather than a theoretical artifact.","Many of the quantitative results are drawn from single papers with specific datasets and assumptions, so field effectiveness will vary; the pattern of vulnerability across domains is stronger evidence than any single number.","The repeated success of lightweight, black-box attacks suggests that attackers do not need white-box access or large compute budgets to threaten deployed security systems.","Defenses borrowed from computer vision, such as distillation and feature squeezing, need re-validation in security settings where perturbing inputs must preserve functional semantics like executability or protocol compliance."],"forward_implications":["Malware and IoT-malware classifiers can be evaded at high rates, up to 100% in graph-based IoT detectors, when attackers use feature-preserving perturbations.","Simple character-level modifications to domain names, as in CharBot and CLETer, can reduce state-of-the-art DGA classifier recall from near 99% to a few percent.","Intrusion detection systems are vulnerable to black-box attacks that require only small packet modifications, and universal perturbations threaten practical speaker verification systems.","Defenses such as ensemble adversarial training, query detection, and feature squeezing lower attack success rates but also introduce trade-offs like increased false positives or reduced clean accuracy.","Encrypted traffic classifiers and website fingerprinting defenses can be evaded or improved using adversarial traces, with bandwidth overhead depending on the technique."],"supporting_citations":[{"why":"Supplies the reported up-to-85% misclassification rate for a neural network malware classifier on DREBIN and evaluation of distillation and retraining defenses.","marker":"[41]"},{"why":"Supplies the COPYCAT results: 98.9% and 96.5% misclassification of Windows and IoT malware in visualization-based detectors.","marker":"[43]"},{"why":"Supplies the DeepDGA result that a random forest DGA detector's AUC dropped from 0.9710 to 0.9319 across adversarial rounds.","marker":"[47]"},{"why":"Supplies the CharBot result that a simple two-character modification evades DGA classifiers, with LSTM.MI detecting only about 55% of samples.","marker":"[48]"},{"why":"Supplies the Tiki-Taka result that time-based feature perturbations achieve up to 35.7% attack success rate against NIDS and that ensemble adversarial training reduces ASR.","marker":"[55]"},{"why":"Supplies the result that a black-box NIDS attack on Kitsune achieved 94.31% success with less than 0.005% packet byte modification.","marker":"[57]"},{"why":"Supplies the VMask voiceprint mimicry results against VGGVox and Microsoft Azure speaker verification, including a 95% grey-box success rate.","marker":"[59]"},{"why":"Supplies the feature squeezing defense mechanism that the survey presents as a general detection approach for adversarial examples.","marker":"[70]"}],"fun_headline_variants":["Adversarial inputs fool security AI up to 100% of the time","Crafted perturbations defeat malware and intrusion detectors","Security deep learning fails against tiny adversarial tweaks","Nearly all security models fall to adversarial examples"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions rest on the premise that the quantitative results reported in the cited papers are accurate and that the 71 included references fairly represent the field, since no systematic search or inclusion criteria are provided.","fun_headline_variants_meta":{"raw":{"variants":["Adversarial inputs fool security AI up to 100% of the time","Crafted perturbations defeat malware and intrusion detectors","Security deep learning fails against tiny adversarial tweaks","Nearly all security models fall to adversarial examples"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00045,"raw_usage":{"total_tokens":2214,"prompt_tokens":840,"completion_tokens":1374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":1311}},"tokens_in":456,"tokens_out":1374,"duration_ms":9201,"temperature":1.0,"reasoning_tokens":1311,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:00:00.398158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic replication study that reruns the key reported experiments—such as COPYCAT on Windows malware, CharBot against FANCI, Tiki-Taka against CSE-CIC-IDS2018, and VMask against VGGVox—under the original conditions and finds materially lower attack success rates, or that shows production-grade security models resist the described perturbations, would weaken the survey's central claim.","supporting_citations":[{"cited_title":"COPYCAT: Practical Adversarial Attacks on Visualization-Based Malware Detection","cited_arxiv_id":"1909.09735","evidence_quote":"Supplies the COPYCAT results: 98.9% and 96.5% misclassification of Windows and IoT malware in visualization-based detectors."},{"cited_title":"DeepDGA: A dversarially- tuned domain generation and detection,","cited_arxiv_id":null,"evidence_quote":"Supplies the DeepDGA result that a random forest DGA detector's AUC dropped from 0.9710 to 0.9319 across adversarial rounds."},{"cited_title":"CharBot: A simple and effective method for evading DGA classiﬁers,","cited_arxiv_id":null,"evidence_quote":"Supplies the CharBot result that a simple two-character modification evades DGA classifiers, with LSTM.MI detecting only about 55% of samples."},{"cited_title":"Tiki-taka: A ttacking and de- fending deep learning-based intrusion detection systems,","cited_arxiv_id":null,"evidence_quote":"Supplies the Tiki-Taka result that time-based feature perturbations achieve up to 35.7% attack success rate against NIDS and that ensemble adversarial training reduces ASR."},{"cited_title":"Adv ersarial attacks against network intrusion detection in IoT systems ,","cited_arxiv_id":null,"evidence_quote":"Supplies the result that a black-box NIDS attack on Kitsune achieved 94.31% success with less than 0.005% packet byte modification."},{"cited_title":"V o iceprint mimicry attack towards speaker veriﬁcation system in smart home,","cited_arxiv_id":null,"evidence_quote":"Supplies the VMask voiceprint mimicry results against VGGVox and Microsoft Azure speaker verification, including a 95% grey-box success rate."}],"review_version":1}