{"id":"d5544d73-3359-4d2c-b55c-20a2a98afdf2","arxiv_id":"2602.08690","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 66 DRL-for-cybersecurity papers, the authors identify 11 recurring methodological pitfalls—averaging 5.8 per paper—and demonstrate their impact in four environments.","lead":"This paper maps 11 common methodological failures in deep-reinforcement-learning cybersecurity research, finds them in most of the 66 papers it reviewed, and shows in controlled experiments how they inflate or distort results. A generalist should read it because it explains why many published DRL security agents may not work when deployed.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prevalence estimates rest on a self-derived, unvalidated taxonomy with only moderate per-pitfall coding reliability; holdout re-coding is needed before '5.8 pitfalls per paper' can be treated as a literature-level fact.","rationale":"The reader's weakest assumption identifies coding reliability as the main vulnerability, and I agree: the quantitative central claim is exactly the part that requires stable, valid measurement. The paper's own reported kappa values and the fact that the taxonomy was derived from a pilot subset of the same corpus make this a real, load-bearing concern rather than a speculative one. The qualitative taxonomy and the recommendation sections are valuable regardless, and the case studies provide suggestive evidence of impact, but the advertised prevalence numbers ('average of 5.8', 'every paper contains at least two pitfalls', per-pitfall ranges) are the paper's headline contribution and are not yet independently validated. A CONDITIONAL verdict is appropriate: accept the paper's framing and contributions, but require either external holdout re-coding or a clear downgrade of the prevalence claims to corpus-specific observations. This does not undermine the overall SoK contribution, which is why I would not move to REJECT; it merely prevents the quantitative claims from being over-interpreted as stable literature-level statistics.","tokens_in":29363,"tokens_out":6083,"duration_ms":77651,"concrete_test":"Take a stratified random sample of 20 papers from the existing 66-paper corpus and an additional 20 DRL4Sec papers published 2018–2025 that meet the same inclusion criteria but were not used to develop the taxonomy. Recruit two external coders who are blind to the authors' classifications; train them only on the published pitfall definitions (Sections 5–8, Appendix A) and have them independently code all 40 papers. Compute per-pitfall Cohen's kappa and compare the external per-paper pitfall counts and average against the reported 5.8 using a two-sample mean test. If per-pitfall kappa for M-MS, M-PO, E-GA, or D-UA falls below 0.6, or if the external average is more than one pitfall lower, or if the fraction of papers with at least two pitfalls drops below 95%, the prevalence claims should be explicitly reworded as corpus-specific rather than literature-wide.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The two central quantitative claims — 'every paper contains at least two pitfalls' and 'average of 5.8 pitfalls per paper' — depend entirely on the pitfall coding of the 66-paper corpus. Section 3 states that the taxonomy itself was produced from a pilot review of the 37 April-collection papers and then applied to the full 66-paper corpus. The measurement instrument is therefore fit to part of the same data it is used to describe, and no external validation or pre-registration is reported. Overall Cohen's kappa is 0.712, but Table 4 shows only moderate agreement on several constituent pitfalls: M-MS 0.579, M-PO 0.583, E-GA 0.594, D-UA 0.570. The headline counts combine 'present' and 'partially present' into a single 'pitfall' category, so even modest coding instability in these moderate-agreement categories can materially shift per-paper totals and the reported average. Additionally, several papers in the corpus are authored by members of the review team; without a conflict-of-interest analysis or a stratified comparison of author-affiliated versus non-affiliated papers, systematic leniency or severity bias cannot be ruled out. The qualitative finding — that convergence evidence is frequently missing, variance is often unreported, and partial observability is widely unaddressed — is robust. But the precise prevalence statistics that anchor the abstract and conclusion are not yet established as stable, generalizable facts about DRL4Sec literature.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This SoK paper identifies 11 methodological pitfalls in applying deep reinforcement learning to cybersecurity, organized across modeling, training, evaluation, and deployment. It reports a systematic review of 66 papers (2018–2025), claiming that every paper exhibits at least two pitfalls, with an average of 5.8 per paper, and that the most common pitfall (policy convergence) appears in 71.2% of papers. The paper further demonstrates the practical impact of the pitfalls through controlled experiments in autonomous cyber defense (MiniCAGE), adversarial malware creation (AutoRobust), and web security testing (Link and SQiRL), and provides recommendations for each pitfall. Appendices include search details, reviewer agreement, and MDP/hyperparameter specifications.","tokens_in":29716,"tokens_out":7087,"duration_ms":82371,"significance":"If the prevalence estimates are reliable, this is a valuable contribution to a growing and methodologically uneven literature. The qualitative findings — that convergence evidence is frequently missing, variance is often unreported, partial observability is widely unaddressed, and environments are often simplified — are credible and align with broader reproducibility concerns in DRL. The paper's strengths include a two-reviewer protocol with reported agreement, case studies using 20 runs with confidence intervals, and detailed appendices with MDP specifications and hyperparameters. However, the central quantitative claims rest on a taxonomy developed from a pilot subset of the same corpus and on coding categories with only moderate inter-rater reliability on several pitfalls. These issues need to be addressed before the headline prevalence numbers can be treated as established facts about the DRL4Sec literature.","major_comments":[{"comment":"The 11-pitfall taxonomy was induced from a pilot review of the 37 April-collection papers and then applied to the full 66-paper corpus that includes those same 37 papers. This is a circular measurement design: the taxonomy is fit to part of the data it is used to describe. No external validation, pre-registration, or holdout analysis is reported. Since the December 2025 collection (29 papers) was not used in taxonomy development, the paper should report whether the 'every paper contains at least two pitfalls' and 'average 5.8' statistics replicate in that held-out subset as a validation check. Without such evidence, the abstract's prevalence claims are not yet established as stable, generalizable facts.","section":"Section 3 (Review Methodology)"},{"comment":"Several constituent pitfalls have only moderate inter-rater reliability (M-MS κ=0.579, M-PO κ=0.583, E-GA κ=0.594, D-UA κ=0.570). The headline counts combine 'present' and 'partially present' into a single 'pitfall' category — for example, 71.2% for policy convergence is 42.4% present plus 28.8% partially present, and 28.8% for underlying assumptions is 15.2% plus 13.6%. Moderate coding stability in these categories can materially shift per-paper totals and the reported average of 5.8. The paper should report prevalence under a stricter 'present-only' threshold, present the distribution of per-paper pitfall counts, and discuss how sensitive the average is to plausible coding error.","section":"Section 3, Table 4"},{"comment":"A number of papers in the 66-paper corpus are authored or co-authored by members of the review team (e.g., refs. [7, 15, 48, 51, 88, 91, 131, 132, 133]). The manuscript does not disclose this or provide a stratified comparison of author-affiliated versus non-affiliated papers. This is a potential source of systematic leniency or severity bias in coding. The paper should disclose the overlap and, at minimum, report a sensitivity analysis with those papers excluded or a comparison of their prevalence scores against the rest of the corpus.","section":"Section 3 / Corpus"},{"comment":"The 'Distinct States' versus 'Original' Link comparison is used to argue that removing modeling pitfalls 'significantly' improves performance. However, the 95% confidence intervals overlap substantially (Original 75.2 [59.4, 90.9]; Distinct 85.3 [76.7, 94.0]), and no significance test or paired comparison is reported. The 10.1 percentage-point increase is not statistically supported by the reported data. The authors should either provide a paired bootstrap or other appropriate test, or soften the claim to a suggestive result. The same issue appears in some MiniCAGE and SQiRL comparisons where confidence intervals overlap.","section":"Section 5.4, Table 1"}],"minor_comments":[{"comment":"The text reports an overall Cohen's kappa of 0.712, but Table 4 does not include the overall kappa. Adding it would make the summary agreement easier to verify.","section":"Section 3, Table 4"},{"comment":"The spelling of 'SQiRL' is inconsistent: both 'SQIRL' and 'SQiRL' appear. Also 'W A VSEP' should be 'WAVSEP'.","section":"Throughout"},{"comment":"The stacked bar chart is dense; the three severity categories are shown only via a legend, and the individual pitfall labels are small. Consider adding the exact percentages and counts in a companion table, or using a labeled dot plot.","section":"Figure 4"},{"comment":"The text states that removing 'unnecessary partial observability' avoids M-MC and M-PO. Since the original environment is described as having valid Markovian transitions, the issue is better characterized as a conflated observation function causing partial observability, not a direct Markov violation. Clarify this distinction.","section":"Section 5.4"},{"comment":"The recommendation to use N≥5 while also citing N≥20 and N≥50 for robust confidence intervals may appear contradictory. Clarify that N≥5 is a pragmatic minimum, not a substitute for a power analysis.","section":"Section 6.4"},{"comment":"Several cross-condition comparisons in the MiniCAGE deployment table have overlapping confidence intervals (e.g., Mixed training vs B-line evaluation: -40.2 [-47.1, -33.4] vs -19.0 [-20.9, -17.2] actually do not overlap, but others do). Reporting paired significance tests would strengthen the claims about degradation under changed assumptions.","section":"Section 8.3, Table 3"},{"comment":"The 'Security Implications' paragraphs are somewhat repetitive across pitfalls; tightening them would improve readability and make the taxonomy easier to scan.","section":"Sections 5–8"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be influential: it addresses a real and widely felt problem, and the qualitative findings are credible. However, the quantitative prevalence claims are load-bearing and currently rest on a taxonomy developed from the same corpus it is applied to, with moderate reliability on several categories and no disclosure of author overlap. A major revision that adds a holdout or sensitivity analysis, present-only prevalence, a per-paper pitfall count distribution, conflict-of-interest handling, and stronger case-study statistics would substantially increase my confidence. I do not think rejection is warranted, because the issues appear addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first SoK that treats DRL-for-cybersecurity pitfalls as a distinct category, and it earns its place. The 11-pitfall taxonomy is well organized across modeling, training, evaluation, and deployment; the prevalence analysis of 66 papers is a real dataset; and the case studies in MiniCAGE, AutoRobust, Link, and SQIRL actually demonstrate the pitfalls with 20 runs and confidence intervals, not just assert them. Appendix B gives MDP specifications and hyperparameters, which is more than most papers in this area bother with. The recommendations are concrete. Credit where due.\n\nI'm less worried about the stress-test than the note suggests, though not dismissive. Deriving the taxonomy from a pilot subset and then applying it to the full corpus is standard qualitative practice, and the two-reviewer protocol with conservative disagreement resolution is a reasonable floor. But the headline numbers—5.8 pitfalls per paper, 71.2% lacking convergence—are only as stable as the coding, and Table 4 shows moderate agreement on some constituent pitfalls (M-MS, M-PO, E-GA, D-UA). Since present and partially present are collapsed in the counts, those moderate categories can shift the average. That doesn't undermine the qualitative story: convergence evidence is often missing, variance is frequently unreported, partial observability is widely unaddressed. It does mean the precise prevalence statistics should be read with a margin, and the authors would strengthen the paper by reporting a sensitivity analysis and making the coding sheets public.\n\nThe author-overlap issue is real. Several corpus papers are by the review team; that is not disqualifying in a young subfield, but without a stratified check it is hard to rule out systematic leniency. I would ask for that in revision. I would also like the review instrument and the per-pitfall coding released; right now the paper is reproducible only in principle for the case studies, not for the prevalence claims.\n\nOverall: a solid, useful SoK. The central claim—that DRL4Sec results are often artifacts of simplified environments and incomplete evaluation—holds up. It should go to serious peer review. I would cite it, and I'd bring it to reading group to discuss the taxonomy with students. My main editorial ask: treat the prevalence numbers as approximately true, address author-overlap head-on, and publish the coding data.","headline":"A genuinely useful SoK with a first DRL-specific pitfall taxonomy and honest case studies; just don't over-trust the exact prevalence numbers, and the author-overlap in the coded corpus needs addressing.","tokens_in":30180,"tokens_out":2992,"would_cite":true,"duration_ms":36235,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematization-of-knowledge paper claims that deep reinforcement learning research in cybersecurity is systematically undermined by 11 recurring methodological pitfalls—every one of the 66 papers it reviews exhibits at least two, with","keywords":["deep reinforcement learning","cybersecurity","systematic review","methodological pitfalls","MDP specification","evaluation bias","reproducibility","policy convergence"],"falsifier":"An independent team re-codes the same 66 papers with a pre-registered coding manual for the 11 pitfalls; if the average pitfall count drops well below 5.8 per paper, or inter-rater agreement falls below substantial, the central prevalence claim does not reproduce. A narrower check: in a case-study environment, replace the trained policy with random actions; if random actions already match trained performance, the paper's own gain-attribution warning is confirmed.","tokens_in":29300,"feed_emoji":"🛡️","tokens_out":4246,"duration_ms":46829,"temperature":0.7,"pith_summary":"Deep reinforcement learning is being applied to cybersecurity tasks such as network defense, malware evasion, fuzzing, and attack simulation, but this review argues that much of that literature is methodologically unsound. The authors define 11 recurring pitfalls spanning how security problems are modeled as Markov decision processes, how agents are trained, how results are evaluated, and how systems are assumed to behave in deployment. Reading 66 papers published from 2018 to 2025, they find that every paper exhibits at least two pitfalls, with an average of 5.8 per paper. They then run controlled experiments in three representative security domains to show that these pitfalls can degrade performance or inflate apparent success. If the review is right, a large share of published DRL-for-cybersecurity results do not currently support the deployable-security conclusions drawn from them.","feed_headline":"Every DRL security paper contains at least two pitfalls","feed_subtitle":"Sixty-six papers average 5.8 methodological flaws, so reported gains may not survive real deployment.","key_machinery":"The central object is an 11-pitfall taxonomy organized by the four stages of applying DRL—environment modeling, agent training, performance evaluation, and system deployment—with each pitfall given a short definition, such as 'the environment is inherently a POMDP but is treated as fully observable.' The taxonomy does the work of making methodological quality measurable: two independent reviewers coded each of the 66 papers for each pitfall as present, partially present, or not present, yielding the prevalence statistics. A second mechanism is a set of controlled ablations in three security domains—autonomous network defense, adversarial malware creation, and web security testing—where the a","core_discovery":"The paper's central claim is that the DRL-for-cybersecurity literature contains systematic, quantifiable methodological failure modes rather than isolated mistakes. It identifies 11 pitfalls: incomplete MDP specification, incorrect MDP modeling, unaddressed partial observability, missing hyperparameter reporting, absent variance analysis, undemonstrated policy convergence, weak motivation for using DRL, misattributed performance gains, oversimplified environments, unrealistic deployment assumptions, and unhandled non-stationarity. Across 66 significant papers, 71.2% lack clear evidence of policy convergence, 66.7% neglect variance analysis, 60.6% fail to address partial observability, and 40","pith_inferences":["The same pitfalls likely apply to adjacent work the review excluded, such as multi-agent RL security systems and non-deep RL security agents, so the findings probably generalize beyond the 66-paper corpus.","A concrete extension would be to convert the 11 pitfalls into a pre-registered reporting checklist or model card for DRL-for-cybersecurity papers, then measure whether adoption changes measured pitfall prevalence over time.","The authors' lenient coding rule—ambiguous cases were scored as less severe—means the headline prevalence figures are plausibly lower bounds; a stricter reviewer would likely count more pitfalls, not fewer.","The case-study pattern suggests a cheap validity test for any new DRL-for-cybersecurity result: compare the trained agent against a random-action policy in the same environment; if random actions already capture most of the performance, the contribution is in the environment, not in the learning."],"forward_implications":["Reported gains in DRL-for-cybersecurity papers should not be taken at face value until convergence, variance, and baseline performance are shown; otherwise improvements may come from environment design rather than from the learned policy.","A minimum reporting standard follows: complete MDP definitions, hyperparameters, training curves, multiple-seed variance with confidence intervals, and ablations against random or simpler baselines.","Treating security tasks as fully observable when they are POMDPs is widespread; recurrent policies or explicit POMDP formulations should become the default in such settings.","Deployment claims should be stress-tested, because agents trained against fixed adversaries or fixed action orders degrade sharply when those assumptions change, so training should include non-stationarity.","The review implies that the field's empirical evidence base for deployable DRL-based security is weaker than the volume of publications suggests."],"fun_headline_variants":["DRL security papers average 5.8 pitfalls each","Systematic flaws plague 66 DRL security papers","Study: DRL for cybersecurity is riddled with pitfalls","11 pitfalls found across 66 DRL security studies","Most DRL security research repeats the same mistakes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole quantitative argument depends on the authors' 11 pitfall definitions being applied consistently and correctly by two reviewers across 66 papers; if the definitions are unstable across coders, the average of 5.8 pitfalls per paper loses its quantitative meaning.","fun_headline_variants_meta":{"raw":{"variants":["DRL security papers average 5.8 pitfalls each","Systematic flaws plague 66 DRL security papers","Study: DRL for cybersecurity is riddled with pitfalls","11 pitfalls found across 66 DRL security studies","Most DRL security research repeats the same mistakes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000528,"raw_usage":{"total_tokens":2362,"prompt_tokens":703,"completion_tokens":1659,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":447,"completion_tokens_details":{"reasoning_tokens":1582}},"tokens_in":447,"tokens_out":1659,"duration_ms":10656,"temperature":1.0,"reasoning_tokens":1582,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T03:09:49.259426+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent team re-codes the same 66 papers with a pre-registered coding manual for the 11 pitfalls; if the average pitfall count drops well below 5.8 per paper, or inter-rater agreement falls below substantial, the central prevalence claim does not reproduce. A narrower check: in a case-study environment, replace the trained policy with random actions; if random actions already match trained performance, the paper's own gain-attribution warning is confirmed.","supporting_citations":[],"review_version":1}