{"id":"5e302264-a6d6-42c5-9c5c-78a52a2ab9c4","arxiv_id":"2604.04952","paper_version":6,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"An open-source dual-score ML NIDS on commodity hardware reports F1=0.9985 on CTU-13 Neris while Suricata (50k rules) alerts zero times and Zeek scores F1=0.042 under matched conditions.","lead":"The paper presents ML Defender (aRGus NDR), a low-cost open-source network intrusion system with embedded machine learning that reports near-perfect botnet detection on the CTU-13 Neris dataset. It also claims that under the same conditions Suricata raised zero alerts and Zeek almost none, arguing the three detection styles are complementary.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Manuscript body is a different paper; the NIDS central claim cannot be audited from the supplied text.","rationale":"The reader correctly diagnosed abstract-only access and the wrong full-text body, set UNVERDICTED with low confidence, and identified the CTU-13 / Suricata interpretation as the weakest assumption. That diagnosis is exact: without the actual NIDS manuscript there is no train/test hygiene, no feature list, no rule-set dump, and no multi-scenario results to audit. My concern is the same one, sharpened: the mismatch is not a minor packaging issue but a complete absence of the paper whose claim is under review. No new technical soft spot inside a missing argument can be invented. The concrete test is simply to obtain the correct document (or confirm its absence); until then the verdict remains UNVERDICTED and the reader's assessment stands unchanged.","tokens_in":11579,"tokens_out":514,"duration_ms":4617,"concrete_test":"Confirm whether the correct PDF/source for arXiv:2604.04952 (ML Defender / aRGus NDR) is available; if the body remains the trailer-synthesis survey, the NIDS claim stays uncheckable and the paper cannot be accepted on this package. If the real NIDS manuscript appears, re-run the Suricata offline experiment on the exact Neris pcap with the stated ET Open rule counts and report alert counts plus any rule-match logs.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The strongest claim (three-paradigm comparison on CTU-13 Neris with F1=0.9985 vs Suricata zero alerts vs Zeek F1=0.042) is stated only in the abstract of arXiv:2604.04952. The CACHEABLE full-manuscript body under that paper_id is an unrelated survey on generative AI for video trailer synthesis (arXiv:2604.04953). No sections, tables, feature definitions, train/test protocol, rule-set verification, or code artifacts for ML Defender exist in the provided text. Therefore the load-bearing premises the reader already flagged—CTU-13 Neris representativeness for modern ransomware/DDoS, Suricata miss as architectural rather than signature-age, and the two FPs as pure VirtualBox artifacts—cannot be checked at all. The central claim is not merely weakly supported; it is not present in the manuscript body that was supplied for review.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The abstract of arXiv:2604.04952 claims an open-source C++20 NIDS (ML Defender / aRGus NDR) with embedded ML inference on commodity hardware (150–200 USD), a six-component eBPF/XDP–ZeroMQ–Protobuf pipeline, and a dual-score Fast Detector + Random Forest architecture. On CTU-13 Neris it reports F1=0.9985, Precision=0.9969, Recall=1.0000 (2 FP in 12,075 benign flows, attributed to VirtualBox artifacts), and a three-paradigm comparison under identical conditions in which Suricata 6.0.10 with 50,010 ET Open rules yields zero alerts, Zeek 8.1.2 yields 14 detections (F1=0.042), and aRGus NDR achieves near-perfect recall, motivating a taxonomy of signature / scripted-behavioral / ML-behavioral decision architectures. The full manuscript body supplied for review, however, is an unrelated survey on generative AI for video trailer synthesis (arXiv:2604.04953), with no NIDS architecture, features, training protocol, Suricata/Zeek setup, tables, or code artifacts for the claimed system.","tokens_in":11818,"tokens_out":1067,"duration_ms":13161,"significance":"If the abstract’s results were substantiated in a matching manuscript—open-source embedded ML NIDS at low cost, reproducible three-way comparison against Suricata and Zeek on the same capture, and a clear complementarity argument—the work would be of practical interest for resource-constrained organizations and of methodological interest for NIDS evaluation design. Those contributions cannot be assessed from the supplied full text, which does not contain them. No machine-checked proofs, NIDS code, feature definitions, or falsifiable experimental protocol for ML Defender appear in the manuscript body under review.","major_comments":[{"comment":"Manuscript identity failure: the titled abstract (ML Defender / aRGus NDR, cs.CR) does not correspond to the full manuscript body, which is a survey on generative AI for video trailer synthesis (different title, authors, venue claims, and arXiv lineage). No section, equation, table, or figure in the supplied body addresses eBPF/XDP, ZeroMQ, Protocol Buffers, dual-score Fast Detector + Random Forest, CTU-13 Neris evaluation, Suricata 6.0.10, or Zeek 8.1.2. The central claims of the abstract are therefore not present for technical review.","section":null},{"comment":"Because the NIDS evaluation is absent from the body, load-bearing premises stated only in the abstract cannot be audited: (i) train/test protocol and feature set on CTU-13 Neris; (ii) confirmation that Suricata’s zero alerts reflect architectural limits rather than signature age, rule enablement, or replay fidelity (the abstract’s “DAY 148” offline check is not documented); (iii) attribution of the two false positives solely to VirtualBox artifacts; (iv) representativeness of a 2011 IRC botnet capture for modern ransomware/DDoS in hospitals and schools. Without these, the superiority and complementarity claims cannot be verified.","section":null},{"comment":"The abstract’s “first three-paradigm experimental comparison” and the proposed taxonomy (signature / scripted behavioral / ML behavioral) require identical-condition methodology, rule-set inventory, Zeek script/notice configuration, and decision-layer analysis. None of this material exists in the supplied manuscript, so the taxonomy is an ungrounded assertion relative to the text under review.","section":null}],"minor_comments":[{"comment":"Abstract alone: CTU-13 Neris is a single, dated scenario; even a correct manuscript would need multi-dataset or modern-traffic evidence before production claims for hospitals/schools.","section":null},{"comment":"Abstract alone: free parameters (Fast Detector thresholds, fusion rule, RF hyperparameters/features) are not specified; reproducibility would require them plus release artifacts under the stated MIT license.","section":null},{"comment":"The trailer-synthesis body that was supplied has its own presentation issues (OCR-garbled equations, placeholder figure captions, mixed arXiv IDs in the header), but those are irrelevant to the NIDS claims and are not the basis of this recommendation.","section":null}],"recommendation":"reject","confidential_remarks":"The supplied package pairs the abstract of 2604.04952 with the full text of an unrelated paper (trailer GenAI survey, consistent with 2604.04953). This looks like a corpus/packaging error rather than a reviewable submission. I recommend the editor verify the correct PDF/source for 2604.04952 before any further review cycle; until a matching NIDS manuscript is provided, there is nothing to accept or revise on the stated claims. I have not attempted to score the trailer survey as if it were the submission."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing you need to know: under the paper_id for ML Defender we were given the full text of a different paper (generative AI for video trailers). So the headline three-paradigm result on CTU-13 Neris exists only as an abstract claim. We cannot check train/test hygiene, features, fusion rule, or the Suricata offline experiment.\n\nWhat would be new if the missing paper matches the abstract is not the dual-score RF or eBPF/XDP pipeline themselves—those are established—but a matched comparison under identical conditions: Suricata 6.0.10 + 50k ET Open rules producing zero alerts (with a packet-level offline check that rules for IRC/botnet/trojan were live), Zeek 8.1.2 logging the full botnet profile yet only 14 correct detections (F1 0.042), and the authors’ system at F1 0.9985 / Recall 1.0 with two FPs they attribute to VirtualBox. Framing that as a taxonomy of where knowledge is encoded (signature vs scripted behavioral vs ML behavioral) and arguing complementarity is a clean systems point. Open-source C++20 on 150–200 USD hardware aimed at hospitals and schools is also a real engineering target.\n\nSoft spots, in proportion: the evaluation is a single 2011 IRC botnet scenario. That is thin support for modern ransomware/DDoS claims. Extreme metrics always need leakage and threshold scrutiny we cannot perform. Calling the comparison “first” is hard to verify without a literature section. The two-FP VirtualBox story is plausible but uncheckable here. None of these are fatal if the code and multi-scenario results exist; they are ordinary applied-security risks.\n\nWho it is for: people building or deploying NDR for under-resourced orgs, and anyone who wants a concrete Suricata/Zeek baseline on a public capture. If the MIT code and full paper appear and reproduce, I would read it and might cite the comparison. As packaged, a serious editor should not desk-reject on abstract alone—the problem and numbers are sharp enough for referee time—but the body mismatch means we cannot treat the result as established yet. Engage only after the actual manuscript and artifacts are in hand.","headline":"The abstract promises a useful cheap open-source ML NIDS and a clean Suricata/Zeek/ML bake-off, but the supplied manuscript body is an unrelated video-trailer survey, so the central claim cannot be audited.","tokens_in":12501,"tokens_out":566,"would_cite":false,"duration_ms":5317,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"An open-source ML network IDS on $150–200 hardware detects all CTU-13 Neris botnet flows while Suricata’s 50,010 rules fire none and Zeek barely alerts.","keywords":["network intrusion detection","botnet detection","embedded machine learning","eBPF/XDP","Suricata","Zeek","CTU-13","open-source NIDS"],"falsifier":"Replay a current ransomware or DDoS traffic corpus that contains flows known to match active Suricata/ET Open rules; if Suricata then fires and aRGus either misses or floods false positives, the superiority and complementarity claims on that dataset collapse.","tokens_in":12396,"feed_emoji":"🛡️","tokens_out":1016,"duration_ms":15145,"temperature":0.7,"pith_summary":"Hospitals, schools, and small organizations face ransomware and DDoS without enterprise budgets. This paper presents ML Defender (aRGus NDR), an open-source C++20 network intrusion detection system with embedded machine-learning inference that runs on ordinary hardware costing about 150–200 USD. On the CTU-13 Neris capture it reports F1 = 0.9985, precision = 0.9969 and perfect recall, with only two false positives among more than twelve thousand benign flows. Under identical conditions Suricata 6.0.10 with the full ET Open rule set produces zero alerts, and Zeek 8.1.2 produces only fourteen correct detections. The authors use that three-way comparison to define a taxonomy of decision architectures—signature, scripted behavioral, and ML behavioral—and argue the three layers are complementary rather than competitive.","feed_headline":"ML IDS on $150 gear catches all Neris bots; Suricata fires none","feed_subtitle":"Open-source dual-score detector hits F1 0.9985 while 50k rules and Zeek scripts largely miss under the same replay.","key_machinery":"A six-component pipeline over eBPF/XDP, ZeroMQ and Protocol Buffers that feeds a dual-score Fast Detector + Random Forest classifier; the architecture encodes detection knowledge in learned behavioral scores rather than hand-written signatures or scripts.","core_discovery":"Under identical conditions on CTU-13 Neris, the dual-score Fast Detector plus Random Forest pipeline of aRGus NDR reaches F1 = 0.9985 and recall = 1.0, while Suricata with 50,010 ET Open rules generates zero alerts (confirmed offline on 323,154 packets with hundreds of IRC, botnet/C2 and trojan signatures active) and Zeek 8.1.2 yields only fourteen correct detections (F1 = 0.042). These results establish a taxonomy of signature, scripted-behavioral and ML-behavioral decision architectures that encode network knowledge at different layers and can operate together.","pith_inferences":["If the same dual-score pipeline is retrained on recent ransomware C2 and encrypted-tunnel features, the cheap appliance model could become a practical default for school and hospital edge networks.","The Suricata zero-alert result under full rule load suggests evaluation benchmarks should report both signature-hit rate and architectural miss rate, not only ML F1.","Because Zeek already saw the full botnet profile, a lightweight bridge that turns its structured logs into aRGus feature vectors could give hybrid detection without new packet capture.","The taxonomy implies that future NIDS papers should state explicitly which encoding layer (signature, script, or learned score) carries the decision, so comparisons stop mixing incomparable systems."],"forward_implications":["Resource-constrained sites can deploy a full-recall botnet detector on commodity hardware without enterprise licenses.","Signature engines that observe the same traffic may still produce zero alerts, so pure rule-based NIDS cannot be assumed sufficient.","Scripted behavioral systems can log the complete botnet profile yet still fail to raise alerts, showing that telemetry alone is not detection.","Organizations can run Suricata signatures, Zeek telemetry and an ML behavioral classifier side-by-side as complementary layers rather than substitutes.","The MIT-licensed C++20 codebase removes the cost barrier that previously kept ML NIDS out of small networks."],"fun_headline_variants":["Dual-score ML NIDS hits F1 0.9985 on Neris; Suricata zero alerts","Embedded ML on $150 hardware recalls all Neris bots where Suricata fails","aRGus NDR F1 0.9985 vs Suricata zero and Zeek 0.042 on same Neris","Open-source ML NIDS catches every Neris bot; 50k rules catch none","ML behavioral classifier tops signatures and scripts on CTU-13 Neris"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That a 2011 IRC-based botnet capture (CTU-13 Neris) and the authors’ reading of Suricata’s silence and of two VirtualBox-related false positives are enough to support claims about defending modern ransomware and DDoS in hospitals and schools.","fun_headline_variants_meta":{"raw":{"variants":["Dual-score ML NIDS hits F1 0.9985 on Neris; Suricata zero alerts","Embedded ML on $150 hardware recalls all Neris bots where Suricata fails","aRGus NDR F1 0.9985 vs Suricata zero and Zeek 0.042 on same Neris","Open-source ML NIDS catches every Neris bot; 50k rules catch none","ML behavioral classifier tops signatures and scripts on CTU-13 Neris"]},"model":"grok-4.5","effort":"low","cost_usd":0.00478,"raw_usage":{"total_tokens":1508,"prompt_tokens":971,"num_sources_used":0,"completion_tokens":105,"cost_in_usd_ticks":47800000,"prompt_tokens_details":{"text_tokens":971,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":432,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":971,"tokens_out":105,"duration_ms":13410,"temperature":1.0,"reasoning_tokens":432,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T13:43:15.789070+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Replay a current ransomware or DDoS traffic corpus that contains flows known to match active Suricata/ET Open rules; if Suricata then fires and aRGus either misses or floods false positives, the superiority and complementarity claims on that dataset collapse.","supporting_citations":[],"review_version":2}