{"id":"9da70220-1ad7-41d6-899d-b8150d9bfeb5","arxiv_id":"2505.11551","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of 85 learning-based CAN bus intrusion detection papers finds that combined known-unknown detection and federated learning are the least developed areas, while most papers ignore deployment metrics.","lead":"This paper reviews 85 studies of machine-learning systems that detect cyberattacks inside a car's internal network, the CAN bus, and sorts them by the type of attack they catch. It points to two under-developed areas, systems that handle both known and new attacks and privacy-preserving federated learning, which makes it a useful map for vehicle security researchers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coverage self-assessment is not independently anchored: Table 3's Scopus query uses TITLE-ABS-KEY while the reported automatic-search numbers come from title-only filtering, and no recall check quantifies papers missed by the title filter.","rationale":"The reader's weakest_assumption identifies the same coverage concern, and the reader's verdict is CONDITIONAL. I agree with that assessment. The survey is not fatally flawed; the search methodology is documented, the taxonomy is useful, and the internal inconsistencies (10 vs 11 combined papers; the sentence about trainable parameters appearing in the wrong section; the discrepancy between the narrative on Seo et al. and Table 6) are correctable. The stress-test pass confirms the coverage risk is the most load-bearing issue because the survey's contribution is precisely a quantitative map of the research area (38/27/11/9). The title-only filter is a legitimate, documented choice, and snowballing is a standard mitigation, but the absence of a recall bound or a PRISMA-style audit trail means the main counts are not independently checkable. The claim of being 'first to provide a comprehensive review' is also empirically fragile, as the related-work table shows prior surveys that cover ML and DL (Rajapaksha 2023, Lampe and Meng 2023a/2023b, Almehdhar 2024) and FL-related work; the distinguishing feature is the combined known-unknown and FL coverage. The concrete test I propose settles the concern by measuring the recall gap directly. No ad hominem or theatrical language is needed; the finding is that the survey should be accepted conditionally, with the authors asked to quantify or bound the title-filter recall loss and to fix the internal inconsistencies.","tokens_in":54420,"tokens_out":1982,"duration_ms":17916,"concrete_test":"Re-run the Section 4.1.1 automatic search without the title filter for the period up to January 2025, using Google Scholar or Scopus with the same query families, and screen the results by abstract for learning-based CAN IDS papers. Count how many papers satisfy the inclusion criteria but lack any of the six title keywords. If the missed set is non-negligible (e.g., more than 5-10% of the 85 papers or enough to change any of the 38/27/11/9 counts), the reported gap statistics and the 'first comprehensive survey' claim would need to be weakened or the dataset list supplemented with a PRISMA-style flow diagram and an appendix listing all 85 included papers.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is the first-comprehensive-survey claim and the 38/27/11/9 gap statistics. The weakest point is the completeness of the Section 4.1.1 title-only Google Scholar filter. The paper says 53,890 results with 'anywhere in the article' were reduced to 95 with 'in the title of the article,' then manual search and snowballing were used to recover missed papers, but no quantification of how many relevant papers lack the title keywords is given. If a meaningful fraction of learning-based in-vehicle IDS papers do not have 'CAN bus', 'in-vehicle', 'intrusion detection', 'anomaly detection', 'unknown attacks', or 'federated' in the title, the 38/27/11/9 counts and the conclusion that combined known-unknown detection and FL are under-researched would be unreliable. The manual search is itself also title-focused in the reported examples (ACM and IEEE Xplore queries use 'Document Title'), and the Scopus example uses TITLE-ABS-KEY with different term combinations ('in-vehicle' is not in the Scopus query, and 'unknown attacks' is, which may change recall). The snowballing step mitigates but does not bound the residual omission. As an additional internal check, the paper reports '10' combined known-unknown papers in Section 5.5 but '11' in Sections 4.1.2 and 8, and Section 5.5.1 says Hoang and Seo rely on CAN ID as a singular feature while Table 6 lists 'ID' and 'Payload' for Seo. The paper does not provide a PRISMA-style flow count of papers excluded at each filtering stage, nor a list of the 85 included papers, so the reader cannot reconstruct or audit the final set. Because the survey's entire value is its map of the field, this coverage risk is the load-bearing concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews learning-based intrusion detection systems (IDSs) for in-vehicle networks, focusing on machine learning (ML), deep learning (DL), and federated learning (FL) approaches. The authors followed a three-stage search strategy (Google Scholar automatic search, manual library search, and forward/backward snowballing) through January 2025, yielding 85 papers that are categorized into known-attack detection (38 papers), unknown-attack detection (27), combined known-unknown attack detection (reported as 11 in Sections 4.1.2 and 8, but as 10 in Section 5.5), and FL-based IDSs (9). The paper also reviews evaluation metrics in terms of performance, time, and memory requirements, and outlines limitations and future research directions. The authors claim to the best of their knowledge that this is the first comprehensive review of ML, DL, and FL-based IDSs for in-vehicle networks.","tokens_in":54732,"tokens_out":9413,"duration_ms":70913,"significance":"If the search is complete, the paper provides a useful map of the field: a documented search protocol, a clear known/unknown/combined taxonomy, detailed tables linking datasets, algorithms, input features, and model sizes, and a dedicated analysis of FL-based IDSs including non-IID and client-selection limitations. The systematic categorization and the review of evaluation metrics beyond accuracy are genuine strengths that would help practitioners assess deployability. The central 'under-researched areas' conclusions—combined known-unknown detection and FL—are falsifiable and important for guiding future work, but they rest on the completeness of the paper collection and on the internal consistency of the reported counts, which currently require attention.","major_comments":[{"comment":"The central claim of comprehensiveness and the gap statistics (38/27/11/9) depend on the completeness of the search, but the paper never quantifies the recall of its title-only filter. Table 2 shows 53,890 'anywhere in the article' results reduced to 95 with 'in the title of the article,' yet the manual search examples are also title-focused (ACM and IEEE use Document Title) and the Scopus query uses TITLE-ABS-KEY with a different term set. Without a PRISMA-style flow count of papers excluded at each stage and without a recall check (for example, by comparing against the reference list of a recent independent survey), the claim that these numbers represent the state of the art up to January 2025 is not fully supportable. Please add per-stage exclusion counts and quantify the fraction of relevant papers that the title filter misses.","section":"4.1.1 Data Sources and Search Strategy"},{"comment":"The number of combined known-unknown attack detection papers is inconsistent across the manuscript: Sections 4.1.2 and 8 state 11, Table 6 contains 11 rows, but Section 5.5 states '10 papers' and lists 10 references (excluding Althunayyan et al. 2024a), even though that work is reviewed in Section 5.5.3 and appears in Table 6. Since the total count of 85 is the paper's headline result, this discrepancy must be resolved.","section":"5.5 Known and Unknown Attacks Detection"},{"comment":"The sentence after Table 5 reads 'Among these studies, only three [Song et al., Fenzl et al., Le et al.] measure the trainable parameters,' which is verbatim from Section 5.3 (after Table 4) and is incorrect in the unknown-attack context: Table 5 itself lists model sizes or parameter counts for Sun et al., Thiruloga et al., Kukkala et al., Kristianto et al., Longari et al., Rajapaksha et al., and Kim et al. The duplication misstates the evidence on model-size reporting in the unknown-attack literature and should be corrected or removed.","section":"5.4.3 / Table 5"},{"comment":"The selection flow is not traceable: the automatic search yielded 95 title-filtered papers, yet the text says 'The papers from the automatic search were reduced from 10 to 9 after filtering by reading the full text.' The paper does not explain how 95 became 10 before the full-text step. The same applies to the manual (1,831 to 59) and snowballing (57 to 17) stages; a per-stage exclusion count is needed to support the reported totals.","section":"4.1.2 Selection Strategy"}],"minor_comments":[{"comment":"The section heading 'Evaluation Metics' should be corrected to 'Evaluation Metrics'.","section":"5.6"},{"comment":"In the sentence 'Koscher et al. [Chockalingam et al.(2016)...] highlight the feasibility of executing various types of wireless attack injections,' the narrative author name and the cited reference do not match; please change the narrative name to Chockalingam et al. or replace the citation.","section":"2.5.2"},{"comment":"The phrase 'without transfer-ring row data' should read 'without transferring raw data.'","section":"6.1"},{"comment":"The note 'DL: Deap Learning' should be corrected to 'Deep Learning.'","section":"Table 6 note"},{"comment":"The text states that Hoang et al. and Seo et al. use only CAN IDs as input features, but the checkmark placement in Table 6 is ambiguous because the M-C, ID, and Payload columns are not visually separated in the rendering; please ensure the table formatting clearly aligns with the text to avoid reader confusion.","section":"5.5.1 / Table 6"}],"recommendation":"major_revision","confidential_remarks":"The inconsistency between the Section 5.5 count (10) and the Sections 4.1.2/8 count (11) appears to hinge on the inclusion of the authors' own paper (Althunayyan et al. 2024a), which is also presented in Section 6.2 as addressing FL limitations and is used in Table 8. I recommend asking the authors to verify the counts and to ensure the survey's gap analysis does not over-rely on their own work; this is a normal consistency check, not an accusation of bias. Additionally, the 'first comprehensive survey' claim should be re-examined in light of Almehdhar et al. 2024, which already touches on FL-based IDSs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my take on arXiv:2505.11551.\n\nThe paper is a survey, not a research result, and it does the job a good survey should: it documents a structured search, classifies 85 papers into known-attack, unknown-attack, and combined known-unknown detection, adds a dedicated federated learning section, and tabulates evaluation metrics with attention to deployment constraints like model size and latency. The known/unknown/combined framing is genuinely useful—most prior surveys don't separate combined detectors—and the FL section, while small (9 papers), is a reasonable synthesis that points at non-IID data and client selection as open problems. That's real value for someone entering the area or looking for gaps.\n\nThe soft spots are real but mostly editorial. The count of combined known-unknown papers is 10 in Section 5.5 and 11 in Sections 4.1.2 and 8; that should be reconciled. The sentence about \"only three papers\" measuring trainable parameters appears in Section 5.4.3 but cites Song, Fenzl, and Le—all known-attack papers—so it's clearly a copy-paste from Section 5.3. The bigger concern, which the stress-test flags correctly, is coverage: the automatic Google Scholar step reduced 53,890 hits to 95 by filtering to title-only keywords, and while manual search and snowballing were used to recover missed papers, the authors don't say how many relevant papers lack those title terms. That doesn't sink the survey—snowballing from 85 papers is plausible mitigation—but it means the 38/27/11/9 counts should be read as approximate, not exact, and the claim of being \"first comprehensive\" review is stronger than the evidence supports. A PRISMA-style flow diagram and a list of included papers would have made the set auditable; their absence is a miss.\n\nThe self-citation is worth a note but not a flaw: the authors' own H-FL paper appears in the FL section as a solution to standard FL limitations, which is a mild conflict of interest in a gap analysis, but the categorization itself doesn't depend on it.\n\nWho this is for: a graduate student or researcher starting in-vehicle IDS work gets a decent map and a clear gap list. A referee should engage with it; the fixes are mechanical. My recommendation: send it to peer review with a request for minor revision—reconcile the counts, fix the duplicated sentence, and add a supplementary list of the 85 papers with a recall caveat. It will be more useful after that cleanup, but it's already a legitimate contribution.","headline":"A useful map of 85 learning-based in-vehicle IDS papers with a defensible known/unknown/combined taxonomy; the coverage claims need a caveat and the manuscript needs a cleanup pass.","tokens_in":55342,"tokens_out":2257,"would_cite":true,"duration_ms":21978,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps 85 learning-based intrusion detection systems for in-vehicle CAN networks and argues that combined known-unknown detection and federated learning are the least developed approaches.","keywords":["CAN bus","intrusion detection system","in-vehicle network","machine learning","deep learning","federated learning","anomaly detection","connected autonomous vehicles"],"falsifier":"Rerun the literature search without the title-only restriction, searching abstracts and full texts instead, and count the additional relevant learning-based CAN-bus IDS papers; if that count is large enough to change the 38/27/11/9 distribution, the paper's under-researched-area conclusions would need to be revised.","tokens_in":54199,"feed_emoji":"🚗","tokens_out":9161,"duration_ms":79184,"temperature":0.7,"pith_summary":"This survey tries to establish a reliable map of learning-based intrusion detection systems (IDSs) for the Controller Area Network (CAN) bus—the internal data bus that connects a vehicle's electronic control units—and to identify where the research is thinnest. It reviews 85 papers that use machine learning, deep learning, or federated learning to detect cyberattacks inside vehicles, and organizes them by whether they detect known attacks, unknown attacks, or both. The paper argues that known-attack detection is a mature research area, while IDSs that can classify known attacks and still catch novel ones, together with federated-learning IDSs, are the least developed. It also argues that most evaluations report accuracy-style metrics while neglecting the inference time, memory footprint, and real-time constraints that decide whether a detector can actually be put into a car. If this map is right, research effort should shift toward combined known-unknown detectors and privacy-preserving federated approaches.","feed_headline":"85-paper survey finds two blind spots in car-network security","feed_subtitle":"Known-attack detection is mature; detecting both known and unknown attacks, plus federated learning, lags behind.","key_machinery":"The machinery carrying the argument is a three-part attack-detection taxonomy—known, unknown, and combined known-unknown attacks—cross-cut by input-feature type (CAN ID, payload, full CAN frame), together with a systematic search protocol of automatic title-filtered Google Scholar queries, manual library searches, and forward and backward snowballing that yields the 85-paper corpus. The taxonomy does the analytical work: it turns 'detect both known and unknown attacks' into a visible category, and the feature cross-cut exposes blind spots such as CAN-ID-only detectors being unable to catch payload-manipulation attacks. A second mechanism is the evaluation-metric rubric, which sorts reported measures into performance, time complexity, memory, and other categories and turns the paper's normative claim—that deployable IDSs must be judged on latency and footprint as well as accuracy—into a concrete checklist.","core_discovery":"The paper's central claim is that learning-based in-vehicle IDS research splits into three attack-detection regimes—known, unknown, and combined known-unknown attacks—and that its 85 collected papers distribute as 38, 27, and 11 across these regimes, with 9 additional papers on federated learning. Within each regime it further groups systems by input features: CAN ID only, payload only, or full CAN frames. It finds that known-attack detectors are mostly supervised deep-learning classifiers trained on CAN frames, while unknown-attack detectors are mostly unsupervised reconstruction or prediction models that profile normal traffic. The paper also claims to be the first survey to review machine learning, deep learning, and federated learning for in-vehicle networks together under a structured search strategy, and its metric review concludes that deployment-relevant measures such as model size, latency, and memory use are widely omitted.","pith_inferences":["Editorial inference: the title-only Google Scholar filter means the 85-paper count is better read as a lower bound; a fuller abstract and full-text search would likely surface additional relevant papers that could shift the 38/27/11/9 distribution.","Editorial inference: a multi-stage architecture—supervised classification first, unsupervised anomaly detection as a backstop—appears several times in the survey and can be tested as a general template; its main open question is whether the unsupervised stage stays reliable on vehicle makes whose normal traffic differs from the training fleet.","Editorial inference: applying the survey's deployment-metric checklist uniformly to the reviewed systems would probably shrink the list of deployable solutions sharply, since many papers report only accuracy and F1.","Editorial inference: the taxonomy implies a testable ordering—that payload-only and full-frame detectors should beat ID-only detectors on spoofing and fuzz attacks—but the surveyed papers rarely compare those feature regimes on identical data, so a controlled benchmark across feature types would be a direct follow-up."],"forward_implications":["Known-attack detection is the most crowded of the three regimes, so new work that only adds another supervised classifier on the same datasets will add little.","If the survey's counts are correct, combined known-unknown detection and federated-learning IDSs are the two least developed areas and the natural focus for future research.","Evaluation practice needs to include inference time, detection latency, and model size; the paper notes a vehicle-level IDS should process each packet in under roughly 10 ms to meet real-time safety requirements.","Most federated-learning IDS studies assume clients hold identically distributed data, and only a few simulate non-IID conditions, so current FL evaluations are not yet realistic for real vehicle fleets.","Because centralized training dominates the surveyed work, privacy, communication overhead, and single-point-of-failure concerns remain open, which is the motivation the paper gives for hierarchical or personalized FL."],"supporting_citations":[{"why":"Supplies the systematic literature-review protocol that defines the survey's search, inclusion, and snowballing steps.","marker":"Kitchenham and Brereton 2013"},{"why":"The closest prior AI-based in-vehicle IDS survey; establishes the baseline taxonomy and limitations this survey claims to extend.","marker":"Rajapaksha et al. 2023a"},{"why":"Defines the GAN-based combined known-unknown detection approach and provides the Car-Hacking dataset used across the reviewed corpus.","marker":"Seo et al. 2018"},{"why":"Provides the OTIDS dataset and attack scenarios (DoS, fuzzy, spoofing) that anchor much of the known-attack and anomaly-detection literature reviewed.","marker":"Lee et al. 2017"},{"why":"Defines FedAvg, the aggregation algorithm assumed in most of the federated-learning IDS papers the survey reviews.","marker":"McMahan et al. 2017"},{"why":"Provides the Car Hacking: Attack & Defence Challenge 2020 dataset used by many reviewed known-attack, unknown-attack, and FL-based systems.","marker":"Kang et al. 2021"},{"why":"A prior survey of federated learning for connected vehicles that, the paper notes, omits in-vehicle IDSs, supporting the claimed gap.","marker":"Chellapandi et al. 2023"}],"fun_headline_variants":["Survey: 85 papers on car-network IDS reveal gaps in unknown-attack detection","Learning-based car intrusion detection survey flags federated learning lag","85-paper IDS review: known attacks covered, combined and FL fall short","In-vehicle network IDS survey: machine learning matures, federated lags","Car security survey: detection of known attacks solid, unknown and FL weak"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's counts and gap conclusions stand on the assumption that its search strategy—especially the Google Scholar step that keeps only papers whose titles contain the chosen keywords—recovered essentially all relevant learning-based in-vehicle IDS work published up to January 2025.","fun_headline_variants_meta":{"raw":{"variants":["Survey: 85 papers on car-network IDS reveal gaps in unknown-attack detection","Learning-based car intrusion detection survey flags federated learning lag","85-paper IDS review: known attacks covered, combined and FL fall short","In-vehicle network IDS survey: machine learning matures, federated lags","Car security survey: detection of known attacks solid, unknown and FL weak"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1358,"prompt_tokens":931,"completion_tokens":427,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":326}},"tokens_in":547,"tokens_out":427,"duration_ms":4436,"temperature":1.0,"reasoning_tokens":326,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:13:34.436723+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the literature search without the title-only restriction, searching abstracts and full texts instead, and count the additional relevant learning-based CAN-bus IDS papers; if that count is large enough to change the 38/27/11/9 distribution, the paper's under-researched-area conclusions would need to be revised.","supporting_citations":[],"review_version":1}