{"id":"c928bde3-b71e-4849-9b3c-6df7ff6cd8cb","arxiv_id":"2411.14278","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A systematic review and taxonomy of adaptive anomaly detection for cyber-physical systems, with the main finding that few methods combine real-time data processing with model adaptation.","lead":"This paper systematically reviews 65 papers on adaptive anomaly detection for cyber-physical systems and organizes them by attack type, application, learning paradigm, data management, and algorithms. It finds that most current methods adapt only one of two needed components, data processing or model updating, and lays out research priorities.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 'rarely both' is not operationalized and may be an artifact of loose inclusion criteria that select for self-described adaptivity rather than for both AAD components.","rationale":"The reader identified corpus representativeness as the weakest assumption, which is related to but distinct from my concern. My concern is more specific: even if the corpus were representative of papers matching the search queries, the headline finding depends on an operationalization of 'data processing' versus 'model adaptation' that the paper never provides, and on inclusion criteria that may select for self-described adaptivity rather than for the two-component AAD definition. This makes the main quantitative conclusion internally under-supported and potentially circular. The paper still has value: the taxonomy, per-paper summaries, and future-directions discussion are useful, and no derived predictions are made, so the circularity burden is low. However, the central 'rarely both' claim should be treated as a hypothesis that requires a transparent re-coding and sensitivity analysis. I therefore keep the reader's CONDITIONAL verdict rather than moving to ACCEPT or REJECT; the condition should be that the authors provide the operational coding and the 2×2 distribution, or explicitly soften the claim to a qualitative observation about the reviewed set rather than a field-level gap.","tokens_in":44627,"tokens_out":4916,"duration_ms":48315,"concrete_test":"Construct an explicit coding rubric: 'data processing' = online/incremental/streaming data handling (e.g., sliding windows, chunk-by-chunk processing, prequential evaluation); 'model adaptation' = an explicit model update mechanism (e.g., retraining, incremental learning, drift-triggered updating). Independently re-code all 47 research papers from their full texts, blinded to the paper's classifications, and build the 2×2 table of data-processing-only, model-adaptation-only, both, and neither. If 'both' is not rare (e.g., >20%) or a large share fall in 'neither,' the headline finding fails. Then re-run the screening restricted to papers whose full text explicitly implements both components; if the 'both' proportion changes materially, the finding is an artifact of the loose inclusion criteria.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central finding is that reviewed works 'focused on a single aspect of adaptation (either data processing or model adaptation) but rarely in both at the same time.' This claim is load-bearing because it defines the review's contribution and motivates its future-research agenda. However, it is not supported by the presented methodology. In Section I, AAD is defined as requiring two components: (1) near real-time data processing and (2) a predefined learning mode for model adaptation. Yet the inclusion criteria in Section IV.A.1 only require 'a clear focus on adaptive, cognitive, or autonomous mechanisms for anomaly detection in CPS'—not both components. As a result, many included papers are classified in Tables II–VI with explicit weaknesses such as 'Lack of continual learning' or 'Offline learning,' meaning they are not AAD under the paper's own definition. The corpus is therefore a sample of papers that self-describe as adaptive, not a sample of systems satisfying the two-component AAD definition. The 'rarely both' finding may then be an artifact of selection: if the screening step already admits papers doing only one aspect, it is unsurprising that few do both. Additionally, the paper never defines 'data processing' and 'model adaptation' operationally, nor does it present the 2×2 cross-tabulation (data processing only, model adaptation only, both, neither). Figure 4 shows distributions over application, learning paradigm, offline/online, and algorithm, but not the two components named in the headline finding. Section VI.A states the conclusion without such evidence. The internal inconsistency in Section VI.A—ICS reported as 24 of 47 papers (51%) while Figure 2(b) reports 41.5%—further reduces confidence in the quantitative tallies that underlie the finding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic literature review (SLR) of adaptive anomaly detection (AAD) for cyber-physical systems. The authors retrieved 397 candidate papers from five databases and, after screening and snowballing, analyzed 47 research papers and 18 survey papers published between 2013 and November 2023. They introduce a taxonomy organized by attack type, CPS application, learning paradigm, data management, and algorithm, and summarize each reviewed paper in per-application tables. The paper's central claims are that it is the first SLR on AAD in CPS and that reviewed works focus on either data processing or model adaptation but rarely integrate both.","tokens_in":44886,"tokens_out":3375,"duration_ms":31886,"significance":"If the central finding is supported, the review would provide a useful organizational map of an emerging subfield and a concrete gap statement that could motivate future research. The paper ships a substantial amount of structured information: detailed per-paper tables with datasets, attacks, strengths, and weaknesses; a taxonomy linking applications, learning paradigms, and algorithm families; and a set of future-research directions. The authors also explicitly acknowledge limitations of the search and selection process (Section VI-B), which is a sign of methodological transparency. The review is not built on derived equations or fitted parameters, so the usual circularity concerns do not apply; the taxonomy is used to organize the corpus, which is normal practice for a literature review. However, the value of the contribution depends on whether the corpus is representative and whether the headline claim about the rarity of integrated AAD is actually measured by the presented methodology.","major_comments":[{"comment":"The data extraction process is said to be 'inspired by Gama et al. [89]', but reference [89] is Loeffel's PhD thesis on adaptive machine learning algorithms for data streams, not a work by Gama and colleagues. The concept-drift adaptation framework cited elsewhere as [34] is Gama et al.'s survey. This citation error makes it harder to trace the origin of the extraction template. Please verify the intended reference and correct it.","section":"Section IV-A.5 and References"}],"minor_comments":[{"comment":"The phrase 'the former' is used twice in the discussion of supervised anomaly detection: 'The former is related to a significantly lower amount of anomalous data versus normal data' and 'The former entails generating a representative proportion of data samples with anomaly labels.' The second 'former' should be 'latter' or the sentence should be rephrased.","section":"Section II-B"},{"comment":"The list of keywords is typeset inconsistently: some terms use semicolons, some use 'AND' in uppercase, and some are capitalized oddly (e.g., 'adaptive cyber security ; dynamic adaption AND cyber security'). Please standardize the formatting of query terms and consider presenting them in a monospaced or quoted style.","section":"Section IV-A.3"},{"comment":"The caption says 'Figure 2(b) displays the application distribution' but the text in Section IV-B says 'Figure 2(b)' while referring to ICS as 41.5%. The percentages 41.5%, 27.7%, 13.8%, 9.2%, and 7.7% do not sum to 100% (they sum to 99.9%, which is fine for rounding), but the ICS value is called out as questionable due to the inconsistency with Section VI-A.","section":"Figure 2"},{"comment":"The tables use inconsistent capitalization for algorithms (e.g., 'Hhull' in Table II for one-class scaled convex hull, 'Extra-Threes' in Table VI for Extra-Trees, 'ANFIS' defined only in text). A final proofreading pass for typographical consistency in table entries is needed.","section":"Tables II-VI"},{"comment":"The sentence 'The model is first trained offline where parameters are optimized using a GA' appears to be missing a closing detail about how the GA was applied (e.g., the fitness function and the number of generations). This is a minor clarity issue in the description of Feng et al. [160].","section":"Section V-B.1"},{"comment":"Several references appear truncated or informal, e.g., 'Y . Zhang and H.-L. Lu' is fine, but 'Jiao et al. [40]' is titled 'Cyberattack-resilient load forecasting with adaptive robust regression' in the text while the reference list entry is 'Cyberattack-resilient load forecasting with adaptive robust regression'—please ensure consistency between in-text titles and the reference list. Also check the spelling of 'International Journal of Forecasting' in [166].","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a clear scope and a large amount of structured content, but the central claim ('rarely both') is currently not demonstrated by the reported methodology, and there is an internal inconsistency in the ICS share. The authors should be asked to add a cross-tabulation of the two AAD components, fix the percentage discrepancy, and provide an online appendix with search logs and the full list of included papers. The claim of being the first SLR on AAD in CPS could also be supported by a more explicit statement of how the search covers the intersection of AAD and CPS, since several related surveys listed in Table I are close in scope. I do not see a load-bearing error that would require rejection; the issues are fixable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the AAD-in-CPS SLR. Bottom line: it's a useful map of 47 research papers and 18 surveys, with careful per-paper tables and a sensible taxonomy, but its headline finding — that prior work focuses on either data processing or model adaptation, rarely both — isn't actually demonstrated by the methodology. The inclusion criteria only require a 'clear focus on adaptive, cognitive, or autonomous mechanisms', not the two-component AAD definition given in Section I. So the corpus is selected for self-described adaptivity, and the 'rarely both' result is at least partly an artifact of that selection. The paper never shows the 2x2 cross-tab of the two components, and Section VI.A states the conclusion without that evidence. There's also an internal inconsistency: ICS is reported as 41.5% of the corpus in Section IV-B but as 24 of 47 (51%) in Section VI-A.\n\nWhat it does well: the SLR process is documented (Okoli's method, snowballing with a 99th-percentile citation threshold), the taxonomy covers attack types, applications, learning paradigm, data management, and algorithms, and the per-paper tables (strengths/weaknesses, datasets, attacks) are genuinely useful for practitioners and newcomers. The future-research section is sensible, especially the points on evaluation metrics and low detection latency. The authors also acknowledge selection bias in Section VI.B, which is honest.\n\nThe inconsistency and the unoperationalized headline finding are real soft spots, but they're fixable with a revision: define the two components operationally, add the cross-tab, and recompute the percentages. I don't think the central argument collapses — the review still gives a fair organizational picture of the field — but the quantitative claims need support before anyone should repeat them.\n\nWho's it for: practitioners and new researchers wanting a structured entry point to AAD in CPS. It deserves serious peer review; conditional accept with major revision.","headline":"A useful but flawed SLR: the taxonomy and per-paper tables are valuable, but the headline 'rarely both' finding is not supported by the stated methodology and the reported statistics are internally inconsistent.","tokens_in":45465,"tokens_out":1795,"would_cite":false,"duration_ms":17319,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic review claims to be the first to map adaptive anomaly detection in cyber-physical systems, and finds that most methods adapt either data processing or the detection model, but rarely both.","keywords":["adaptive anomaly detection","cyber-physical systems","systematic literature review","concept drift","online learning","intrusion detection","taxonomy","streaming data"],"falsifier":"Rerun the same review with a broader keyword set and a lower citation threshold, and count the share of included methods that implement both near real-time data processing and model adaptation; if that share is well above what the paper reports, the 'rarely both' finding fails.","tokens_in":44418,"feed_emoji":"🛡️","tokens_out":4681,"duration_ms":41087,"temperature":0.7,"pith_summary":"This paper is a systematic literature review of adaptive anomaly detection (AAD) for cyber-physical systems (CPS), and its central claim is that it is the first such review. The authors gathered 397 candidate papers, narrowed them to 65 (47 research and 18 survey papers) from 2013 to 2023, and classified them with a new taxonomy built on attack types, CPS application, learning paradigm, data management, and algorithms. Their headline finding is that most reviewed methods handle only one side of adaptation—either fast data processing or model updating—and that combining both in a single system is rare. If the finding holds, it gives the field a map and identifies a concrete gap: few existing detectors can update themselves online while still making low-latency decisions.","feed_headline":"Most adaptive attack detectors adapt data or model, not both","feed_subtitle":"A systematic review of 65 cyber-physical security papers finds near real-time processing and model updating are rarely combined.","key_machinery":"The review's load-bearing object is its taxonomy, which splits AAD into near real-time data processing versus model adaptation and then cross-classifies each reviewed paper by CPS application, learning paradigm (supervised, unsupervised, reinforcement), data management (offline vs online), and algorithm family. The classification uses the concept-drift and online-learning vocabulary of Gama et al. [34]—learning modes, adaptation methods, ensemble updating—and the data-stream definition from [89]. This machinery is what turns a pile of papers into the 'rarely both' finding, because each paper is tagged on both adaptation dimensions.","core_discovery":"The paper argues that AAD in CPS requires two components—near real-time data processing and a predefined learning mode for model adaptation—and that the reviewed literature rarely integrates them. Of the 47 research papers, roughly three-quarters use supervised learning, nearly two-thirds train and test offline, and only about a third report online evaluation. The paper's taxonomy organizes the field by application (industrial control, vehicles, power grid, IoT, smart grid), learning paradigm, and algorithm type, and it concludes that an online unsupervised ensemble method is the most suitable AAD configuration for resource-constrained CPS. A direct corollary the authors draw is that future work should target low detection latency, consistent metrics such as the Matthews correlation coefficient, adaptive thresholding, and protection against adversarial manipulation of the detectors themselves.","pith_inferences":["The review's own numbers imply that the 'rarely both' gap may be an evaluation-infrastructure problem: online evaluation is harder and less standardized, so methods that couple both components are less likely to be reported even if they exist.","The 99th-percentile citation threshold for snowballing likely biases the sample toward highly cited, mostly offline academic methods; a lower threshold would probably surface more operational online systems in smaller venues.","A testable next step would be to build a shared streaming benchmark from existing CPS datasets and measure detection latency along with accuracy, which would reveal whether the field's bottleneck is algorithmic or evaluative."],"forward_implications":["If the central finding is correct, the most productive design target for AAD in CPS is a detector that couples streaming data processing with online model updates rather than optimizing one side alone.","The taxonomy gives new researchers a way to locate any AAD method within a few minutes and see which combinations of application, paradigm, and adaptation strategy are unexplored.","Practitioners selecting an AAD method should expect that most published detectors are evaluated offline and may not meet low-latency deployment constraints.","The paper's recommendation of online unsupervised ensembles is a concrete candidate architecture for future CPS defenses, one that could detect unseen attacks without expensive labels.","Standardizing on metrics such as precision, recall, false positive and negative rates, and the Matthews correlation coefficient would make future AAD comparisons directly head-to-head."],"supporting_citations":[{"why":"Supplies the concept-drift, learning-mode, adaptation, and ensemble vocabulary used to tag every paper in the taxonomy.","marker":"[34]"},{"why":"Provides the data-stream and online-learning definitions that ground the review's notion of near real-time processing.","marker":"[89]"},{"why":"Provides the standalone systematic review process that structures the search, selection, and extraction.","marker":"[129]"},{"why":"Supplies the forward and backward snowballing guidelines used to expand the initial keyword search.","marker":"[130]"},{"why":"Supplies the anomaly detection definition and one-class classification baseline used in categorization.","marker":"[10]"},{"why":"Supplies the availability, integrity, and confidentiality attack taxonomy used to classify attack types.","marker":"[47]"}],"fun_headline_variants":["Most adaptive CPS detectors adapt data or model, not both","Systematic review of 65 CPS papers finds rare dual adaptation","Adaptive anomaly detection in CPS rarely updates data and model together","CPS adaptive detection: most adapt data OR model, few do both","Online unsupervised ensemble recommended for adaptive CPS detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings depend on the manually curated corpus—built from chosen keyword queries, a 99th-percentile citation cutoff for snowballing, and author-side quality screening—being representative of the AAD-in-CPS literature; the paper itself acknowledges this selection bias in Section VI-B.","fun_headline_variants_meta":{"raw":{"variants":["Most adaptive CPS detectors adapt data or model, not both","Systematic review of 65 CPS papers finds rare dual adaptation","Adaptive anomaly detection in CPS rarely updates data and model together","CPS adaptive detection: most adapt data OR model, few do both","Online unsupervised ensemble recommended for adaptive CPS detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001229,"raw_usage":{"total_tokens":5037,"prompt_tokens":918,"completion_tokens":4119,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":4036}},"tokens_in":534,"tokens_out":4119,"duration_ms":26928,"temperature":1.0,"reasoning_tokens":4036,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:20:05.325307+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the same review with a broader keyword set and a lower citation threshold, and count the share of included methods that implement both near real-time data processing and model adaptation; if that share is well above what the paper reports, the 'rarely both' finding fails.","supporting_citations":[{"cited_title":"A guide to conducting a standalone systematic literature review,","cited_arxiv_id":null,"evidence_quote":"Provides the standalone systematic review process that structures the search, selection, and extraction."},{"cited_title":"Guidelines for snowballing in systematic literature studies and a replication in software engineering,","cited_arxiv_id":null,"evidence_quote":"Supplies the forward and backward snowballing guidelines used to expand the initial keyword search."}],"review_version":1}