{"id":"71c29beb-b64d-4b62-8793-0972563612c1","arxiv_id":"2508.12428","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"A one-class GRU forecasting framework with adaptive residual windows and modified SHAP detects replayed false data injections in Purdue's PUR-1 reactor signals with 93% point-level accuracy and under 1% false positives.","lead":"This paper tests a four-module framework that uses a GRU forecaster, adaptive residual thresholds, modified SHAP, and hand-built rules to detect false data injection attacks on a university research reactor's digital signals. A generalist should care because it is a concrete, real-data demonstration of a passive defense against replay attacks on nuclear instrumentation, with source attribution, though the evaluation has significant limitations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Detection metrics are not prospectively validated: thresholds and rules are tuned on the same datasets used to report performance, and Table 6's accuracy is internally inconsistent with its precision/recall and the stated attack duration, so the 0.93/0.01 claim is not yet established.","rationale":"The framework is a plausible and well-described pipeline with real PUR-1 data and public code, so I do not think rejection is warranted. The central problem is that the evaluation protocol does not separate threshold selection from performance estimation. This is the same issue the reader identified, so I partially agree. It is load-bearing because every headline number in the abstract depends on thresholds whose selection procedure is not described. The Table 6 precision/recall/accuracy inconsistency adds a further, independent reason to treat the reported numbers cautiously: it suggests the metrics may be computed at different levels or with different positive labels than the text implies, so even the in-sample numbers are not auditable. A conditional verdict is appropriate: the approach is reasonable and the paper is transparent about many implementation details, but the 0.93/0.01 performance claim should be re-evaluated with a held-out threshold-selection protocol and a clear per-second confusion matrix before it can be accepted.","tokens_in":16931,"tokens_out":7627,"duration_ms":85117,"concrete_test":"Recompute Table 6 with a strict per-second labeling where positives are only the injected intervals (e.g., seconds 120-190 of each FDI event) and select thresholds using only a held-out subset of events (e.g., tune on 60% of attack and normal events, test on the remaining 40%). Then verify that accuracy = (TP+TN)/(TP+TN+FP+FN), precision, recall, and F1 all come from the same confusion matrix. If held-out accuracy falls below 0.93, or if the reported accuracy cannot be reproduced from the precision/recall and positive fraction, the central detection claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is that the reported detection accuracy is an in-sample number. The adaptive-window thresholds (0.07 over 5 s and 0.04 over 60 s) and the rule thresholds (0.7 cm, 2%/s, 0.75 std) are described as 'optimized' and 'developed from experience' but are selected and evaluated on the same six datasets, so Table 6 measures fit to the test set rather than prospective performance. The paper does not document a validation fold for threshold selection, a cross-event split, or an independent attack dataset. Additionally, Table 6 contains an internal inconsistency that makes the headline metric ambiguous: for the FDI rows, precision=1, recall=0.8066, and F1=0.8929 are mutually consistent, but the reported accuracy of 0.9348 implies about one third of the dataset is positive, whereas the stated attack construction (120 s normal + 70 s injected, plus 10 s transition, in 1560 s datasets) gives only a few percent positive seconds. The text also says false positives occur at FDI initiation, which contradicts precision=1. Without a precise statement of the ground-truth labeling (point-level vs event-level) and a recomputation, the >0.93 accuracy and <0.01 FPR cannot be verified. The origin-identification claim is also only qualitative (Figure 14, Table 7) and tied to the specific false-scram scenario.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript describes a one-class explainable AI framework for detecting replay-based false data injection (FDI) attacks in nuclear reactor sensor signals. The framework consists of a GRU forecasting model trained only on normal operation data, a residual analysis module with short- and medium-term rolling-average thresholds, a modified WindowSHAP module for feature attribution, and rule-based correlation checks. The authors evaluate on data from the PUR-1 research reactor, using a normal dataset and five evaluation datasets (transients, scrams, and three FDI scenarios). They report over 93% accuracy with under 1% false positives, successful differentiation of FDIs from normal transients and scrams, and identification of the falsified signals' origin.","tokens_in":17261,"tokens_out":7557,"duration_ms":78948,"significance":"If the empirical results are correct, this work would provide a practical, passive defense against replay FDI in digital instrumentation and control systems, with the notable strengths of one-class training (no attack data required), real-world PUR-1 data, explainability via a nonstationary adaptation of SHAP, and publicly available code and data. The combination of forecasting residuals with two complementary time windows is a sensible detection strategy, and the rule-based correlation layer offers an interpretable complement. However, the current evidence does not establish the headline accuracy and false-positive claims because the evaluation is in-sample, the metrics in Table 6 are internally inconsistent, and the source-identification result is only qualitative.","major_comments":[{"comment":"The reported detection performance is an in-sample estimate. The GRU was selected as the primary model based on its performance on the evaluation datasets (Table 3, which includes the same datasets later used for accuracy reporting), and the adaptive-window thresholds (error threshold 0.07 over a 5-second window and 0.04 over a 60-second window) are described as 'optimized' without any held-out validation split or independent tuning procedure. Similarly, the rule thresholds (0.7 cm, 2%/s, 0.75 std) are said to be 'developed from experience with PUR-1 data' and are applied to the same datasets. A proper protocol is needed: a validation fold for threshold selection, a cross-event split, or pre-specified thresholds, before the >0.93 accuracy and <0.01 false-positive claims can be accepted.","section":"Model Selection and Results (Table 6)"},{"comment":"The numbers in Table 6 are internally inconsistent. With Precision=1 and Recall=0.8066 on the FDI datasets, the Accuracy value of 0.9348 implies that approximately 33.7% of the dataset is positively labeled (because accuracy = 1 - (1 - Recall) * prevalence when false positives are zero). However, the attack design described in the data-collection section is 120 seconds of normal operation, a 10-second transition, and 70 seconds of injected shutdown data in 1560-second datasets, i.e., about 5% positive seconds. In addition, the text states that false positives usually occur at the initiation of the FDI, which contradicts Precision=1. The ground-truth labeling (point-level, event-level, or windowed) must be stated precisely and all metrics recomputed, because Table 6 as reported cannot be verified.","section":"Table 6 and attack construction (Data collection)"},{"comment":"The source-identification and FDI-differentiation claims are not quantitatively evaluated. Table 7 reports the fraction of anomaly points that break each correlation rule, but this is not a sensor-level attribution accuracy or a confusion matrix for FDI versus non-FDI. Moreover, the rules were developed from experience with PUR-1 data and target exactly the inconsistencies that the authors inserted into the FDI datasets (e.g., neutron change rate deviating from hand-calculated change rate, control rod position changes inconsistent with active states), so Table 7 partly reflects the construction of the attack data rather than an independent test. The authors should provide a held-out attack scenario or an explicit evaluation of attribution accuracy, and should temper the claim that the framework 'identifies the origin of the falsified signals' accordingly.","section":"Rule-based correlations and Table 7"},{"comment":"The data accounting is inconsistent. The text states that 265,000 seconds were collected, with 200,000 seconds used for the normal dataset and 'the remaining 65,000 datapoints' used for additional datasets. Table 1, however, lists 19,800 + 1,560 + 1,560 + 1,560 + 1,560 = 26,040 points for datasets #2 through #6, which does not match 65,000 or the total of 265,000. This discrepancy affects all per-second rate computations and must be resolved.","section":"Data collection and Table 1"}],"minor_comments":[{"comment":"The caption states 'SHAP contribution for SS2 position in dataset where it is falsified FDI-A,' but in FDI-A only neutron counts are falsified; the reference should likely be to FDI-C or be corrected.","section":"Figure 13 caption"},{"comment":"The column headers 'FDI #1' and 'FDI #2' do not match the dataset names FDI-A and FDI-B used elsewhere; please align the terminology.","section":"Table 3"},{"comment":"The definition of successful FDI detection gives epsilon < ||Delta X(t)|| <= epsilon_H, where epsilon_H is defined as the minimum deviation required to cause harm; this appears to be a typo, since detection before harm should require ||Delta X(t)|| < epsilon_H. Please clarify.","section":"Module 2, Definition"},{"comment":"The reference 'Lawson-Jenkins, n.d.' is incomplete; please provide a full citation with year and source.","section":"References"},{"comment":"Several equations contain Unicode artifacts (e.g., subscript characters rendered as Latin letters) that will not typeset correctly; please ensure all mathematics uses proper notation.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The data-count discrepancy and the Table 6 arithmetic suggest the authors may be using a ground-truth labeling different from the one described in the text. It would be worth asking them to provide the exact labeling and a recomputed confusion matrix. The in-sample threshold selection is the most serious issue, as it directly affects the central quantitative claim; a proper validation split should be required."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what you should know. This is a real-data replay-attack detection study on PUR-1, a licensed research reactor, with code and data on GitHub. The pieces are familiar—GRU forecasting, residual thresholds, WindowSHAP, hand-built rules—but the combination, applied to nonstationary multivariate reactor signals, is new. The one-class training and the inclusion of scrams and transients as normal variation are the right choices.\n\nThe paper does several things well. It trains only on normal data, which is the correct setup for detecting unseen attacks. It evaluates on separate normal datasets and reports low false positives there. The implementation is open and the writing is clear.\n\nThe problem is that the headline numbers are not yet trustworthy. The 0.93 accuracy / <0.01 FPR are reported on the same datasets used to tune the thresholds. The window thresholds are 'optimized' and the rule thresholds are 'developed from experience with PUR-1 data'; there is no held-out fold or independent attack set. So the metrics describe fit to this test set, not prospective performance.\n\nTable 6 also has an internal inconsistency. Precision=1, recall=0.8066, and F1=0.8929 are consistent with each other, but accuracy=0.9348 is not consistent with an attack that lasts ~70 seconds out of 1560. With those precision/recall values, accuracy of 0.9348 implies over 30% of the dataset is positive, not the ~5% from the attack construction. Either the accuracy or the ground-truth labeling is wrong. The text also mentions false positives at FDI initiation, which contradicts precision=1. This needs to be fixed with a clear definition of point-level vs event-level labels.\n\nThe source-identification claim is qualitative. Figure 14 shows patterns, but there is no quantitative metric or baseline comparison. Table 7 is suggestive, but the rules explicitly flag the exact inconsistencies inserted into the attack datasets, so part of that power is built in.\n\nThe data description also says 65,000 datapoints were reserved for evaluation, but Table 1 sums to 26,040.\n\nWho should read this: anyone working on FDI detection or nuclear cybersecurity. It is a useful testbed demonstration, but not yet a validated detector. It deserves a serious referee, not a desk reject. I'd recommend major revision: separate threshold selection from evaluation, recompute Table 6 with a clear ground-truth definition, and add a quantitative source-identification evaluation.","headline":"Real-data replay attack demo on a licensed reactor, but in-sample threshold tuning and an inconsistent Table 6 mean the headline accuracy is not yet established.","tokens_in":17819,"tokens_out":3876,"would_cite":false,"duration_ms":38351,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural forecaster trained only on normal reactor data, with dual residual windows, detects replayed false data injections at over 93% accuracy and under 1% false positives.","keywords":["false data injection","replay attack","nuclear reactor cyber security","anomaly detection","gated recurrent unit","explainable AI","SHAP","residual analysis"],"falsifier":"Split the 65,000 seconds of non-normal recordings into a tuning set and a held-out test set, tune the two residual thresholds and the three rule thresholds on the tuning set only, then measure accuracy and false-positive rate on the held-out set; if the held-out numbers fall below 0.93 accuracy or above 0.01 false positives, the claimed generalization is not established.","tokens_in":16733,"feed_emoji":"☢️","tokens_out":9794,"duration_ms":89219,"temperature":0.7,"pith_summary":"The paper claims that a monitoring system built from a gated recurrent unit (GRU), trained exclusively on normal operating data from a research reactor, can detect replay-based false data injections in real time. The detection signal is a forecast residual: the GRU predicts neutron counts five seconds ahead, and a point is flagged when either a five-second or sixty-second rolling average of the prediction error crosses a threshold. The paper further claims that this windowed residual analysis distinguishes injected false scrams from genuine scrams and fast transients, and that a modified SHAP analysis plus rule-based correlations identifies which sensors were falsified. On test recordings of one-, two-, and five-signal replay attacks, the reported accuracy is above 0.93 with fewer than 1% false positives. If these claims hold, a passive monitor could be layered onto digital reactor instrumentation without watermarking or modifying process data.","feed_headline":"Replay attacks on reactor signals caught at 93% accuracy","feed_subtitle":"One-class AI trained only on normal data flags injected sensor readings with under 1% false positives.","key_machinery":"The load-bearing mechanism is a two-time-scale residual check on a one-class forecaster. A GRU—a gated recurrent unit, chosen after comparing ANN, RNN, and LSTM variants—maps a 30-second window of eight reactor signals to a forecast of the neutron count five seconds ahead; because it is trained only on normal data, its error rises when injected signals violate the learned relationships among neutron counts, neutron change rate, and control-rod positions. The residual check labels a point anomalous if either the 5-second rolling average of absolute error exceeds 0.07 or the 60-second rolling average exceeds 0.04, so abrupt injections trip the short window and subtle or prolonged ones accumulate in the long window. Around this core sit the interpretability modules: a modified WindowSHAP occlusion scheme that replaces occluded values with a moving baseline (current power level, zero change rate, zero rod motion) rather than a global mean, and rule-based correlations that flag inconsistencies between rod position and rod active state, between neutron counts and reported change rate, and between change-rate variance and rod motion.","core_discovery":"Working with 265,000 seconds of recorded operations from a fully digital research reactor, the authors build a one-class predictive model: a GRU that takes a sliding 30-second window of eight reactor signals and forecasts the neutron count five seconds ahead, trained only on 200,000 seconds of normal operation. The evaluation sets include three replay-attack scenarios that falsify one, two, or five signals to make the reactor appear to be scramming while it is actually still at power, alongside genuine scram and fast-transient recordings. The central finding is that a dual-window residual check—flagging any point whose five-second mean absolute error exceeds 0.07 or whose sixty-second mean absolute error exceeds 0.04—achieves over 93% accuracy on the attack datasets and under 1% false positives on the normal, transient, and scram datasets. The paper also reports that the modified SHAP explanations show falsified signals carrying negative contributions while unaffected signals stay near zero, and that rule-based correlations between control-rod motion, rod position, neutron count, and reported change rate break consistently for injected data.","pith_inferences":["Extension beyond the paper: the threshold values (0.07 over 5 s, 0.04 over 60 s) and the rule cutoffs (0.7 cm, 2%/s, 0.75 standard deviation) were selected and evaluated on the same recordings, so a true out-of-sample test would tune them on a disjoint subset before scoring.","Extension beyond the paper: the correlation rules encode how control-rod motion drives neutron population in this particular reactor; porting the framework to a plant with different control mechanisms would require rewriting those rules, though the residual and SHAP components would transfer.","Extension beyond the paper: the replay episodes tested last 70 seconds and start from genuine operation; shorter, smoother, or longer injection profiles may be harder or easier to catch, and the dual-window thresholds would need retuning per profile.","Extension beyond the paper: combining the SHAP attribution with a physics-based state estimator could extend this passive approach to other integrity attacks such as sensor drift or scaling attacks, which the paper names as future work."],"forward_implications":["An operator-facing monitor of this type could run on existing sensor streams in a digital reactor control room, adding a passive detection layer with no watermarking or control perturbation.","Because the forecaster trains on normal operation only, the method does not require labeled attack data; a new plant could deploy it after collecting routine operating logs.","The dual-window residual design gives two detection speeds: near-immediate response to abrupt injections and slower, accumulating evidence against gradual or subtle ones.","The SHAP and rule-based layers would let operators see not only that an alarm fired but which sensor channels are implicated, supporting faster diagnosis and shorter outage."],"supporting_citations":[{"why":"Proves that a standard χ2 failure detector on a linear time-invariant state estimator cannot asymptotically detect replay attacks, defining the evasion problem this framework targets.","marker":"Mo and Sinopoli (2009)"},{"why":"Adds a second χ2 detector to distinguish replay attacks from other anomalies, the classification task this paper extends with rules and SHAP.","marker":"Zhao and Smidts (2020)"},{"why":"Provides the comparative evaluation of time-series anomaly detection methods that motivates the forecasting-residual approach.","marker":"Schmidl et al. (2022)"},{"why":"Introduces LSTM, establishing the recurrent memory architecture that the paper's GRU competes against.","marker":"Hochreiter and Schmidhuber (1997)"},{"why":"Introduces the GRU architecture selected as the forecasting module after hyperparameter tuning.","marker":"Cho et al. (2014)"},{"why":"Supplies the Shapley-value explanation framework that the paper modifies with a moving baseline for time-series data.","marker":"Lundberg and Lee (2017)"},{"why":"Provides WindowSHAP, the windowed Shapley approximation adapted here for non-stationary reactor signals.","marker":"Nayebi et al. (2023)"},{"why":"Demonstrates SHAP-based explainability for nuclear power plant anomaly diagnosis, the interpretive template this paper follows.","marker":"Park et al. (2022)"}],"fun_headline_variants":["One-class AI flags reactor replay attacks with 93% accuracy","AI trained on normal data catches 93% of reactor signal spoofs","Fake reactor signals exposed: 93% detection, <1% false alarms","Neural net spots replay attacks in reactor data with 93% accuracy","Reactor replay attacks caught at 93% with under 1% false positives"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The detection thresholds and rule parameters were chosen and measured on the same datasets, so the reported accuracy and false-positive rates may not hold once the thresholds are fixed before seeing attack data.","fun_headline_variants_meta":{"raw":{"variants":["One-class AI flags reactor replay attacks with 93% accuracy","AI trained on normal data catches 93% of reactor signal spoofs","Fake reactor signals exposed: 93% detection, <1% false alarms","Neural net spots replay attacks in reactor data with 93% accuracy","Reactor replay attacks caught at 93% with under 1% false positives"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000427,"raw_usage":{"total_tokens":2229,"prompt_tokens":1031,"completion_tokens":1198,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":1099}},"tokens_in":647,"tokens_out":1198,"duration_ms":8645,"temperature":1.0,"reasoning_tokens":1099,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:22:16.092639+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Split the 65,000 seconds of non-normal recordings into a tuning set and a held-out test set, tune the two residual thresholds and the three rule thresholds on the tuning set only, then measure accuracy and false-positive rate on the held-out set; if the held-out numbers fall below 0.93 accuracy or above 0.01 false positives, the claimed generalization is not established.","supporting_citations":[],"review_version":1}