{"id":"a4dca433-1c44-4901-ad83-323889f562e9","arxiv_id":"2602.21524","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":9,"one_line_summary":"A threat analysis claims quantum 'harvest now, decrypt later' attacks on nuclear OT systems succeed 8-78% of the time under current defenses, falling below 1% after full PQC migration.","lead":"This paper argues that nuclear power plants' 60-80 year lifecycles make them uniquely exposed to 'harvest now, decrypt later' quantum attacks, and claims attack success probabilities of 8-78% under current defenses. It proposes a forensics-first, post-quantum cryptography migration framework and seven validation criteria to reduce residual risk below 1%.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline success probabilities are arithmetic products of internally inconsistent expert-judgment priors; the 8–78% and <1% figures are not supported by evidence.","rationale":"The reader's weakest-assumption identification is exactly the load-bearing fault: the quantitative model is a product of unvalidated and internally conflicting expert-judgment priors. My independent check confirms the paper provides at least three incompatible prior triples for the same QUANTUMSCAR scenario, and the defense 'validation' section provides acceptance criteria rather than empirical results. The reader's REJECT verdict therefore remains appropriate; I do not see a need to alter it. In good faith, the paper does have non-quantitative value: the Purdue-level vulnerability mapping, MITRE ATT&CK extensions, and PQC migration guidance are plausible and worth discussion. But the central advertised contribution—precise attack success probabilities and residual-risk reductions—is not secure, and no machine-checked proof or reproducible dataset is supplied to rescue it.","tokens_in":29072,"tokens_out":3344,"duration_ms":27486,"concrete_test":"Produce a complete audit table of every probability triple in the manuscript: Table IV, Fig. 3, §V.B.1, Table XI, Fig. 5, §VI.B.1, and §VII.F. For each triple, recompute the product interval exactly as the paper defines it and compare with the stated overall success range. Then perform a one-at-a-time sensitivity analysis replacing each component with a neutral base rate (e.g., 0.5) and with literature-derived base rates for OT intrusions and ICS exploitation. If the three QUANTUMSCAR prior sets generate disjoint or inconsistent overall intervals, or if the headline 8–78% and <1% residual ranges are not reproduced unless one cherry-picks among the priors, the quantitative claims are unsupported and must be reported only as illustrative scenario exercises, not as validated risk estimates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim rests entirely on the conditional chain ℙ(Success)=ℙ(S1)·ℙ(S2|S1)·ℙ(S3|S1∩S2), used to produce the advertised 8–78% baseline success and <1% residual risk (Abstract; §V.B.1; §VI.B.1). Every factor is an author-assigned 'expert-judgment prior' or 'realistic probability assessment' with no elicitation protocol, dataset, or calibration. The paper is internally inconsistent about these priors for the same QUANTUMSCAR scenario: Table IV gives ℙbaseline=5–15%; Fig. 3's component ranges (0.4–0.7, 0.4–0.6, 0.3–0.5) multiply to roughly 5–21%; and §V.B.1 uses (0.85–0.98, 0.75–0.92, 0.55–0.75) to obtain 35–68%. These are not sensitivity variants of a single model; they are different models with different priors, and the abstract/conclusion select the highest values. The claimed reduction to 1–8% (SL-4) and <1% (PQC) is obtained by substituting smaller author-chosen ranges, while the 'operational validation' in §VII is a list of acceptance criteria (Table XIX), not measurements demonstrating the reduction. The qualitative HNDL/PQC-migration recommendation may survive, but the quantitative headline—the paper's claimed novel contribution—is not load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that nuclear power plants face a serious quantum-computing threat because their 60–80 year lifecycles exceed the projected arrival of cryptographically relevant quantum computers. It proposes a forensics-first framework, presents two attack scenarios (QUANTUMSCAR and QUANTUMDAWN), models their success probabilities using a conditional-probability chain, and claims that ISA/IEC 62443 SL-4 plus full PQC migration reduces residual risk from 8–78% to below 1%. It also contributes six proposed MITRE ATT&CK for ICS technique extensions and a set of operational validation criteria.","tokens_in":29577,"tokens_out":3581,"duration_ms":36859,"significance":"If the quantitative results were well-supported, this would be a valuable prioritization tool for nuclear OT cybersecurity. The paper also contains useful, generally sound qualitative material: the Purdue-level vulnerability survey, the PQC performance and side-channel tables, and the defense-in-depth control objectives in Section VII are consistent with existing standards and practical migration guidance. However, the central quantitative contribution — the headline 8–78% success probabilities and the claimed reduction to <1% — rests entirely on author-assigned expert-judgment priors that are never justified with an elicitation protocol, data, or calibration. The same scenario receives three different probability sets in different parts of the paper, and the 'operational validation' is a list of acceptance criteria rather than a demonstration. These are load-bearing problems for the paper's stated novelty.","major_comments":[{"comment":"The QUANTUMSCAR scenario is assigned three mutually inconsistent probability sets. Table IV lists ℙ_baseline = 5–15%; Fig. 3 multiplies [0.4,0.7]×[0.4,0.6]×[0.3,0.5] ≈ 5–21%; and §V.B.1 uses [0.85,0.98]×[0.75,0.92]×[0.55,0.75] ≈ 35–68%. The abstract and conclusion adopt only the highest set. This is not a sensitivity range for a single model; it is a switch between different models with different priors, and the paper gives no criterion for preferring one over the other. Because these numbers are the paper's central quantitative claim, the inconsistency is fatal to that claim.","section":"§V.B.1, Table IV, Fig. 3"},{"comment":"The same inconsistency appears for QUANTUMDAWN: Table XI states ℙ_baseline = 7–18%; Fig. 5 states 7–18%; while §VI.B.1 gives 8–34% for 'typical' and 17–50% for 'targeted' facilities. The 'Multiple Attempts' transformation further inflates the range to 41–88%. The paper does not explain whether these are alternative scenarios or alternative prior sets for the same scenario. The reader cannot reproduce the advertised 8–78% envelope or the <1% residual from any single, consistently stated input set.","section":"§VI.B.1, Table XI, Fig. 5"},{"comment":"The model ℙ(Success)=ℙ(S1)·ℙ(S2|S1)·ℙ(S3|S1∩S2) is a straightforward product of the selected phase probabilities. Every factor is an 'expert-judgment prior' or 'realistic probability assessment' with no documented elicitation procedure, no supporting dataset, and no sensitivity analysis over the prior ranges. The paper's quantitative conclusions are therefore arithmetic tautologies: changing the prior changes the output in exactly the way intended. Without a calibration basis, claims such as '8–78% under current defenses' and 'below 1% under PQC' are not supported by evidence.","section":"§V.B.1 and §VI.B.1 (conditional chain)"},{"comment":"The paper calls Section VII an 'operational validation' of the defense framework, but Table XIX provides only acceptance criteria (e.g., conformance rates, latency budgets, TVLA |t|<4.5). No measured telemetry, test results, or observed residuals are reported. The claim that full PQC migration and SL-4 reduce success from 8–78% to <1% is reasserted using substituted prior ranges (e.g., Fig. 3 SL-4 values), not derived from the validation tests. The criteria may be reasonable design requirements, but they do not validate the quantitative risk-reduction claim.","section":"§VII, Table XIX"}],"minor_comments":[{"comment":"Typo: 'ViRTUal' should be 'Virtual'. Also, the dagger footnote in Table II defines severity codes but the legend in the caption is somewhat cramped; the relationship between attack codes and severity is hard to parse on first reading.","section":"§III-B"},{"comment":"The six 'MITRE ATT&CK for ICS technique extensions' T1001–T1006 are proposed by the authors, not yet adopted by MITRE. The paper should say 'proposed extensions' rather than implying they are already part of a 'standardized vocabulary', which overstates their status.","section":"§V.C / §X (Tables X, XVII)"},{"comment":"The contribution list says Section VII provides 'seven quantitative tests demonstrating systematic risk reduction', but the tests are acceptance criteria, not demonstrations. The wording should be aligned with the actual content of Section VII to avoid confusion.","section":"§I.A, contribution 4"}],"recommendation":"reject","confidential_remarks":"The paper has a useful qualitative core — the Purdue-level attack survey, the PQC performance data, and the control objectives — but the advertised quantitative contribution is not reproducible or internally consistent, and the 'validation' does not validate. These are not cosmetic issues; they concern the central claim. If the authors reframe the paper as a qualitative threat analysis and defense checklist, a different contribution could be considered, but as submitted the quantitative headline is unsupported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing before you read it. The qualitative framing — HNDL plus forensic integrity as an operational safety requirement in long-lived nuclear OT — is genuinely useful, and the six MITRE ATT&CK for ICS extensions (T1001–T1006) are a reasonable start on a vocabulary the field lacks. But the paper's central quantitative claim (8–78% attack success, reduced to <1% with PQC) is not supported. It is the arithmetic product of author-chosen 'expert-judgment priors,' and the paper itself contains three mutually inconsistent probability sets for the same QUANTUMSCAR scenario: Table IV says 5–15% baseline; Figure 3's component ranges multiply to roughly 5–21%; Section V.B uses (0.85–0.98)×(0.75–0.92)×(0.55–0.75) to get 35–68%. The abstract and conclusion pick the highest numbers. The 'operational validation' in Section VII is a list of acceptance criteria (Table XIX), not measurements.\n\nWhat is actually new and worth credit: the forensics-first positioning, the level-by-level vulnerability analysis across the Purdue model, the concrete PQC migration guidance aligned with ISA/IEC 62443 and NIST, and the side-channel/performance tables. The lifecycle asymmetry argument (60–80 year assets vs. CRQC arrival) is well made and the threat is real.\n\nThe soft spot is not that probabilities are uncertain — that would be fine. It is that these numbers are presented with a false precision that the underlying inputs do not justify, and the paper contradicts itself about the inputs in different sections. The residual-risk reductions (1–8%, <1%) are obtained by replacing the priors with smaller ranges, not by any stronger evidence. So the quantitative headline is not load-bearing. The qualitative story is.\n\nWho this is for: regulators, plant operators, and OT security people who want a structured migration checklist and a plausible picture of what quantum-enabled attacks could look like. They should ignore the probabilities and treat the case studies as narratives.\n\nMy recommendation: engage with the paper — I'd send it to a serious referee rather than desk reject, because the topic is important and the qualitative framework is substantial. But the revision path is non-optional: the authors should either replace the invented priors with a single, fully documented model (with a real elicitation protocol and a consistent parameter set), or explicitly relabel the quantitative centerpiece as illustrative sensitivity analysis. As it stands, the internal inconsistency is the load-bearing flaw.","headline":"The qualitative HNDL/forensic framework for nuclear OT is useful, but the headline success probabilities are arithmetic products of internally inconsistent expert priors and should not be trusted.","tokens_in":30002,"tokens_out":3779,"would_cite":false,"duration_ms":69517,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that quantum-enabled attacks on nuclear power plants have a realistic success probability of 8–78% under current defenses, and that a defense-in-depth migration to post-quantum cryptography can drive residual feasibility b","keywords":["quantum computing threats","post-quantum cryptography","industrial control systems","nuclear power plant security","harvest-now decrypt-later","forensic integrity","probabilistic attack modeling","cryptographic diversity"],"falsifier":"Recompute the products using one documented set of phase probabilities. The paper itself gives 5–15% (Table IV), roughly 5–21% (Figure 3), and 35–68% (Section V.B) for the same QUANTUMSCAR baseline, so a reader can determine which numbers survive by checking which prior set is reproducible; if the credible priors are near the low ends, the upper bound of the claimed range collapses.","tokens_in":29006,"feed_emoji":"☢️","tokens_out":6479,"duration_ms":58156,"temperature":0.7,"pith_summary":"This paper tries to establish that nuclear power plants face a quantum threat unlike other infrastructure: their 60–80 year operating lives stretch past the expected arrival of cryptographically relevant quantum computers, so data and signatures harvested today become decryptable and forgeable within the asset's lifetime. It models two attack campaigns, QUANTUMSCAR and QUANTUMDAWN, as three-phase conditional chains and claims success probabilities of 8–78% under current defenses, 1–8% after reaching the top industrial security level with cryptographic diversity, and below 1% after full migration to post-quantum cryptography. It also argues that forensic integrity is operationally essential, not merely evidentiary, because quantum-forged but cryptographically valid evidence makes sabotage indistinguishable from equipment failure or operator error. A sympathetic reader would care because the paper gives concrete, quantified stakes to a migration that is often framed in abstract terms.","feed_headline":"Up to 78% success seen for quantum attacks on nuclear plants","feed_subtitle":"Harvested plant traffic becomes decryptable before a reactor's license expires; PQC migration cuts residual risk below 1%.","key_machinery":"The load-bearing object is the conditional probability chain ℙ(Success) = ℙ(S1) × ℙ(S2|S1) × ℙ(S3|S1∩S2), where S1 is the harvest-now collection phase, S2 is quantum weaponization (factoring RSA-2048 and forging certificates/signatures via Shor's algorithm), and S3 is execution plus forensic obfuscation. The paper feeds expert-judgment phase probabilities for baseline, SL-4, and full-PQC postures into this product to produce the 8–78%, 1–8%, and below-1% numbers. The chain also yields a sensitivity ordering, ∂P/∂S1 > ∂P/∂S2 > ∂P/∂S3, used to argue that early-phase defenses — disrupting collection and breaking cryptographic monoculture — are the highest-leverage interventions.","core_discovery":"The paper's central claim is that quantum-enabled attacks on nuclear OT are feasible within realistic timelines, quantified through two named scenarios: QUANTUMSCAR (35–68% success for typical deployments, 51–78% for targeted facilities) and QUANTUMDAWN (8–34% and 17–50% respectively). The mechanism is harvest-now, decrypt-later: adversaries collect encrypted traffic and signed firmware today, wait for a CRQC, apply Shor's algorithm to recover RSA/ECC private keys, and then both sabotage safety systems and forge valid-looking forensic evidence. The paper claims that a defense-in-depth migration — hybrid key exchange, post-quantum signatures for code and logs, authenticated time synchronizati","pith_inferences":["If the conditional-chain structure holds, the same three-phase model should transfer to other long-lived critical infrastructure — power grids, hydro dams, pipelines — where harvest-now risks compound over decades; nuclear plants are only the most extreme case because of their safety functions and evidence requirements.","The paper's forensic-paradox argument implies that cryptographic diversity is not just a security control but an evidence-integrity control: regulators may at some point require independent trust anchors across safety and control domains simply so that post-incident investigations can distinguish accident from sabotage.","A testable extension is to apply the framework to plants whose field protocols already use Grover-resistant symmetric authentication, such as DNP3-SA; the model would predict that the dominant risk shifts almost entirely to key-distribution infrastructure, which could be validated by comparing attack-success estimates against real-world key-management incident rates.","The harvest-now framing suggests a new metric for data-retention policy: any encrypted log or historian archive kept longer than the remaining time to CRQC arrival should be treated as potentially public, which would change how long nuclear operators retain sensitive operational data."],"forward_implications":["If the model is right, RSA/ECC-protected plant communications and firmware collected today are vulnerable to retroactive decryption and forgery within a reactor's operating life, so the risk exists even if no plant is ever 'hacked' in the classical sense before CRQCs arrive.","Current defensive postures are quantitatively insufficient: single-attempt success is modeled at 8–78%, and persistent adversaries pushing multiple attempts raise modeled success to 41–99% for targeted facilities.","Reaching the top industrial security level with cryptographic diversity, not necessarily full PQC replacement, cuts modeled success to 1–8%, giving operators an intermediate goal with a concrete risk reduction.","Full migration to hybrid key exchange with post-quantum signatures for code, logs, and time synchronization is modeled to drive residual feasibility below 1%, which would make quantum-enabled sabotage a minor tail risk rather than a primary one.","The six new technique identifiers (T1001–T1006) give defenders a shared vocabulary to detect and discuss quantum cryptanalysis, HNDL collection, forged-evidence manipulation, timing attacks, authenticated persistence, and certificate forgery in industrial control systems."],"fun_headline_variants":["Quantum attacks on nuclear plants: 78% success possible","Harvest-now quantum threat: 78% success on reactors","PQC migration cuts nuclear quantum attack risk to <1%","Nuclear plants: 78% quantum attack success, 1% after PQC"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline probabilities are products of three expert-judged phase-success rates — collection, quantum key-breaking, execution — that the paper asserts rather than measures; if those priors are wrong, every downstream number scales with them.","fun_headline_variants_meta":{"raw":{"variants":["Quantum attacks on nuclear plants: 78% success possible","Harvest-now quantum threat: 78% success on reactors","PQC migration cuts nuclear quantum attack risk to <1%","Nuclear plants: 78% quantum attack success, 1% after PQC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000648,"raw_usage":{"total_tokens":2867,"prompt_tokens":858,"completion_tokens":2009,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":1934}},"tokens_in":602,"tokens_out":2009,"duration_ms":13858,"temperature":1.0,"reasoning_tokens":1934,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T20:59:57.636021+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the products using one documented set of phase probabilities. The paper itself gives 5–15% (Table IV), roughly 5–21% (Figure 3), and 35–68% (Section V.B) for the same QUANTUMSCAR baseline, so a reader can determine which numbers survive by checking which prior set is reproducible; if the credible priors are near the low ends, the upper bound of the claimed range collapses.","supporting_citations":[],"review_version":1}