{"id":"98ec49ae-c2a0-42b4-b30e-4a388bfec5f5","arxiv_id":"2608.08280","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of ML/RL/QML applications in five specialized QKD areas, with a Tier I/II/III framework for how close learned components may sit to the security proof.","lead":"This paper surveys how machine learning, reinforcement learning, and quantum machine learning are being applied to specialized quantum key distribution (QKD) settings, such as satellite links, drones, 6G networks, and one-sided device-independent security. It organizes the work into five themes and argues that the safest ML roles sit above or beside the quantum security proof, while inside-proof uses remain unproven.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's quantitative backbone is unreliable: most headline gains are simulation-only or self-cited, and the QLSTM result is internally inconsistent (93.7% vs 94.7%), so the claim of 'clearest, lowest-risk wins' is not yet supported.","rationale":"Agreeing with the reader's weakest assumption, the most load-bearing point is the reliability of the quantitative evidence underlying the survey's synthesis. The paper's own Table XX acknowledges that most headline results are simulation-only, and the reference list's 'vERIFIED' annotations are not independent verification. The QLSTM discrepancy (93.7% vs 94.7%) is a concrete, checkable instance of the problem; it shows the survey can cite inconsistent numbers for the same work without comment. Because Section XI converts these numbers into a risk-tier map, an unreliable number directly undermines the central claim about where ML delivers 'clearest, lowest-risk wins.' The concern does not overturn the survey's structural insight—that above-proof and beside-proof roles are lower risk than inside-proof certification—but it does mean the claim of substantial, quantifiable improvements is conditional on verification. The reader's CONDITIONAL verdict is therefore appropriate; no change needed. The one additional insight is to treat the QLSTM inconsistency and the 'vERIFIED' annotations as required fixes before the survey can be relied upon.","tokens_in":40479,"tokens_out":8555,"duration_ms":75509,"concrete_test":"Perform a provenance audit of the 15 quantitative headline claims in Tables XVI and XVIII: for each, record (a) simulation vs. experimental data, (b) whether the cited source overlaps with the survey authors' prior works, and (c) whether the stated number matches the cited source. Start with the QLSTM claim by retrieving [114] and [159] and comparing accuracy and attack-class lists; if the two versions disagree or the number is not reproducible, the survey's quantitative synthesis is not reliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ML/RL/QML yields substantial, quantifiable improvements in five specialized QKD areas depends on the reported gains being accurate and representative. The paper itself, in Table XX, shows most of those gains come from simulation-only evaluations, without link-wise splitting or uncertainty reporting; Section XIII even lists these as common pitfalls, yet the surveyed results are not screened against them. A concrete internal inconsistency exposes the fragility: Section VIII-C and Table XIV report ~93.7% accuracy over five attack types for the hybrid QLSTM, citing the IET version [114], while the arXiv version [159] of the same work is annotated with 94.7% accuracy over four attack types. Two different numbers appear for one result, with no reconciliation, in a reference list that also carries nonstandard 'vERIFIED'/'CONFIRM authors' annotations. Since Section XI synthesizes the quantitative gains into the risk-tier conclusion, any inflated or non-reproducible number changes the map of where learning is ready to deploy. The survey's own prior works (e.g., [37], [38], [51]–[56]) contribute several headline rows; their independence is not established. Without independent verification or a documented provenance audit, the evidence base for the paper's central claim remains unsecured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews machine learning (ML), reinforcement learning (RL), and quantum machine learning (QML) applied to specialized and emerging QKD scenarios beyond conventional point-to-point fiber links, organized into five thematic pillars: adaptive protocol and parameter support; free-space, satellite, UAV, and HAP-assisted QKD; QKD for IoT, 6G, and quantum-secured federated learning; QML-assisted QKD functions; and steerability-aware, one-sided-device-independent QKD security estimation. For each theme it provides a problem/conventional-solution/ML-solution structure, per-theme and cross-theme comparison tables, and consolidated quantitative gains. It also proposes a three-tier risk classification (above-proof, beside-proof, inside-proof) for placing learned components relative to QKD security proofs, and identifies open challenges including dataset scarcity, generalization, interpretability, and trustworthy QML. The paper's central claim is that learning delivers its clearest, lowest-risk wins when it improves an estimate or decision supporting adaptive, non-terrestrial, or application-driven QKD without touching the security proof directly.","tokens_in":40700,"tokens_out":2240,"duration_ms":23107,"significance":"If the survey's conclusions are reliable, it provides a valuable and much-needed map of where ML/RL/QML can be safely and effectively deployed in specialized QKD settings, distinguishing low-risk decision support from security-critical certification. Its strengths include a clean five-theme taxonomy, a consistent per-theme structure, a useful Tier I/II/III security-placement framework, and an unusually honest evaluation-quality critique in Section XIII and Table XX that acknowledges most surveyed results are simulation-only and lack link-wise splitting or uncertainty reporting. The paper also explicitly identifies open problems (OP-1 to OP-10) and argues for open datasets and standardized benchmarks, which are constructive contributions to the community. However, the quantitative evidence underpinning the central claim is presented largely at face value from heterogeneous sources, several of which are the authors' own prior works, and at least one reported result is internally inconsistent; this currently limits the confidence with which the survey's synthesized conclusions can be accepted as a reliable guide.","major_comments":[{"comment":"The QLSTM result is reported inconsistently: Section VIII-C, Table XIV, and Section XI-D state '93.7% accuracy over five attack types' citing the IET version [114], while the arXiv version [159] of the same work is annotated in the reference list as 94.7% accuracy over four attack types. Because this number is used as a headline quantitative gain for Theme IV and feeds directly into the paper's conclusion about QML-assisted attack detection, the inconsistency is load-bearing and must be reconciled or explicitly flagged as an unresolved discrepancy before the survey's quantitative claims can be considered reliable.","section":"Section VIII-C and Table XIV vs reference [159]"},{"comment":"The paper identifies in Section XIII common evaluation pitfalls—leakage-free splitting, class imbalance, uncertainty and calibration—and Table XX scores representative works as mostly simulation-only, without link-wise splitting, without uncertainty reporting, and without field validation. Yet Sections X and XI present the same headline gains (e.g., >98% protocol-selection accuracy, 93.7% QLSTM detection, ~0.96 steerability regression, orders-of-magnitude speedups) at face value, without screening or qualifying them against these criteria. The survey should apply its own evaluation-quality criteria when reporting each headline gain, or clearly state that these numbers are unvalidated literature claims, so that the risk-tier conclusions are not built on unsecured quantitative ground.","section":"Section XIII vs Sections X-XI"},{"comment":"The claim that ML-assisted LLO phase recovery gave 'a roughly fourfold improvement in distance over the 25 km commercial baseline' compares the 100 km result of [108] with 25 km systems from [164], [165] that differ in many hardware aspects beyond the use of ML. This improvement is not attributable specifically to the learned component, and the comparison conflates technological progress over roughly fifteen years with the ML contribution. The statement should be reworded to report the demonstrated reach of [108] and [126] without attributing the distance gain to ML unless a controlled comparison exists.","section":"Section XI-A (CV-QKD reach comparison)"},{"comment":"A substantial number of the headline rows in Table XVI and the consolidated tables come from the authors' own works (e.g., [37], [38], [39], [51], [52], [53], [54], [55], [56]), and their independence is not assessed or discussed. Since the survey's central synthesis claims that specific ML roles are 'ready to deploy,' the manuscript should include a provenance or self-citation disclosure, ideally marking which quantitative entries are from the authors' own papers versus independently replicated results, and discuss any potential bias in the conclusions that rely on those entries.","section":"Reference list and Table XVI"}],"minor_comments":[{"comment":"The reference list contains nonstandard annotations such as 'vERIFIED' and 'CONFIRM authors' (e.g., [33], [108], [114], [122], [142], [149], [159], [189]). These appear to be internal verification notes and should be removed or converted into a standard editorial footnote, as they are not part of a formal reference entry.","section":"Reference list"},{"comment":"OptiQKD appears twice with the same arXiv identifier: as [118] in Section V-C and as [120] in Section VII-A, with slightly different descriptions. The duplicate reference should be merged and cited consistently.","section":"Section V-C and Section VII-A"},{"comment":"The sentence 'the hybrid QLSTM raises the bar to ~93.7% over a harder five-class problem spanning unknown attack types' is unclear because the QLSTM result concerns five known attack classes, not necessarily unknown attack types; please clarify the relationship to the DBSCAN-based unknown-attack detection reported in [157].","section":"Section XI-B"},{"comment":"Equation (1) writes the asymptotic secret fraction with 'r' while the surrounding text and Eq. (2) use 'R' for the key rate; unify the notation for readability.","section":"Section II-E"},{"comment":"The key-assignment optimization problem in Eq. (5) would benefit from a brief definition of the utility function u_k and the path set P_k, which are introduced only implicitly in the surrounding text.","section":"Section VII-B, Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a solid organizational contribution and an unusually honest evaluation-quality section, but the quantitative backbone needs an explicit provenance audit and reconciliation of the QLSTM inconsistency before the survey can serve as a reliable reference. The heavy reliance on the authors' own prior works for several headline rows is a concern that should be addressed transparently. I recommend major revision rather than rejection because the central tier-based framework is defensible and the issues identified are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a real reference survey, not a padded rehash. The five-theme taxonomy (adaptive support, non-terrestrial links, IoT/6G/FL, QML, steerability) is sensible, and the Tier I/II/III split—above, beside, inside the security proof—is genuinely useful. It gives practitioners a clear way to think about where ML is safe. The consolidated tables and ten-item roadmap are also legitimately helpful.\n\nWhere it earns credit: the paper does not oversell QML. Section VIII explicitly says no hardware-validated advantage exists, and Section XIII/Table XX openly admit most surveyed results are simulation-only with no link-wise splitting. That kind of self-criticism in a survey is rare and should be acknowledged.\n\nSoft spots, in order of importance. First, the quantitative backbone is not solid enough to carry the weight the conclusion puts on it. Most headline gains are from simulation-only studies, several are the authors' own prior works, and the survey does not screen the numbers against its own evaluation-quality criteria. The QLSTM inconsistency is a concrete symptom: 93.7% accuracy over five attack types in Section VIII-C/Table XIV citing [114], but 94.7% over four attack types in [159], with no reconciliation. One number is wrong or the source descriptions differ; either way, a reader cannot tell which number is the real reported result.\n\nSecond, self-citation concentration is real. The free-space/HAP and IoT-attack themes lean heavily on the authors' own papers. Self-citation is not disqualifying—the topics are their specialty—but the paper should disclose it and, where possible, mark which headline rows are independent.\n\nThird, the 'vERIFIED' and 'CONFIRM authors' annotations in the reference list are nonstandard and unexplained. They look like incomplete provenance notes and should either be removed or turned into a proper reproducibility appendix.\n\nDoes the central argument hold? Mostly. The qualitative claim—lowest-risk wins come from improving estimates and decisions without touching the security proof—is well supported by the tier analysis and does not depend on any single number. That is the durable takeaway.\n\nWho is this for? Practitioners entering non-terrestrial or application-driven QKD, and referees needing a map of the area. It deserves serious peer review; the taxonomy and roadmap are worth publishing after the quantitative claims are audited and the provenance issues fixed.","headline":"Useful five-theme map and risk-tier framing, but the reported quantitative gains need a provenance pass before the survey's stronger conclusions can be trusted.","tokens_in":41232,"tokens_out":1936,"would_cite":true,"duration_ms":18573,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Dd"],"model":"deepseek-v4-flash","headline":"This survey claims that machine learning, reinforcement learning, and quantum machine learning improve five specialized QKD areas, and that their lowest-risk, most reliable role is improving estimates or decisions that sit beside or above…","keywords":["quantum key distribution","machine learning","reinforcement learning","quantum machine learning","adaptive protocols","free-space optics","satellite QKD","quantum steering"],"falsifier":"Take the highest-profile classification claims, the protocol-selection accuracy above 98% and the steerable-weight regression near 0.96, and re-run them with leave-one-link-out or leave-one-device-out splitting on the original simulated data; if accuracy drops to near chance or well below the reported figures, the survey's picture of where ML helps most would be substantially overstated.","tokens_in":1682,"feed_emoji":"🔐","tokens_out":2344,"duration_ms":67002,"temperature":0.7,"pith_summary":"The survey tries to establish which roles machine learning, reinforcement learning, and quantum machine learning can usefully play in five specialized QKD settings: adaptive protocol and parameter support; free-space, satellite, UAV, and high-altitude-platform links; QKD for IoT, 6G, and quantum-secured federated learning; QML-assisted QKD functions; and steerability-aware one-sided device-independent security. It argues that in every theme learning pays off most reliably when it improves an estimate or decision that supports the QKD system, such as predicting a channel, choosing a protocol, tracking a phase, or pricing a key rate, without being asked to certify security on its own. If that is right, practitioners get a usable map of where ML is safe to deploy today, where it is promising but unproven, and where it must stay out of the proof. A sympathetic reader would care because non-terrestrial and application-driven QKD are exactly the regimes where channel conditions move too fast for static optimization and where a learned component must not weaken the security guarantee.","feed_headline":"ML's safe role in QKD is decision support, not security proofs","feed_subtitle":"Across adaptive, free-space, 6G, and steerability-aware QKD, learning wins where it sharpens estimates without entering the proof.","key_machinery":"The organizing machinery is the five-theme taxonomy plus a three-tier placement rule that classifies every learned component by its relation to the security proof: above the proof for protocol selection, routing, resource allocation, and key assignment; beside the proof for phase, polarization, channel, and key-rate estimators that feed a proven rate formula; and inside the proof for learned detectors and steerability claims that would certify secrecy. The survey's argument runs on this placement rule, treating the same random-forest or neural-network tool as safe in the first two tiers and dangerous in the third unless its output is a certified conservative lower bound. The five themes are adaptive protocol and parameter support; free-space, satellite, UAV, and HAP-assisted QKD; QKD for IoT, 6G, and quantum-secured federated learning; QML-assisted QKD functions; and steerability-aware one-sided device-independent security.","core_discovery":"The paper's central claim is a stratification of ML's role in specialized QKD. In five thematic areas it identifies a consistent pattern: learned components deliver large speedups and high classification or regression accuracy, including protocol-selection accuracy above 98%, a five-class QLSTM attack-detection accuracy near 93.7%, Strehl-ratio prediction with mean absolute percentage error in the low single digits, steerable-weight regression accuracy near 0.96, and ML-assisted carrier recovery that extends local-local-oscillator CV-QKD to 100 km, when they act as surrogates or estimators feeding a proven rate formula or as optimizers above the proof. The same evidence is read as a warning: learned attack detectors, key-rate predictions, or steerability estimates used directly in a secrecy claim are high-risk unless they are built with provably conservative bounds, as the paper reads the composable excess-noise estimator as showing. The conclusion is that the clearest and lowest-risk wins sit beside or above the proof, not inside it.","pith_inferences":["A direct consequence the authors leave implicit is that the near-ceiling accuracy reported for protocol selection and steerability classification is likely inflated by simulation-only, non-link-split evaluation; re-running with proper splits would probably lower the scores while preserving the ranking of the three tiers.","The same placement rule could be exported to adjacent problems such as post-quantum cryptography migration or classical optical-network control, wherever a learned estimator feeds a certified margin.","A concrete next experiment suggested by the survey's own roadmap is an integrated pipeline that jointly trains phase recovery, reconciliation decoding, and parameter optimization toward composable secret bits per second, which the paper lists as an open problem without demonstrating it."],"forward_implications":["New QKD deployments can treat learned protocol selectors and RL schedulers as deployable efficiency layers today, because errors in those roles cost throughput, not secrecy.","Physical-layer learned estimators, such as ML phase recovery and SOP prediction, can extend reach and availability without weakening security, making them the most immediate candidates for field integration.","Learned components that claim to detect attacks or certify steerability should be used only as monitors, or with provably conservative bounds, until certification requirements are met.","Future research should prioritize open non-terrestrial datasets, link-wise evaluation splits, and composable conservative estimators over further accuracy chasing.","QML's role in QKD remains speculative: no hardware-validated advantage over classical ML exists, so the safest reading is to treat QML as a monitoring and optimization layer."],"supporting_citations":[{"why":"Supplies the neural-network parameter-prediction result that anchors Theme I's claim that a learned surrogate can replace slow decoy-state optimization with orders-of-magnitude speedup.","marker":"[33]"},{"why":"Provides the random-forest protocol-selection result above 98% accuracy that anchors the adaptive protocol-selection claim.","marker":"[121]"},{"why":"Provides the hybrid QLSTM five-class attack-detection result near 93.7% accuracy that anchors the QML-assisted attack-detection claim.","marker":"[114]"},{"why":"Provides the neural-network steerable-weight regression near 0.96 accuracy that anchors the steerability-surrogate claim for one-sided device-independent QKD.","marker":"[163]"},{"why":"Provides the ML-assisted carrier-recovery field result of 100 km LLO CV-QKD that anchors the physical-layer estimation gains.","marker":"[108]"},{"why":"Provides the composable excess-noise neural-network estimator that anchors the claim that conservative learned components can enter a finite-key security proof.","marker":"[130]"},{"why":"Provides the graph-attention and LSTM deep-RL key-provisioning result that anchors the Theme III claim that learning reduces keystore exhaustion in dense networks.","marker":"[149]"},{"why":"Provides the composable XOR-relay HAP QKD architecture used to anchor non-terrestrial, 6G-oriented deployment as a first-class specialized theme.","marker":"[37]"}],"fun_headline_variants":["ML sharpens QKD estimates, but leaves proofs to math","In QKD, ML is a fast estimator, not a security prover","Adaptive, free-space QKD: ML helps, but doesn't prove","For QKD, ML's wins are in estimation, not certification","QKD survey: ML accelerates links without replacing proofs"],"cache_read_input_tokens":43392,"weakest_assumption_plain":"The survey's conclusions depend on the reported quantitative gains in the primary literature being accurate and representative, even though most of them come from simulation-only evaluations with no link-wise train/test splitting and several are the authors' own prior works.","fun_headline_variants_meta":{"raw":{"variants":["ML sharpens QKD estimates, but leaves proofs to math","In QKD, ML is a fast estimator, not a security prover","Adaptive, free-space QKD: ML helps, but doesn't prove","For QKD, ML's wins are in estimation, not certification","QKD survey: ML accelerates links without replacing proofs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00034,"raw_usage":{"total_tokens":1939,"prompt_tokens":1076,"completion_tokens":863,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":692,"completion_tokens_details":{"reasoning_tokens":772}},"tokens_in":692,"tokens_out":863,"duration_ms":8454,"temperature":1.0,"reasoning_tokens":772,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:10:56.509873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the highest-profile classification claims, the protocol-selection accuracy above 98% and the steerable-weight regression near 0.96, and re-run them with leave-one-link-out or leave-one-device-out splitting on the original simulated data; if accuracy drops to near chance or well below the reported figures, the survey's picture of where ML helps most would be substantially overstated.","supporting_citations":[],"review_version":1}