{"id":"752066e8-649d-44b0-b3fd-62a68ea40e3f","arxiv_id":"2412.17826","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that catalogs ML and DL algorithms for optical fiber, network, and wireless systems, with quantitative tables of reported gains and comparisons to conventional methods.","lead":"This paper surveys machine learning and deep learning work in optical communications, organizing about 282 prior studies across fiber, networking, and wireless settings. It maps algorithms to applications and compiles reported performance gains from the cited literature.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table entries are internally inconsistent (e.g., [104] 99.95% vs 95-99%; [131] off-domain), so the quantitative comparisons that anchor the survey's contribution are not reliable as printed.","rationale":"The strongest claim is that the survey gives readers an up-to-date, algorithm-centric map of 282 studies with quantitative/qualitative comparisons. That claim depends on the reliability of the tabulated entries and on the comparability of the numbers across rows. My stress test found concrete internal evidence that this condition is not met: [104] is reported as 95-99% accuracy in Table V but 99.95% in the text; [131] appears in the OFC table while the text frames it as an OCN fault-localization method; and [174] shows MAE 0.18 dB in Table VI but 0.05 dB in the text. These discrepancies directly affect the quantitative comparisons the survey promises. They do not necessarily mean the entire table set is wrong, but they put the burden on the authors to demonstrate that the tables are faithful transcriptions and that the numbers are comparable across heterogeneous setups. Because the reader's verdict already conditions acceptance on fixing editorial inconsistencies and stating scope, and because verifying versus original sources is part of that process, I do not change the verdict. If a full audit of the tables were to reveal many more such errors, a more severe verdict would be warranted.","tokens_in":44420,"tokens_out":6935,"duration_ms":67530,"concrete_test":"Check the specific discrepancy in Table V, ref [104]: retrieve the original paper and record its reported accuracy. If it is 99.95% (as the survey's text states) and not '95-99%', the table transcription is wrong. Extend this source-checking to a stratified random sample of 20 entries from Tables III–VIII, verifying each reported performance number and its setup (simulation vs experiment, modulation, rate, reach). If ≥2 of 20 entries are misreported or mis-categorized, the survey's numerical tables are not a trustworthy basis for the comparative conclusions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central value is its quantitative comparison tables (Tables III–VIII), which claim to summarize performance, complexity, and conditions for 282 studies. For these comparisons to support the paper's conclusions, the entries must be faithful and comparable. Internal evidence shows this condition fails. Table V, row [104], reports '95-99% accuracy' for QoT estimation, while Section III.A states the same work achieves '99.95%'. Table III includes [131], a linear-regression fault-localization study that Section III discusses under OCN, placing it in the wrong domain. Table VI, row [174], gives MAE of 0.18 dB in the table but the text reports 0.05 dB for the same result. These are not cosmetic; they alter the reported gains and the location of the work in the taxonomy. If such errors occur elsewhere, the 'quantitative/qualitative comparisons'—the survey's claimed novelty—cannot be relied on without a full audit.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys applications of machine learning (ML) and deep learning (DL) in three optical communication domains: optical fiber communication (OFC), optical communication networking (OCN), and optical wireless communication (OWC). It organizes roughly 282 cited works according to ML/DL algorithm families, provides schematic figures of each algorithm type, and presents six comparison tables (Tables III–VIII) that summarize performance, complexity, train/test data type, objective, input features, metrics, and adopted algorithms. The paper also reviews motivation, discusses challenges, and outlines future research directions.","tokens_in":44541,"tokens_out":4107,"duration_ms":42242,"significance":"If the comparison tables are faithful to the cited sources, this survey would provide a useful, up-to-date, algorithm-centric map of ML/DL in optical communications, with broader domain coverage than several prior surveys, particularly through its inclusion of OWC. The organizational contribution is real: grouping by algorithm rather than only by application helps readers compare methods. However, the survey's central value is explicitly claimed to be its quantitative/qualitative comparison tables, so the reliability of those tables is load-bearing. The internal inconsistencies identified below directly affect that reliability and must be resolved before the survey can serve its stated purpose.","major_comments":[{"comment":"The text states that the SVM-based QoT estimator in [104] achieves 99.95% accuracy in lightpath classification, while Table V reports 95-99% accuracy for the same reference. These numbers are mutually incompatible. Because Table V is one of the key quantitative comparison tables that the paper presents as a contribution, this discrepancy is not cosmetic. Please correct the entry and audit other table-text pairs for the same reference; for example, reference [133] is reported in Table V as detecting 95% of malicious nodes, while the text in Section III.A reports up to 90% accuracy for the K-means-based scheme.","section":"Section III.A / Table V, row [104]"},{"comment":"Reference [131] is listed in Table III as an ML application in OFC, with a fault-localization objective based on optical time-domain reflectometry measurements, but the text discusses the same work in Section III.A under OCN regression algorithms. This is a domain misplacement in the paper's central taxonomy. Since the survey's contribution includes a clear algorithm/domain classification, the placement must be corrected and other table entries checked for similar cross-domain inconsistencies.","section":"Table III, row [131] / Section III.A"},{"comment":"Table VI reports for reference [174] a mean absolute error of 0.18 dB for OSNR estimation, while the text in Section III.B states that the same reference achieved a MAE of 0.05 dB. Both cannot be correct for the same reported result. The discrepancy changes the stated performance gain and undermines confidence in the quantitative summaries; please reconcile the entry against the cited work.","section":"Table VI, row [174] / Section III.B"},{"comment":"Several branches of the taxonomy have empty reference slots, including PCA in OFC, policy-based and value-based RL in OFC, ensemble learning in OFC, hierarchical clustering in OCN, ICA in OCN, and PCA in OWC. If no works exist in these categories, the survey should state that explicitly; if works were omitted, the 'comprehensive' claim is weakened. As printed, the unexplained gaps leave the central organizational figure incomplete.","section":"Fig. 1"}],"minor_comments":[{"comment":"The heading 'Deap Learning' contains a typo; it should be 'Deep Learning'.","section":"Section II.B heading"},{"comment":"The paragraph beginning 'hese RL-based approaches...' is missing the initial T; it should be 'These RL-based approaches'.","section":"Section IV.A"},{"comment":"The text 'Receive Rewrard and New State' contains a typo; it should be 'Receive Reward and New State'.","section":"Fig. 13"},{"comment":"The phrase 'OFC, OCW, and OCN' uses 'OCW', which is not defined in the acronym list; this should be 'OWC'.","section":"Section I.C"},{"comment":"The acronym table lists both 'Principal Component PC' and 'Principal Component Analysis PCA'; this is redundant and potentially confusing.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a literature survey, so the standard verification burden applies to faithful summarization rather than to original derivations. The identified table-text inconsistencies are fixable in scope, but they are not isolated: they concern the very tables that the paper presents as its main contribution. I recommend that the authors audit every row of Tables III-VIII against the cited sources, not only the examples listed in this report, before the manuscript is reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The survey does one thing well: it organizes a large, scattered literature into a clear algorithm-by-application taxonomy, and it is one of the few reviews to give OWC full coverage alongside OFC and OCN. For someone entering this area, the structure and the 282-reference bibliography are genuinely useful. The tables that summarize reported performance, complexity, and data type per paper are a good idea, and they do give a quick sense of what has been tried.\n\nBut the stress-test note is right: the tables are not reliable as printed. [104] is described in the text as achieving 99.95% accuracy for QoT estimation, but Table V says 95–99%. [174] appears as MAE 0.18 dB in Table VI and 0.05 dB in the text. [131] is a fault-localization study that Section III treats as part of OCN, yet it sits in the OFC table. Fig. 1 has empty reference slots for several branches (PCA, policy/value RL in OFC; hierarchical and ICA in OCN). These are not merely typos: the paper's stated contribution is \"quantitative/qualitative comparisons,\" and if the numbers and placement are off in the few spots I spot-checked, I cannot trust the rest. The blame is not that the authors invented results; the entries presumably come from the cited papers. But the survey does not provide its search protocol or inclusion criteria, so selection bias is also unquantifiable.\n\nThere are also smaller editorial problems — \"Deap Learning,\" \"hese,\" repetitive future-directions prose — that indicate a rushed preprint. None of this changes the taxonomy, which is sound. The core idea — organize by ML/DL family, cross three application domains, list reported gains — is a legitimate and needed service. No new methods or data are claimed, so there is no circularity issue.\n\nI would send this to peer review, but with a clear demand: audit every table entry against its source, fix the internal contradictions, fill or remove the blank slots in Fig. 1, and state the search and selection methodology. After that, the survey becomes a solid entry point. As printed, it is a good skeleton with unreliable flesh.\n\nI would not cite the tables in my own work without checking original sources, but I would point people to the paper for its coverage and organization. It deserves a serious referee, not a desk reject.","headline":"A useful, algorithm-centric survey of ML/DL in optical communications whose quantitative tables are inconsistent enough that they need an audit before the paper can be trusted.","tokens_in":45094,"tokens_out":2515,"would_cite":true,"duration_ms":27057,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The survey claims that 282 prior studies, organized by algorithm, reveal where machine learning and deep learning improve optical fiber, network, and wireless systems over conventional baselines.","keywords":["machine learning","deep learning","optical fiber communication","optical communication networking","optical wireless communication","optical performance monitoring","nonlinear equalization","reinforcement learning"],"falsifier":"Randomly sample a dozen quantitative entries in Tables III through VIII, read the original papers behind the cited markers, and recompute or repeat each measurement under the stated conditions; finding that the reproduced values or their experimental conditions diverge materially from the table entries would undercut the survey's comparative conclusions.","tokens_in":44205,"feed_emoji":"📡","tokens_out":5194,"duration_ms":53263,"temperature":0.7,"pith_summary":"This paper is a survey: it tries to establish, in one place, how machine learning and deep learning are being used across optical fiber communication, optical communication networking, and optical wireless communication. It organizes roughly 282 prior studies into an algorithm-first taxonomy, and it summarizes their reported performance, complexity, data, and metrics in six comparison tables. A sympathetic reader would care because the survey turns a scattered literature into a map: for each type of task, it shows which algorithm families have been tried and what quantitative gains they claim over conventional methods. The paper further argues that deep learning is underexplored relative to machine learning, that optical wireless communication has been neglected by earlier surveys, and that algorithm-centric comparison reveals gaps and future directions.","feed_headline":"One survey maps 282 machine-learning results in optical links","feed_subtitle":"Algorithm-by-algorithm tables show where learning beats conventional signal processing and networking.","key_machinery":"The organizing machinery is a two-level taxonomy: three application domains, namely optical fiber communication, optical communication networking, and optical wireless communication, each subdivided into machine learning (supervised, unsupervised, and reinforcement learning) and deep learning (DNN, RNN, CNN, and DRL). Six summary tables, Tables III through VIII, carry the argument; each row reports performance, complexity, train/test data type, objective, input features, metrics, and adopted algorithm for the reviewed studies. This structure is what lets the survey compare quantitative gains across heterogeneous setups.","core_discovery":"The paper's central claim is that the literature on machine learning and deep learning for optical communication can be organized by algorithm rather than only by application, and that doing so exposes consistent patterns. Supervised and unsupervised learning methods, including support vector machines, artificial neural networks, k-nearest neighbors, clustering, principal and independent component analysis, regression, and ensemble learning, are widely used for nonlinear equalization, detection, quality-of-transmission estimation, and optical performance monitoring. Deep architectures, including deep neural networks, recurrent networks, convolutional networks, and deep reinforcement learning, push reported performance further in many cases while often claiming lower complexity than classical nonlinear equalizers. The survey asserts, with quantitative entries in Tables III through VIII, that these approaches improve Q-factor, bit error rate, OSNR estimation accuracy, spectrum utilization, blocking probability, throughput, and positioning error relative to conventional baselines. The discovery, if the survey is right, is not any single algorithm but the landscape itself: which algorithm families are attached to which optical tasks, and where the reported gains are largest.","pith_inferences":["If the reported quantitative gains are taken at face value, the strongest cross-cutting pattern is that hybrid methods, such as DNN/RNN cascades for equalization and CNN/RF for monitoring, appear most often; this suggests a design heuristic the paper does not state explicitly.","The comparability problem in the tables is the main editorial risk; a standardized benchmark suite with common fiber lengths, modulation formats, and link setups would sharpen the comparisons the survey can only approximate.","The algorithm-centric map highlights gaps, for instance the sparse use of some unsupervised methods in optical networking and optical wireless communication relative to optical fiber communication; these gaps are candidate targets for the next wave of applications.","A testable extension would be to use the survey's table entries as training data for a meta-analysis, regressing reported gains against algorithm family, data rate, distance, and simulation-versus-experiment status to see which factors predict success."],"forward_implications":["For nonlinear equalization in fiber links, the tables show several viable families, including SVM, ANN, clustering, RNN, CNN, and DRL-optimized Volterra equalizers, so a system designer can trade reported performance against complexity.","For monitoring and quality-of-transmission estimation, image-based CNN and feature-based RF/DNN methods report high accuracy for joint OSNR and modulation-format identification, supporting consolidation of several monitoring functions into one model.","For network-level control, RL and DRL agents are reported to reduce blocking probability and improve throughput or spectrum utilization in routing, spectrum assignment, handover, and power allocation problems.","The survey's future-directions argument implies that continual learning, transfer learning, active learning, explainable AI, and open datasets are the next needed steps if these methods are to be deployed in dynamic optical networks."],"supporting_citations":[{"why":"Defines the prior scope of ML for failure management in optical networks, a baseline this survey extends.","marker":"[22]"},{"why":"Shows the existing categorization of ML for routing optimization in SDNs, one of the scopes this survey broadens.","marker":"[23]"},{"why":"Provides an earlier review of AI for optical systems and networks whose coverage this survey updates and re-organizes by algorithm.","marker":"[27]"},{"why":"Overview of ML for optical communications and networking that motivates the algorithm-centric reorganization here.","marker":"[28]"},{"why":"Supplies mathematical foundations of basic ML techniques used to frame the reviewed optical-communication applications.","marker":"[29]"},{"why":"Highlights DL contributions to optical communications, the area this survey treats systematically alongside ML.","marker":"[31]"},{"why":"Reviews ML approaches for coherent optical-OFDM, providing one of the benchmark-DSP comparisons the survey incorporates.","marker":"[32]"},{"why":"Reviews ML for optical performance monitoring, one of the application areas whose table entries the survey compiles.","marker":"[34]"}],"fun_headline_variants":["Survey maps 282 ML results in optical communications","How ML and deep learning reshape optical links","Algorithm-by-algorithm: ML gains in optical systems","282 ML experiments on optical links, one survey","Where machine learning beats optical DSP"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The position rests on the assumption that the performance numbers drawn from the cited studies are accurate and can be fairly compared across different simulators, fiber lengths, modulation formats, and data rates.","fun_headline_variants_meta":{"raw":{"variants":["Survey maps 282 ML results in optical communications","How ML and deep learning reshape optical links","Algorithm-by-algorithm: ML gains in optical systems","282 ML experiments on optical links, one survey","Where machine learning beats optical DSP"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1535,"prompt_tokens":892,"completion_tokens":643,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":575}},"tokens_in":508,"tokens_out":643,"duration_ms":7143,"temperature":1.0,"reasoning_tokens":575,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:55:10.430791+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample a dozen quantitative entries in Tables III through VIII, read the original papers behind the cited markers, and recompute or repeat each measurement under the stated conditions; finding that the reproduced values or their experimental conditions diverge materially from the table entries would undercut the survey's comparative conclusions.","supporting_citations":[],"review_version":1}