{"id":"0e4ddb50-31c4-4487-ab2a-392ab87bb97b","arxiv_id":"2506.03041","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A proposed CNN-based system for automatic OTDR fault classification and localization in rural fiber is claimed to improve detection accuracy from 71.2% to 93.4%, but the evidence is not reproducible.","lead":"This preprint describes an AI system that reads fiber-optic test signals (OTDR traces) to automatically find and classify faults in rural broadband cables. The paper reports that its CNN model beats traditional thresholding on a 10-kilometer test spool, but provides no code, data, or model details.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central performance claim is unsupported because Table 1's metrics lack a defined train/test split and independent test set; the evaluation is circular, so 93.4% accuracy does not establish real-world fault localization.","rationale":"I read the paper's strongest claim as the Table 1 metrics. The condition for that claim to hold is that the evaluation measures generalization to unseen fiber conditions. The manuscript fails to provide the minimum information needed to assess this: no train/test split, no test-set size, no model architecture or hyperparameters, no error bars, no dataset or code, and no independent field deployment. The reader's weakest assumption about representativeness is plausible, but I locate a more immediate problem: internal validity. The paper's own statements are contradictory ('Field testing' vs. 'controlled fiber testbed and synthetic datasets'), and the reused placeholder references with future dates (References 19-30) are independent integrity red flags that reinforce rejection. My concern is not that the numbers are necessarily fabricated, but that they are unreproducible from the text, so no reviewer can verify them. This agrees with the reader in outcome but focuses on missing evaluation protocol rather than transferability; hence 'partial.' The verdict should remain REJECT, as the evidence does not support the central claim.","tokens_in":8218,"tokens_out":3004,"duration_ms":33339,"concrete_test":"Obtain the trained model and the 7,500-trace dataset from the authors, or reconstruct the protocol if released. Perform a spool-stratified or session-stratified split: train on traces from some fiber spool/measurement sessions, then test on traces from a held-out spool/session never seen in training. Compute per-class accuracy and localization error on that held-out set. If the held-out detection accuracy is materially below 93.4% (e.g., <80%) or localization error exceeds the OTDR's stated resolution, the Table 1 metrics overstate generalization. Also require at least three repeated trials to report confidence intervals; if no such split is possible with the released data, the central claim remains unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the 93.4% detection accuracy and 1.4-meter localization error in Table 1. For this to be evidence, the model must be evaluated on traces not used in training, and the experimental protocol must be specified. Neither condition is met. Section III states the CNN is trained on 7,500 labeled OTDR traces, but Section IV reports only that the system was tested using a 10-kilometer fiber optic spool with induced splice, bend, and connector faults; it never states how many traces were in the test set, how the train/test split was made, or whether any training trace came from the same spool and fault-induction session. If the test traces came from the same controlled spool or synthetically augmented variants of it, the reported accuracy can reflect memorization of fault signatures rather than generalization. The absence of model architecture, hyperparameters, error bars, and code makes it impossible to rule out this circularity. The abstract's own wording—'controlled fiber testbed and synthetic datasets'—conflicts with Table 1's 'Field testing' label, further undermining the claim that the metrics transfer to rural deployments. This is not merely an external-validity concern; the internal validity of the evaluation is unverifiable from the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an AI-augmented OTDR fault localization system for rural fiber networks, combining a Raspberry Pi-based acquisition module with a cloud-deployed CNN classifier. The system is trained on 7,500 labeled OTDR traces covering splice, bend, and connector faults, and evaluated on a 10-km fiber spool with artificially induced faults. Table 1 reports that the proposed AI model achieves 93.4% detection accuracy, a 7.1% false positive rate, 1.4-meter average localization error, and 4.2-second average detection time, compared to 71.2%, 18.9%, 5.6 meters, and 11.5 seconds for traditional thresholding. The manuscript claims a scalable, field-deployable tool for rural broadband maintenance aligned with the BEAD program.","tokens_in":8453,"tokens_out":3073,"duration_ms":36567,"significance":"If the reported performance were substantiated, the system could be a useful contribution to rural fiber network diagnostics, where automated OTDR interpretation is an acknowledged operational need. The hardware-software concept is plausible and timely given U.S. broadband expansion initiatives. However, the central performance claim is not supported by the evidence presented: the evaluation is circular (trained and tested within the same synthetic-and-controlled pipeline), no model details or statistical confidence measures are provided, and the reference list contains fabricated entries. As it stands, the paper does not establish a valid empirical contribution, so its significance is currently unverified.","major_comments":[{"comment":"The reported accuracy, false positive rate, and localization error are not accompanied by any description of the test set size, how the train/test split was performed, or whether the test traces were collected independently of the 7,500 synthetic traces used for training. Because the faults induced on the 10-km spool belong to the same three classes as the synthetic training data, the 93.4% detection accuracy may simply reflect the model's fit to the same controlled distribution and cannot be interpreted as evidence of generalization. A proper holdout protocol with independent data collection (different spool configurations, OTDR devices, and environmental conditions) is required to support the central claim.","section":"Section IV, Table 1"},{"comment":"The CNN architecture, loss function, optimizer, hyperparameters, and data augmentation details are entirely unspecified. The paper reports only point estimates in Table 1 with no repeated runs, confidence intervals, or error bars. Since the entire contribution rests on the claimed superiority of the AI model over thresholding, the absence of these implementation and statistical details makes the reported metrics unverifiable and the comparison non-reproducible.","section":"Section III and Section IV"},{"comment":"Table 1 is labeled 'Field testing using 10 km test fiber rolls,' but the abstract describes the evaluation as 'a controlled fiber testbed and synthetic datasets.' The text in Section IV also claims the model performed well 'under variable weather and topological noise conditions,' yet no weather, temperature, humidity, or noise measurements are reported anywhere in the manuscript. This internal contradiction about the evidence base and the unsupported environmental claim together undermine the external validity of the reported metrics for real rural deployments.","section":"Section IV and Abstract"},{"comment":"References 4 through 30 are placeholder citations: they share identical titles but list sequential authors (Author1, Coauthor1, Researcher1, etc.) with publication years extending into the future up to 2037. This makes the related-work positioning impossible to verify and constitutes a serious scholarly integrity problem that is independent of the technical evaluation. The paper must be based on genuine, verifiable prior literature.","section":"Section VII, References"}],"minor_comments":[{"comment":"Figures 1 and 2 are referenced only in the contributions list and are never cited in the experimental section; their captions claim they support training and validation, but no analysis in Section IV refers to them.","section":"Section I and Section IV"},{"comment":"The manuscript does not state the number of OTDR traces collected on the spool, the number of induced faults per fault type, or the OTDR device settings (e.g., pulse width, sampling interval, averaging time), so the baseline comparison cannot be reproduced.","section":"Section IV"},{"comment":"The text has inconsistent spacing, missing punctuation, and the bibliography mixes a few real references with a long placeholder sequence, making the reference list unreliable as a guide to prior work.","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"I recommend rejection. The central evaluation is circular and underspecified, and the reference list contains fabricated entries. Even if the experimental results were genuine and reproducible, the paper would need a complete description of the model and protocol plus a properly independent field or field-like validation to be publishable; the placeholder citations alone are disqualifying for a serious journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper's Table 1 numbers are the whole ballgame, and they're unsupported. The reference list alone is disqualifying.\n\nThe problem is real enough. Rural ISPs do rely on manual OTDR trace reading, and a low-cost edge-to-cloud pipeline that classifies splice, bend, and connector faults would save technician time. The hardware concept (Raspberry Pi, GPS, cloud CNN) is coherent, and the author does cite Al-Khafaji 2021 as prior CNN-OTDR work, so the classification idea is not being oversold as wholly new.\n\nThe soft spots are not minor. There is no CNN architecture, no training procedure, no hyperparameters, no repeated runs, no error bars, and no code or data. Section III says 7,500 labeled OTDR traces were used for training; Section IV describes a 10 km spool with induced faults but never states how test traces were produced or whether they came from the same spool/session as the training data. Given the training data are synthetic and the faults are artificially induced from the same three classes, the 93.4% accuracy in Table 1 could easily be memorization of fault signatures rather than generalization. That is exactly the circularity the stress-test note describes, and the paper does nothing to rule it out. The abstract even says 'controlled fiber testbed and synthetic datasets,' which directly contradicts Table 1's 'Field testing' label.\n\nThe reference list is worse. Entries 4-30 are 'A. Author1, B. Coauthor1, and C. Researcher1' with future publication dates running to 2037. These are placeholders or fabrications. This is not an editorial slip; it means the related-work section is not a genuine engagement with the literature, and it destroys the trust a referee needs.\n\nWhat the paper does well is describe a deployment scenario that matters. But the technical substance is a thin wrapper around a figure and a table. I'd not send this to peer review. Desk-reject, and invite the author to resubmit with an honest evaluation: real OTDR data, a defined train/test split, baseline comparisons, and a clean reference list.","headline":"A system-integration idea with an unverifiable central metric, a circular evaluation, and a reference list full of fabricated placeholders; desk-reject.","tokens_in":9008,"tokens_out":2739,"would_cite":false,"duration_ms":29510,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A CNN trained on OTDR traces can localize and classify splice, bend, and connector faults in rural fiber links with 93.4% accuracy and 1.4-meter average error, outperforming traditional thresholding.","keywords":["optical time-domain reflectometry","fault localization","fiber optics","convolutional neural network","rural broadband","middle-mile infrastructure","predictive maintenance","BEAD program"],"falsifier":"Run the trained model on OTDR traces from a different manufacturer's unit or from an in-service rural fiber route with known fault locations, and compare the classification and localization metrics. If accuracy drops below the claimed 93.4% or localization error exceeds a few meters on traces the model did not train on, the claimed field readiness fails.","tokens_in":8006,"feed_emoji":"📡","tokens_out":4166,"duration_ms":39026,"temperature":0.7,"pith_summary":"The paper claims that a convolutional neural network can automate optical time-domain reflectometer (OTDR) fault analysis for rural fiber networks, replacing manual trace reading by technicians. On a 10-kilometer test spool with induced splice, bend, and connector faults, the AI model reports 93.4% detection accuracy, a 7.1% false positive rate, and an average localization error of 1.4 meters, against 71.2%, 18.9%, and 5.6 meters for a conventional thresholding baseline. The system pairs a Raspberry Pi-based acquisition module with a cloud analytics engine, so the claim is that this low-cost stack gives rural ISPs near-real-time fault localization. If true, it would let small operators move from reactive outage response to proactive maintenance.","feed_headline":"AI pinpoints fiber faults to 1.4 meters","feed_subtitle":"CNN trained on 7,500 OTDR traces beats manual thresholding for rural networks.","key_machinery":"The load-bearing mechanism is a convolutional neural network trained on 7,500 labeled OTDR traces representing splice, bend, and connector faults. The CNN takes backscatter intensity over distance as input and outputs fault type and distance estimate. A Raspberry Pi 4 with an OTDR probe and GPS module handles edge acquisition; the cloud runs the model and serves a dashboard. The CNN's learned features are what distinguish subtle signal degradations that fixed dB-loss cutoffs miss.","core_discovery":"The central claim is that AI-augmented OTDR interpretation is both feasible and substantially better than threshold-based analysis for rural fiber networks. Using 7,500 labeled OTDR traces from a controlled 10-km spool, the proposed CNN classifies three fault types (splice loss, bend faults, connector damage) and estimates fault distance. The reported results are 93.4% detection accuracy, 7.1% false positive rate, 1.4-meter average localization error, and 4.2-second average detection time; the thresholding baseline achieves 71.2%, 18.9%, 5.6 meters, and 11.5 seconds. The abstract and conclusion frame this as enabling proactive maintenance in low-resource environments targeted by the U.S. BEAD program.","pith_inferences":["Because the reported gains come from a single controlled spool and one set of 7,500 synthetic traces, the strongest testable extension is a cross-device field trial: retrain the CNN on traces from several OTDR models and fiber routes, then measure whether the 1.4-meter error holds.","If the model transfers, rural network operators could pair it with drone- or vehicle-based OTDR patrols, turning the dashboard into a preventive maintenance trigger rather than an outage response tool.","The same CNN architecture could be extended to classify gradual degradation like connector contamination or water ingress, which thresholding misses entirely; the paper mentions expanding dataset diversity in future work.","The 4.2-second detection time suggests the classifier could eventually run on the edge device itself, removing the cloud dependency in disconnected rural areas; the paper does not test this."],"forward_implications":["Fiber fault detection no longer requires a specialist to read OTDR traces; a technician with a tablet can get a classified fault and map location.","Fault detection time drops from 11.5 to 4.2 seconds, making real-time or near-real-time monitoring practical for rural links.","Localization error of 1.4 meters on a 10-km link is small enough to send a crew to the right splice case or pole.","The modular, low-cost hardware profile fits the budgets of small ISPs, cooperatives, and state broadband programs.","The training pipeline with synthetic and real data points to a path for expanding to other fault types and network geometries."],"supporting_citations":[{"why":"Supplies a supervised DWDM monitoring approach that the paper positions as unsuitable for rural, decentralized settings.","marker":"Zhang et al. (2022)"},{"why":"Supplies an AI-driven smart city monitoring system that requires high-end sensors, contrasting with the low-cost rural goal.","marker":"Han et al. (2023)"},{"why":"Supplies the CNN-for-OTDR classification approach that this paper extends with edge/cloud integration and real-time diagnostics.","marker":"Al-Khafaji (2021)"}],"fun_headline_variants":["AI fiber fault localization drops to 1.4 meters, cuts false positives","CNN on OTDR traces beats manual thresholding for rural fiber faults","Neural network pinpoints fiber faults to 1.4m, outperforms baselines","AI-augmented OTDR enhances fault detection for US rural broadband"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 7,500 labeled traces and the 10-kilometer spool are assumed to represent real U.S. rural fiber deployments, so the reported 93.4% accuracy and 1.4-meter error would carry over to field conditions; the paper gives no independent field data to confirm this transfer.","fun_headline_variants_meta":{"raw":{"variants":["AI fiber fault localization drops to 1.4 meters, cuts false positives","CNN on OTDR traces beats manual thresholding for rural fiber faults","Neural network pinpoints fiber faults to 1.4m, outperforms baselines","AI-augmented OTDR enhances fault detection for US rural broadband"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000279,"raw_usage":{"total_tokens":1601,"prompt_tokens":834,"completion_tokens":767,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":685}},"tokens_in":450,"tokens_out":767,"duration_ms":8913,"temperature":1.0,"reasoning_tokens":685,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:09:33.903129+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained model on OTDR traces from a different manufacturer's unit or from an in-service rural fiber route with known fault locations, and compare the classification and localization metrics. If accuracy drops below the claimed 93.4% or localization error exceeds a few meters on traces the model did not train on, the claimed field readiness fails.","supporting_citations":[],"review_version":1}