{"id":"e85b842e-29d7-41cb-a62e-16a36ba65587","arxiv_id":"2508.20077","paper_version":2,"verdict":"REJECT","confidence":"LOW","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"ML-MaxProp embeds an XGBoost classifier into MaxProp to predict relay suitability in disaster networks, but the claimed performance gains are not backed by any presented data.","lead":"This paper proposes ML-MaxProp, a routing protocol that adds a machine-learning classifier to the classic MaxProp algorithm for disaster-prone delay-tolerant networks. The authors claim it improves delivery, latency, and overhead, but the simulation results supporting these claims are not actually presented in the manuscript.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central empirical claim is unverifiable: all result figures are absent and no numeric/statistical data are reported, so 'consistently surpasses' is an assertion without evidence.","rationale":"The reader correctly identified the central claim as the empirical superiority of ML-MaxProp and flagged the weak assumption about the training/evaluation loop. My stress-test pass lands on a more fundamental and immediate problem: the manuscript does not actually present any results. Every figure is missing, there are no tables or numeric summaries, and the statistical tests are referenced without any reported output. The central claim is therefore unsupported by the submitted text, regardless of whether the training loop is sound. This is a stronger concern than the reader's 'weakest_assumption' (which concerns the validity of the training data), but it is consistent with the reader's rationale that the paper is 'a claim without derivation.' I agree with the REJECT verdict; the appropriate action is unchanged from the reader's verdict. A conditional accept would require the missing data and artifacts to be supplied and the internal inconsistencies resolved. No ad hominem is intended; the issue is the absence of evidence, not the authors' intent.","tokens_in":9628,"tokens_out":2273,"duration_ms":22904,"concrete_test":"Request the missing Figures 1-8 and the underlying simulation CSV logs from the authors, or independently re-run the described ONE experiments with the stated parameters (node count 50-150, buffer 5-20 MB, TTL 300-3600 s, range 50-150 m, 10 repetitions per configuration). Recompute delivery probability, average latency, overhead ratio, and hop count for MaxProp and ML-MaxProp, and run the paired t-test/Wilcoxon signed-rank tests on per-run values. If ML-MaxProp does not significantly beat MaxProp at p<0.05, or if the recovered numbers contradict Section 6.2's claim that MaxProp achieves near-perfect delivery, the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that ML-MaxProp outperforms MaxProp and Spray-and-Wait on delivery probability, latency, overhead, and hop count, with statistical significance. The only support presented is prose and figure placeholders: none of Figures 1-8 are actually included, no tables or raw numbers appear anywhere, and the paired t-test/Wilcoxon tests mentioned in Section 5.4 are never reported with test statistics, p-values, or effect sizes. This is not a subtle methodological gap; the core comparative evidence is absent. The text is also internally inconsistent: Section 6.2 states that MaxProp 'consistently achieves near-perfect delivery across all node counts' and that Spray-and-Wait achieves 'comparable delivery performance' with minimal overhead, which conflicts with the abstract and Section 7 framing ML-MaxProp's 'consistently exceeded 99.8%' delivery as a significant improvement over baselines. Without the actual data, the claim that an XGBoost classifier trained on labels extracted from baseline MaxProp simulations (Section 5.2) improves on MaxProp's own heuristic cannot be distinguished from the model simply mimicking MaxProp, overfitting to the ONE simulator's SPMBM mobility, or the differences being within noise. The central claim therefore rises or falls on data that the manuscript does not contain.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ML-MaxProp, a DTN routing protocol that augments MaxProp with an XGBoost classifier to predict relay suitability using features such as encounter frequency, buffer occupancy, hop count, message age, and TTL. The authors report that ML-MaxProp outperforms MaxProp, Spray-and-Wait, and Epidemic in the ONE simulator under the Helsinki SPMBM mobility model, with delivery probability exceeding 99.8%, lower latency, and reduced overhead, validated by paired t-tests and Wilcoxon tests. However, the manuscript contains no figures, no numerical results, no statistical test statistics, and no implementation details that would allow the claims to be checked. The evaluation is also in-sample: the classifier is trained on labels extracted from baseline MaxProp simulations and tested on the same simulator. The central comparative claim is therefore unsupported as written.","tokens_in":9993,"tokens_out":2338,"duration_ms":25786,"significance":"If substantiated, an interpretable, low-cost ML extension of MaxProp could be a useful contribution to opportunistic networking in disaster scenarios. The feature set and the idea of embedding a supervised classifier into a utility-based routing heuristic are reasonable, and the use of ONE with a city-scale mobility model is a standard evaluation context. The paper also attempts to address explainability with SHAP/LIME and to include statistical validation. However, none of the promised evidence is present: all figures are missing, no numerical outputs are reported, and the statistical tests are only mentioned. The paper's core empirical contribution is entirely unverifiable, and the training/evaluation loop raises a circularity concern that is not addressed. As a result, the significance of the claimed results cannot currently be assessed.","major_comments":[{"comment":"The central claim—that ML-MaxProp 'consistently surpasses baseline protocols' with delivery probability >99.8%—is unsupported. All eight referenced figures (Figures 1–8) are absent, and no numerical values for delivery ratio, latency, overhead, or hop count appear anywhere. The paired t-test/Wilcoxon results mentioned in §5.4 are never reported (no test statistics, p-values, or effect sizes). Without the empirical data, the paper cannot be evaluated.","section":"§6, abstract, §7.1"},{"comment":"The evaluation is in-sample: the XGBoost model is trained on labeled outcomes drawn from baseline MaxProp simulations in the ONE simulator, then tested on the same simulator (§5.2: 'Data collected from baseline MaxProp simulations provided labeled outcomes'). This creates a risk that the model simply mimics MaxProp's decisions or overfits to the simulator's dynamics. No independent test data, real-world traces, or ablation analysis are provided to establish that the claimed improvements are genuine rather than artifacts of the training/evaluation loop.","section":"§5.2, §5.4"},{"comment":"The results are internally inconsistent. §6.2 states that MaxProp 'consistently achieves near-perfect delivery across all node counts' and that Spray-and-Wait achieves 'comparable delivery performance' with minimal relay cost. The abstract and §7.1, however, claim that ML-MaxProp 'significantly outperforms' these baselines, while also conceding it is 'matching or surpassing MaxProp.' If MaxProp already achieves near-perfect delivery, the claimed improvement is marginal or contradictory, and the 'significant outperformance' framing is misleading.","section":"§6.2 vs. abstract/§7.1"},{"comment":"No experimental parameters are reported beyond ranges (node count 50–150, buffer 5–20 MB, TTL 300–3600 s, range 50–150 m). The manuscript does not state which specific values were used, how many configurations were tested, what the 'default configuration' was, or how the ten repetitions were aggregated (means? confidence intervals?). This lack of detail makes the evaluation irreproducible. Without concrete parameter settings and per-configuration results, the claim of robustness across diverse conditions cannot be verified.","section":"§5.3, §5.4"}],"minor_comments":[{"comment":"All figure placeholders are empty. The manuscript says 'Figure 1 illustrates...', 'As shown in Fig. 2...', etc., but no images or captions are included. This is unacceptable for a submission.","section":"Figures 1–8"},{"comment":"Many references appear irrelevant to the topic (e.g., [4] on VR body pose estimation, [11] on power transformers, [13] on finger character recognition, [17] on SDN flow tables, [18] on sparse tensor accelerators, [21] on quasi-passive walkers, [29] on analog-to-digital converters). This suggests citation padding and undermines confidence in the scholarship.","section":"References"},{"comment":"The model description is underspecified: no dataset size, class balance, feature values, hyperparameters (max_depth, learning_rate), or train/test split statistics are given. The 80/20 split is mentioned, but there is no report of model accuracy, precision, recall, or validation performance.","section":"§5.2"},{"comment":"The text refers to 'this thesis' repeatedly, but the manuscript is formatted as a research article. This suggests the content was adapted from a thesis without proper editing.","section":"§7.1"},{"comment":"Naming is inconsistent: 'ML-MaxProp' is sometimes written 'mlmaxprop' or 'MLMaxProp' (Figure 8 description). Please standardize.","section":"§6.2, Figure 8 description"},{"comment":"The limitations section acknowledges that the model was 'trained solely on simulation-generated data,' but this is a central methodological weakness that should be acknowledged in §5 and addressed explicitly in the evaluation, not relegated to future work.","section":"§7.3"}],"recommendation":"reject","confidential_remarks":"This manuscript is an incomplete draft. The empirical core—all figures, numerical results, and statistical outputs—is missing, so the central claim cannot be verified. The in-sample training/evaluation setup compounds this: the model is trained on the very simulator it is tested against, and no evidence rules out mimicry or overfitting. While the topic is potentially interesting and the idea of combining XGBoost with MaxProp is not without merit, the paper as submitted does not meet the standards for publication. The absence of data is not a local fix; it requires a complete rework of the evaluation and likely additional experiments. I recommend rejection rather than major revision, as the current submission does not contain enough substance to referee."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe core idea—bolt an XGBoost classifier onto MaxProp's forwarding decisions in the ONE simulator—is sensible and squarely in the line of work it cites. The feature set (encounter frequency, buffer occupancy, hop count, age, TTL) is reasonable, and the authors are upfront about the limitation that training data comes from simulations. The limitation section is honest, and the planned use of paired t-tests and Wilcoxon is appropriate.\n\nThat's where the good news ends. As submitted, the paper contains no results. All eight figures are placeholders, there is not a single numeric delivery ratio, latency, overhead, or p-value anywhere in the text. The abstract and conclusion claim delivery probability >99.8% and overhead reduction with p<0.01, but the body only gives prose summaries. This is not a subtle gap: the central claim is unverifiable.\n\nThere's also an internal inconsistency the authors may not have noticed. Section 6.2 says MaxProp \"consistently achieves near-perfect delivery across all node counts\" and Spray-and-Wait \"achieves comparable delivery performance with minimal relay cost,\" which undercuts the abstract's framing that ML-MaxProp \"significantly outperforms\" baselines. If the baselines are already near-perfect, the claimed gain needs to be quantified and shown to be outside noise.\n\nThe training/evaluation loop also bothers me. The model is trained on labels from baseline MaxProp simulations and then evaluated against MaxProp in the same simulator. That is not fatal, but it makes it hard to distinguish genuine improvement from the model learning the simulator's dynamics. Without code or data, I can't tell.\n\nThe citation list is distractingly sloppy—references to power transformers, ADCs, and VR pose estimation have no bearing on DTN routing. That, plus the missing figures, suggests the manuscript was assembled in haste.\n\nMy take: this is a plausible idea that deserves to be tested properly, but this manuscript does not contain the test. It should be returned to the authors to add the actual results, raw data, and code, and to fix the internal consistency and citation issues. As is, it is not ready for peer review.\n\nRecommendation: desk reject with an invitation to resubmit once the evidence is actually present.","headline":"Sensible XGBoost+MaxProp idea, but the manuscript contains no actual results—all figures are placeholders and no numbers are reported—so the central performance claim is unverifiable as submitted.","tokens_in":10430,"tokens_out":2980,"would_cite":false,"duration_ms":30998,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Embedding an XGBoost classifier into MaxProp's forwarding pipeline yields near-perfect delivery, lower latency, and lower overhead in simulated post-disaster DTNs.","keywords":["Delay-Tolerant Networks","Disaster communication","MaxProp","Machine learning","XGBoost","Opportunistic routing"],"falsifier":"Independently rerun the published parameter sweep (50–150 nodes, 5–20 MB buffers, 300–3600 s TTL, 50–150 m range) for, say, 100 seeds and compare the paired delivery-probability difference between ML-MaxProp and MaxProp. If the 95% confidence interval of that difference contains zero, the headline claim of consistent superiority is falsified.","tokens_in":9560,"feed_emoji":"📡","tokens_out":8259,"duration_ms":83565,"temperature":0.7,"pith_summary":"The paper tries to establish that a delay-tolerant routing protocol can be made more reliable in post-disaster settings by replacing MaxProp's static forwarding heuristic with a lightweight XGBoost classifier that predicts, at each contact, whether a relay is worth using. The proposed protocol, ML-MaxProp, is trained offline on labels harvested from baseline MaxProp simulations and evaluated in the ONE simulator with Helsinki mobility. Across node counts, buffer sizes, TTLs, and communication ranges, the authors report delivery probability consistently above 99.8%, while Epidemic routing delivers about 61%, with lower latency and overhead than MaxProp. If true, this would make infrastructure-free mission-critical emergency messaging feasible by adding a cheap, interpretable learning step to an existing protocol.","feed_headline":"MaxProp plus XGBoost: 99.8% delivery in disaster networks","feed_subtitle":"A learned relay-suitability check keeps emergency messages flowing when infrastructure collapses.","key_machinery":"The mechanism is an XGBoost binary classifier embedded in MaxProp's forwarding pipeline. At each contact it consumes five contextual features—encounter frequency, hop count, buffer occupancy, message age, and TTL—and outputs a relay-suitability score; MaxProp's queue management still runs, but the handoff decision is gated by the learned model. The model is trained offline on labeled outcomes extracted from MaxProp's own simulation logs, so it is designed to learn which MaxProp-style forwarding opportunities actually lead to successful delivery.","core_discovery":"The central discovery is that MaxProp's queueing logic can be augmented—without replacing its store-carry-forward machinery—and still produce large gains. ML-MaxProp uses XGBoost to score each potential relay using contextual features: contact frequency, buffer occupancy, hop count, message age, and TTL. It forwards only when the classifier judges the relay suitable. Training labels come from baseline MaxProp simulations: a message is labeled successful if it eventually reaches its destination under MaxProp. At runtime the learned model gates the forwarding decision, and the authors report that this raises delivery probability above 99.8%, from Epidemic's ~61%, while matching or surpassing M","pith_inferences":["The closed training/evaluation loop (labels from baseline MaxProp in the same simulator used for testing) leaves open whether the gains persist across different mobility models; a cross-city retraining experiment would test that.","The same learned-gate idea could be bolted onto other utility-based DTN protocols such as Prophet or Spray-and-Wait, making the contribution a generic prefilter rather than a MaxProp-specific patch.","Because the features are per-contact and cheap to compute, the classifier is a plausible fit for smartphone- or Raspberry-Pi-class emergency nodes; public mobility traces would be the natural next validation.","A noise-feature control—shuffling the classifier's inputs while keeping the pipeline intact—would isolate whether the delivered gains come from the contextual features or from the XGBoost infrastructure itself."],"forward_implications":["Post-disaster DTNs can sustain above-99.8% delivery under constrained buffers, short TTLs, and varied node densities when a learned relay-suitability model is available.","Overhead falls sharply under resource constraints, saving bandwidth and energy in exactly the conditions where emergency networks are most stressed.","The reported gains are statistically significant across paired t-tests and Wilcoxon tests, so they are not presented as single-run artifacts.","TTL and buffer state, rather than raw contact frequency, dominate forwarding decisions, pointing protocol designers toward message-lifetime and storage-aware routing.","ML-MaxProp stays interpretable: SVM classification and SHAP/LIME analysis show its decisions align with domain knowledge and remain distinguishable from MaxProp's."],"supporting_citations":[{"why":"Supplies the MaxProp protocol definition and the baseline that ML-MaxProp must beat.","marker":"[3]"},{"why":"Cited for the paired t-test and Wilcoxon statistical validation methodology used to certify the performance gap.","marker":"[6]"},{"why":"Supplies the premise that lightweight supervised classifiers can predict forwarding success in DTNs.","marker":"[9]"},{"why":"Supplies the disaster-area network context and the ONE/SPMBM simulation design used for evaluation.","marker":"[10]"},{"why":"Provides the survey basis that ML routing approaches often lack realistic evaluation and strong baselines.","marker":"[14]"},{"why":"Supports the map-based social mobility evaluation methodology that grounds the simulation benchmark.","marker":"[28]"}],"fun_headline_variants":["ML-MaxProp: 99.8% delivery when infrastructure fails","ML-guided DTN routing hits 99.8% delivery","XGBoost makes MaxProp smarter for disaster comms","Adaptive ML routing outperforms classic DTNs in crises","AI-powered relay selection keeps disaster networks alive"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The classifier is trained on labels produced by baseline MaxProp inside the same simulator used for testing, so the whole claim assumes those training labels represent what actually makes a relay good at runtime rather than merely reproducing MaxProp's own behavior.","fun_headline_variants_meta":{"raw":{"variants":["ML-MaxProp: 99.8% delivery when infrastructure fails","ML-guided DTN routing hits 99.8% delivery","XGBoost makes MaxProp smarter for disaster comms","Adaptive ML routing outperforms classic DTNs in crises","AI-powered relay selection keeps disaster networks alive"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000709,"raw_usage":{"total_tokens":3043,"prompt_tokens":769,"completion_tokens":2274,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":2191}},"tokens_in":513,"tokens_out":2274,"duration_ms":16348,"temperature":1.0,"reasoning_tokens":2191,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:13:12.417763+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently rerun the published parameter sweep (50–150 nodes, 5–20 MB buffers, 300–3600 s TTL, 50–150 m range) for, say, 100 seeds and compare the paired delivery-probability difference between ML-MaxProp and MaxProp. If the 95% confidence interval of that difference contains zero, the headline claim of consistent superiority is falsified.","supporting_citations":[{"cited_title":"Machine learning based intelligent routing for VDTNs,","cited_arxiv_id":null,"evidence_quote":"Supplies the premise that lightweight supervised classifiers can predict forwarding success in DTNs."},{"cited_title":"A review of applicable technologies, routing protocols, requirements, and architecture for disaster area networks,","cited_arxiv_id":null,"evidence_quote":"Supplies the disaster-area network context and the ONE/SPMBM simulation design used for evaluation."},{"cited_title":"Applications of machine learning in networking: A survey of current issues and future challenges,","cited_arxiv_id":null,"evidence_quote":"Provides the survey basis that ML routing approaches often lack realistic evaluation and strong baselines."},{"cited_title":"Performance Evaluation of DTN Routing Protocols on Map -Based Social Mobility Models for DTN Networks,","cited_arxiv_id":null,"evidence_quote":"Supports the map-based social mobility evaluation methodology that grounds the simulation benchmark."}],"review_version":1}