{"id":"2f692829-c4e1-4e54-8753-5ff864f9edfb","arxiv_id":"2412.00472","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":13,"one_line_summary":"A hybrid CNN-DWT-self-attention pipeline with swarm optimizers reports 98.11% accuracy on ISIC-2016 and 97.95% on ISIC-2017.","lead":"This paper combines discrete wavelet transforms, self-attention, and swarm-based optimizers with pretrained CNNs for skin cancer classification. On two public dermoscopy datasets, the best configurations report about 98% accuracy, roughly 1% above baseline methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ISIC-2017 significance tables (8-11) are numerically identical to the ISIC-2016 tables (3-6), so the claimed statistical support for the headline accuracy gains is not established.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the validity of the evaluation, specifically the ISIC-2017 significance claims. The duplicated p-value tables are concrete, checkable evidence that the statistical support was not independently produced for ISIC-2017. This is sufficient to reject the manuscript as written, even though the proposed pipeline might work better than the text demonstrates. Additional internal inconsistencies reinforce this: the text states DenseNet+Wavelet+MGTO reached 98.87% on ISIC-2016 while Table 1 reports 97.87%, and the conclusion's claimed improvements of 1.1% and 2.05% are not clearly derivable from the reported tables. None of these observations require assumptions about author intent; they are arithmetic and textual contradictions in the manuscript. A one-cell recomputation of the t-test from the published fold accuracies would settle the matter, and the expected outcome is a large discrepancy from the reported values. Therefore the reader's REJECT verdict should stand unchanged.","tokens_in":24438,"tokens_out":4621,"duration_ms":44115,"concrete_test":"Recompute the pairwise t-test for Inception+Wavelet+Fox vs Inception+Wavelet+MGTO on ISIC-2017 using the five fold accuracies in Table 7 (Fox: 0.9725, 0.9795, 0.9700, 0.9680, 0.9690; MGTO: 0.9771, 0.9715, 0.9700, 0.9701, 0.9791) and Eq. (34)-(35). Table 9 reports (-2.6962, 0.0272), identical to the ISIC-2016 cell. If the recomputed statistic and p-value differ materially from those values, the ISIC-2017 significance tables were not computed from the ISIC-2017 fold data, confirming that the statistical comparison is invalid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is an empirical accuracy improvement, and its statistical evidence is the pairwise t-test tables. Tables 3-6 (ISIC-2016) and Tables 8-11 (ISIC-2017) are numerically identical apart from one row header, despite the two datasets having different sizes and different 5-fold accuracies in Tables 2 and 7. If the t-values and p-values had actually been computed from the ISIC-2017 fold results using Eq. (34)-(35), they could not coincide to four decimal places with the ISIC-2016 values. This means the significance comparisons reported for ISIC-2017 were not derived from the ISIC-2017 experiments shown in the paper, leaving the claim that the optimizers improve accuracy by at least 1% without valid statistical support. In addition, Table 1 does not state which split was used for the reported accuracy numbers, so the headline accuracies themselves cannot be independently checked from the text. The numerical duplication is the load-bearing problem: it directly invalidates the significance claims, not merely the presentation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a skin-cancer classification pipeline that combines a pretrained CNN (DenseNet-121, Inception, Xception, or MobileNet), a late discrete wavelet transform (DWT), a self-attention module, and one of three swarm-based optimizers (FOX, IGWO, or MGTO) that tunes a final ANN head. The authors report accuracy up to 98.11% on ISIC-2016 and 97.95% on ISIC-2017 and claim improvements of at least 1% over prior methods. The evaluation is based on 5-fold cross-validation tables and pairwise t-test significance tables, with public code and data availability stated.","tokens_in":24721,"tokens_out":8503,"duration_ms":76997,"significance":"If the reported numbers were reproducible under a clean evaluation protocol, the paper would offer a useful empirical data point: a late DWT plus self-attention and a swarm-optimized head reaching roughly 98% accuracy on two standard skin-lesion benchmarks. The architecture description is readable, and the release of a public repository is a positive feature. However, the contribution is incremental and mostly empirical; the novelty lies in the specific combination of known components rather than in a fundamentally new method. The central claims are currently not supported because the statistical tables and protocol details contain serious problems that must be addressed.","major_comments":[{"comment":"The ISIC-2017 significance tables (Tables 8–11) are numerically identical to the ISIC-2016 tables (Tables 3–6) except for a row label in Table 8. Because the 5-fold accuracies in Tables 2 and 7 differ between the two datasets, the two-sample t-statistics and p-values computed from Eqs. (34)–(35) cannot coincide to four decimal places. This invalidates the statistical support for the claimed ISIC-2017 improvements, and the tables need to be recomputed from the actual ISIC-2017 fold results or removed.","section":"Tables 8–11 and Eqs. (34)–(35)"},{"comment":"The headline accuracy for DenseNet+Wavelet+MGTO on ISIC-2016 is reported as 98.87% in the Results section but as 97.87% in Table 1; this discrepancy must be resolved. In addition, the manuscript never states the train/test split for the Table 1 accuracies; the only split description (before Table 2) says 15% of the train data is used as validation and 65% as training, which is incomplete and does not specify the test set. Without a precise split, the reported accuracies cannot be reproduced.","section":"Table 1 and Results"},{"comment":"The claim that the k-fold results show \"there is no sign of overfitting and underfitting\" is unsupported because only validation-fold accuracies are presented; no training-set accuracy or train/test generalization gap is shown. The authors should either report training and test performance or remove the overfitting claim.","section":"Tables 2 and 7; text before Table 2"},{"comment":"The best result is selected after comparing many model/optimizer combinations in Table 1, but no model-selection protocol is described. If the test set was used to choose the best combination, the reported accuracies are optimistically biased; the authors should describe how the validation set was used for selection and confirm that the final numbers are on an untouched test set.","section":"Results and Table 1"},{"comment":"The conclusion states that the method improves accuracy \"by at least 1.1%\" on ISIC-2016 and by \"2.05%\" on ISIC-2017 relative to not using swarm-based optimizers, but Table 1 does not show a single consistent baseline supporting these numbers; for example, DenseNet+Wavelet+Fox improves over DenseNet by 0.13 percentage points, while MobileNet+Wavelet+Fox improves over MobileNet by 3.13 percentage points. The quantitative gain claim should be recomputed and stated with the specific baseline used.","section":"Conclusion"}],"minor_comments":[{"comment":"The heading \"Conlusion\" should be \"Conclusion\".","section":"Conclusion"},{"comment":"The caption for Figure 1 is inconsistent: it lists C1 and C2 for the Inception block but then says \"B2: Inception Net Block\"; the block labels should be corrected.","section":"Figure 1"},{"comment":"The legends in Figures 2–5 use the label \"AGTO\"; this should be \"MGTO\".","section":"Figures 2–5"},{"comment":"Table 8 uses the row header \"Xception+Wavelet+GWO\"; this should be \"Xception+Wavelet+IGWO\" for consistency with Tables 3 and 10.","section":"Table 8"},{"comment":"The methodology says the swarm optimizers tune \"filters size, kernel size\" even though the trainable head after DWT and self-attention is a dense ANN; clarify whether the optimizers tune pretrained CNN hyperparameters or the ANN weights.","section":"Methodology"},{"comment":"The literature review cites references [5] and [11] as skin-cancer studies, but the titles indicate they concern sickle cell disease; please verify the citations or replace them with appropriate skin-cancer works.","section":"Literature review"}],"recommendation":"major_revision","confidential_remarks":"The duplicated ISIC-2017 significance tables are the most serious issue; I would ask the authors to provide the raw per-fold results for ISIC-2017 and recompute Tables 8–11, and to specify the exact train/test split. If they cannot do so, the paper should not be published. I recommend major revision rather than rejection because the core architecture description and the stated code availability make it possible to verify and correct these points in a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the paper's central claim—that the swarm optimizers add a significant accuracy boost—is not supported by the evidence as written. The ISIC-2017 significance tables (8–11) are numerically identical to the ISIC-2016 tables (3–6) to four decimal places, yet the 5-fold accuracies for the two datasets in Tables 2 and 7 are clearly different. That cannot happen if the t-tests were computed with Eq. (34)–(35) from the reported folds. The text also misreports DenseNet+Wavelet+MGTO as 98.87% on ISIC-2016 when Table 1 says 97.87%, and the train/test split is never defined precisely (the '15% validation / 65% training' sentence is ambiguous and doesn't add to 100%).\n\nCredit where it's due: the pipeline is easy to follow, the authors compare four backbones with and without DWT and three optimizers on two public datasets, and they give per-fold results. Code and data are available on GitHub, which is more reproducible than most papers in this area. The basic idea—a late DWT and attention before a swarm-optimized classifier—is a legitimate incremental extension of prior wavelet-CNN and I-GWO work, so the accurate numbers could be usable if the evaluation were solid.\n\nThe soft spots are not cosmetic. The duplicated tables are load-bearing because the p-values are the only statistical evidence for the headline improvement. The split ambiguity means no one can independently reproduce the reported accuracies. The overfitting claims are assertion only—no train/test gap is reported. These are fixable in principle, but they require re-running the experiments and recomputing the statistics, not just copy edits.\n\nWho gets value from this? Someone doing a literature sweep on swarm-optimizer + wavelet pipelines might skim it, but I wouldn't cite the numbers. A serious editor should not spend referee time on this version. My recommendation: desk reject with an invitation to resubmit after a major overhaul of the evaluation section, or reject outright if the duplicated tables turn out to be more than copy-paste errors.","headline":"The accuracy gains are plausible but the statistical support is not: the ISIC-2017 p-value tables are identical to the ISIC-2016 ones, and the headline accuracy is misquoted.","tokens_in":25235,"tokens_out":3686,"would_cite":false,"duration_ms":35599,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A late discrete wavelet transform plus three swarm optimizers lifts four pretrained CNN classifiers past 98% accuracy on ISIC-2016 and ISIC-2017 skin-lesion benchmarks.","keywords":["skin cancer diagnosis","discrete wavelet transform","swarm optimization","self-attention","transfer learning","ISIC 2016","ISIC 2017","melanoma classification"],"falsifier":"Re-run the paired t-tests on the per-fold accuracies given for ISIC-2017 (Table 7) and compare with Tables 8-11: those tables exactly repeat the ISIC-2016 numbers, so unless the 2017 statistics are recomputed from 2017 folds, the claimed statistical significance is not supported. A simpler check is whether any independent run reproduces the reported 97.95% Inception+Wavelet+MGTO accuracy on the official ISIC-2017 test split.","tokens_in":24265,"feed_emoji":"🩺","tokens_out":6794,"duration_ms":61158,"temperature":0.7,"pith_summary":"This paper argues that a comparatively cheap post-processing pipeline can push already-strong deep classifiers over a clinically meaningful accuracy threshold in skin-cancer screening. The authors insert a discrete wavelet transform after the feature maps of four pretrained CNNs, add a self-attention layer, and then replace the ordinary training rule of a small classification head with swarm-based optimizers (Improved Grey Wolf Optimizer, Fox optimizer, Modified Gorilla Troops Optimizer). On ISIC-2016 they report 98.11% accuracy with MobileNet+Wavelet+Fox and DenseNet+Wavelet+Fox, and on ISIC-2017 97.95% with Inception+Wavelet+MGTO, each at least 1% above the comparison methods. If reproducible, the result matters because early and accurate automated screening could catch melanoma before it advances.","feed_headline":"Wavelets plus swarm optimizers push skin-cancer accuracy past 98%","feed_subtitle":"Late DWT and self-attention lift four pretrained CNNs past prior ISIC 2016/2017 results by at least 1 percent.","key_machinery":"The load-bearing mechanism is the late discrete wavelet transform (DWT) module: it takes the feature maps produced by a pretrained CNN and decomposes each into four sub-bands (LL, LH, HL, HH), preserving low-frequency global structure and high-frequency edges and textures, then concatenates them for a self-attention layer and a dense classification head. Around this, three swarm optimizers—Fox, Improved Grey Wolf Optimizer (IGWO), and Modified Gorilla Troops Optimizer (MGTO)—are used to choose ANN weights and hyperparameters, replacing standard backpropagation-based fine-tuning of the head.","core_discovery":"The central claim is that a 'late' discrete wavelet transform applied to CNN feature maps, followed by self-attention and swarm-optimizer-driven weight tuning, improves binary melanoma classification on standard dermoscopy benchmarks beyond prior art. Specifically, the paper reports that MobileNet + Wavelet + FOX and DenseNet + Wavelet + FOX reach 98.11% accuracy on ISIC-2016, Inception + Wavelet + MGTO reaches 97.95% on ISIC-2017, and these results exceed comparison methods by at least 1%. The paper also claims the wavelet stage and optimizer stage each contribute, with the relative benefit varying by backbone.","pith_inferences":["The wavelet stage likely acts as a frequency-domain regularizer that sharpens edge and texture cues and reduces sensitivity to class imbalance, though the paper's ablations do not isolate this mechanism directly.","A direct comparison against plain backpropagation fine-tuning of the same head would clarify how much of the gain is due to swarm optimization rather than to additional training on the wavelet and attention features.","The ISIC-2017 significance tables numerically duplicate the ISIC-2016 tables, so the statistical-significance claims for the 2017 results need to be re-derived before the improvement can be taken as established."],"forward_implications":["The pipeline can be added to any pretrained CNN by inserting DWT, self-attention, and a swarm-optimized head, so multiple backbones benefit without retraining the trunk.","The reported accuracy gains on both ISIC-2016 and ISIC-2017 suggest the approach transfers across datasets of different sizes and difficulty.","Since the best optimizer differs by backbone (Fox with MobileNet and DenseNet on ISIC-2016, MGTO with Inception on ISIC-2017), the paper implies optimizer choice should be treated as a per-architecture hyperparameter.","The combination offers a route to automated screening tools that could rank suspicious lesions for dermatologist review."],"supporting_citations":[{"why":"Supplies the ISIC-2016 benchmark used for the headline 98.11% results.","marker":"44"},{"why":"Supplies the ISIC-2017 benchmark used for the 97.95% Inception+Wavelet+MGTO result.","marker":"45"},{"why":"Lai et al. is the prior IGWO-based skin-cancer method whose reported accuracy the paper must beat.","marker":"47"},{"why":"Nawaz et al. provides a deep-learning comparison baseline on the same datasets.","marker":"48"},{"why":"InSiNet provides another comparison baseline in the accuracy tables.","marker":"49"},{"why":"Defines the Fox optimizer, one of the three swarm optimizers central to the method.","marker":"37"},{"why":"Defines the Improved Grey Wolf Optimizer (IGWO), the second optimizer used on the classification head.","marker":"40"},{"why":"Defines the Modified Gorilla Troops Optimizer (MGTO), the third optimizer and the best choice for Inception on ISIC-2017.","marker":"42"},{"why":"Defines the DenseNet-121 backbone used in one of the top ISIC-2016 combinations.","marker":"35"},{"why":"Defines the MobileNet backbone used in the other top ISIC-2016 combination.","marker":"36"}],"fun_headline_variants":["Wavelet plus swarm tuning hits 98% in skin cancer detection","Late DWT and swarm optimizers boost skin cancer accuracy","Skin cancer AI exceeds 98% with wavelets and swarms","98% accuracy in skin cancer via wavelet-swarm combo"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every headline comparison assumes that all models were trained and tested on the same fixed data split and that the ISIC-2017 statistical tables were computed on 2017 data; the duplicated p-value tables indicate that this premise is likely violated.","fun_headline_variants_meta":{"raw":{"variants":["Wavelet plus swarm tuning hits 98% in skin cancer detection","Late DWT and swarm optimizers boost skin cancer accuracy","Skin cancer AI exceeds 98% with wavelets and swarms","98% accuracy in skin cancer via wavelet-swarm combo"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000406,"raw_usage":{"total_tokens":2156,"prompt_tokens":1034,"completion_tokens":1122,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":1051}},"tokens_in":650,"tokens_out":1122,"duration_ms":8113,"temperature":1.0,"reasoning_tokens":1051,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:20:54.518203+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the paired t-tests on the per-fold accuracies given for ISIC-2017 (Table 7) and compare with Tables 8-11: those tables exactly repeat the ISIC-2016 numbers, so unless the 2017 statistics are recomputed from 2017 folds, the claimed statistical significance is not supported. A simpler check is whether any independent run reproduces the reported 97.95% Inception+Wavelet+MGTO accuracy on the official ISIC-2017 test split.","supporting_citations":[{"cited_title":"A., Codella, N","cited_arxiv_id":null,"evidence_quote":"Supplies the ISIC-2016 benchmark used for the headline 98.11% results."},{"cited_title":"H., & Lee, S","cited_arxiv_id":null,"evidence_quote":"Lai et al. is the prior IGWO-based skin-cancer method whose reported accuracy the paper must beat."},{"cited_title":"A., Rehman, A., Iqbal, M., & Saba, T","cited_arxiv_id":null,"evidence_quote":"Nawaz et al. provides a deep-learning comparison baseline on the same datasets."},{"cited_title":"C., Turk, V ., Khoshelham, K., & Kaya, S","cited_arxiv_id":null,"evidence_quote":"InSiNet provides another comparison baseline in the accuracy tables."},{"cited_title":"M., & Rashid, T","cited_arxiv_id":null,"evidence_quote":"Defines the Fox optimizer, one of the three swarm optimizers central to the method."},{"cited_title":"R., Gaheen, M","cited_arxiv_id":null,"evidence_quote":"Defines the Modified Gorilla Troops Optimizer (MGTO), the third optimizer and the best choice for Inception on ISIC-2017."}],"review_version":1}