{"id":"e136170d-357d-4b38-9afd-af8f1e558977","arxiv_id":"2411.16008","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Expanding the segmented nodule by 8 mm into surrounding tissue gave the best radiomics classification result, AUC 0.78.","lead":"This study tested whether adding a ring of tissue around a lung nodule, the peritumoral region, improves automated cancer classification on CT scans. It reports that an 8 mm expansion gave AUC 0.78, slightly better than the compared deep learning models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on validation-set AUCs with overlapping confidence intervals and post hoc selection among six expansion distances; without significance testing or reported independent test-set results, the 8 mm improvement may be chance variation.","rationale":"The reader's weakest assumption is the load-bearing one: the observed improvement from 0.73 to 0.78 could reflect chance or selection over expansion distances, and no significance test or multiple-comparison correction is provided. My reading confirms this and adds that the absence of numeric test-set AUCs means even the basic generalizability check is missing from the paper. Methodologically, the pipeline is described clearly, and the peritumoral expansion idea is plausible; the issue is not internal inconsistency but insufficient statistical support for the central claim. Therefore the appropriate disposition is the same conditional verdict: the paper can be accepted only after the authors provide a pre-specified or cross-validated expansion distance, significance tests on an independent test set, and ideally code to reproduce the numbers. I agree with the reader's assessment rather than proposing a harsher verdict, because the data could support the claim if the requested tests are supplied.","tokens_in":5526,"tokens_out":4092,"duration_ms":33960,"concrete_test":"Request the held-out test-set AUC for each expansion distance (0, 2, 4, 6, 8, 10, 12 mm) using the same KNN segmentation and Logistic Regression protocol, along with a paired DeLong or bootstrap test comparing 8 mm versus nodule-only and 8 mm versus 6 mm and 10 mm on that test set. If the test-set 8 mm AUC is not numerically the best and the comparison with nodule-only is not significant after correcting for six comparisons, the central claim that peritumoral expansion up to 8 mm provides critical diagnostic information is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim—that peritumoral expansion up to 8 mm 'significantly enhanced' classification (nodule-only AUC 0.73 vs. 8 mm AUC 0.78)—is not supported by the statistical evidence presented. The reported AUCs appear to come from the validation set (Table 2 is explicitly validation; the peritumoral paragraph does not state a different split), and their 95% confidence intervals (0.63–0.82 vs. 0.69–0.87) overlap substantially. Six expansion distances were evaluated on the same validation data, and 8 mm was selected as the best; no multiple-comparison correction or significance test (e.g., DeLong or bootstrap) is reported. Figure 3 is captioned 'training and testing results,' but the text gives no numeric test-set AUCs for the peritumoral expansions, so the reader cannot verify that the selected 8 mm model was confirmed on an independent split. Because the central claim depends on a small AUC difference selected post hoc, the most parsimonious explanation consistent with the reported numbers is chance variation. The deep-learning comparisons (FMCB 0.71, ResNet50-SWS++ 0.71) are taken from prior work and not re-run under the same protocol, but the load-bearing issue is the within-study statistical comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript evaluates four 3D segmentation methods (Otsu, Fuzzy C-Means, Gaussian Mixture Model, K-Nearest Neighbors) and three classifiers (Random Forest, Logistic Regression, KNN) for radiomics-based lung nodule classification on the Duke Lung Cancer Screening Dataset. It reports that KNN segmentation with Logistic Regression performs best on the validation set (AUC 0.89), and that expanding the segmentation into peritumoral regions improves validation AUC from 0.73 (nodule only) to 0.78 at 8 mm, with some decline at larger expansion distances. The authors compare these results with deep learning baselines from prior work (FMCB, ResNet50-SWS++, Genesis, MedNet3D) and conclude that radiomics with peritumoral expansion outperforms them.","tokens_in":5775,"tokens_out":4258,"duration_ms":37312,"significance":"If properly validated, the hypothesis that peritumoral expansion up to a specific distance adds incremental diagnostic information is valuable for radiomics pipeline design and for interpreting tumor microenvironment effects in CT-based classification. The study uses a public dataset and PyRadiomics, and the multi-segmentation and multi-classifier comparison is a useful contribution. The main limitation is statistical: the central claim rests on a post hoc selected expansion distance, overlapping confidence intervals, and the absence of significance testing or independent test-set confirmation. As presented, the evidence does not establish that the observed AUC difference is real rather than chance variation. The deep-learning comparison is also weakened by reliance on previously published numbers without re-running baselines in the same pipeline.","major_comments":[{"comment":"The claim that incorporating peritumoral regions \"significantly enhanced\" performance is not supported by the reported statistics. The 95% confidence intervals for nodule-only AUC (0.63-0.82) and 8 mm expansion AUC (0.69-0.87) overlap substantially, and no significance test (e.g., DeLong or bootstrap) is provided. Moreover, the 8 mm result was selected post hoc after evaluating six expansion distances on the same validation data; without multiple-comparison correction or a pre-specified hypothesis, the reported improvement is consistent with selection over expansion distances.","section":"Analysis of Peritumoral Expansion on Classification Performance"},{"comment":"There is a numerical inconsistency that needs explanation: Table 2 reports a validation AUC of 0.89 for Logistic Regression with KNN segmentation, while the peritumoral analysis reports an AUC of 0.73 for the original nodule segmentation using the same combination. The manuscript does not clarify whether the two analyses use different feature sets, sample sizes, or preprocessing. This discrepancy undermines the reader's ability to interpret the expansion effect and should be resolved before the central claim can be assessed.","section":"Results & Discussion, Table 2"},{"comment":"Figure 3 is captioned \"training and testing results of Logistic Regression with expanded segmentation,\" but the text reports no numeric test-set AUCs for the peritumoral expansion distances. Since the dataset was split into training, validation, and testing sets, and since the 8 mm distance was selected on the validation set, the authors must report the test-set AUC for the pre-specified 8 mm expansion (and ideally for all expansion distances) to demonstrate that the selected distance was not an artifact of validation-set overfitting. Without this, the central claim is not independently confirmed.","section":"Peritumoral Expansion and Figure 3"},{"comment":"The deep-learning comparison is based on AUC values taken from prior publications rather than models re-run under the same protocol in this study. The confidence intervals for the compared methods overlap (e.g., 8 mm radiomics AUC 0.69-0.87 vs. ResNet50-SWS++ 0.61-0.81 and FMCB 0.60-0.82), so the statement that the radiomics approach demonstrates \"superior classification accuracy\" is not supported by the reported numbers. Re-running the baselines on the same data split or restricting the claim to descriptive observations would be necessary.","section":"Comparison with Deep Learning Models"}],"minor_comments":[{"comment":"The word \"radionics\" appears in the abstract and methods and should be corrected to \"radiomics.\"","section":"Abstract and Methods"},{"comment":"The text refers to \"fig. 2(b)\" but no Figure 2 appears in the manuscript; the reference should be corrected or the figure added.","section":"Analysis of Peritumoral Expansion on Classification Performance"},{"comment":"The caption states that error bars indicate 95% confidence intervals for each model, but confidence intervals are not reported in the text or table for all deep learning models (e.g., Genesis and MedNet3D only have point AUCs).","section":"Figure 4"},{"comment":"The code availability statement says the code \"will be openly available\" rather than providing a permanent repository link or DOI; if the code is intended to support reproducibility, a stable link should be provided.","section":"Dataset and Code Availability"},{"comment":"Minor formatting issues include a stray percent sign in the \"Malignant\" row and inconsistent use of decimal places across percentage values; these should be cleaned up.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a worthwhile question and uses a public dataset, but the central claim is currently under-supported by the statistical evidence. In particular, the post hoc selection of the 8 mm expansion distance, overlapping confidence intervals, and the missing test-set confirmation are load-bearing issues that require a proper revision rather than copy-editing. The deep-learning comparison also leans heavily on previously published numbers, so I would ask the editor to verify that those baselines were evaluated on the same data split and under comparable conditions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a clean, honest exploratory study on a public dataset, but the headline claim doesn't survive contact with the statistics. The paper applies peritumoral expansion radiomics to the Duke lung cancer screening dataset, comparing four segmentation methods and six expansion distances. That's a legitimate benchmark, not a new idea — peritumoral radiomics has been around for years. The contribution here is the systematic sweep on this particular dataset.\n\nWhat the paper does well: it uses a public dataset, reports 95% CIs for the AUCs, and the protocol is straightforward enough to reproduce if the code appears. The observation that AUC rises from 0.73 to 0.78 as you expand from 0 to 8 mm and then falls is a reasonable pattern worth reporting.\n\nThe soft spots are serious. The 8 mm distance was chosen after seeing the results on the same validation data; no multiple-comparison correction or significance test is reported, and the CIs (0.63–0.82 vs 0.69–0.87) overlap substantially. The word 'significantly' in the abstract is not backed by any test. More troubling, Table 2 reports Logistic Regression with KNN segmentation at AUC 0.89 on the validation set, but the peritumoral analysis starts from nodule-only AUC 0.73. The paper never reconciles those numbers, so it's unclear whether the comparisons are on the same split. The deep learning baselines come from prior work, not re-run here, so if the radiomics numbers are from a different split than the DL numbers, the comparison is unfair. Code is promised but not yet available.\n\nThis is a useful exploratory benchmark, not a demonstrated result. It could be made sound with a preregistered expansion distance (or cross-validation), DeLong or bootstrap tests, a clear split protocol, and released code. Right now the central claim is not supported. I'd send it to peer review anyway — the topic is relevant and the fixes are feasible — but I wouldn't cite it as evidence until the statistics are cleaned up. Reading group? Maybe, as a cautionary tale about post hoc selection.","headline":"Useful exploratory benchmark on a public dataset, but the 8 mm claim is post hoc and the statistics don't back the word 'significant'.","tokens_in":6296,"tokens_out":2673,"would_cite":false,"duration_ms":24873,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding an 8 mm shell of tissue around a segmented lung nodule improves radiomics-based cancer classification to an AUC of 0.78, beating the deep-learning baselines compared in the study.","keywords":["radiomics","peritumoral expansion","lung cancer classification","CT segmentation","K-nearest neighbors","logistic regression","tumor microenvironment","deep learning baselines"],"falsifier":"Run the same KNN-segmentation and logistic-regression pipeline on an independent external lung CT dataset with the expansion fixed at 8 mm, and compare against nodule-only features with a permutation test over expansion distances; if the 8 mm AUC gain is within the chance distribution of best-of-six distances, or the confidence intervals still overlap, the central claim would be unsupported.","tokens_in":5313,"feed_emoji":"🫁","tokens_out":7908,"duration_ms":66562,"temperature":0.7,"pith_summary":"Peritumoral tissue is not noise: this paper claims that a shell of tissue around a lung nodule carries information about whether the nodule is cancerous, and that the best shell thickness is about 8 mm. Using CT scans from an open lung-screening dataset, the author segments nodules with four algorithms, extracts quantitative imaging features, and trains simple classifiers. The best combination, K-nearest-neighbors segmentation with logistic regression and an 8 mm expansion, reaches an area under the ROC curve of 0.78 on validation data, above the 0.71 of the deep-learning baselines compared. The reason to care is practical: if true, a transparent radiomics pipeline with a fixed expansion shell could strengthen lung cancer screening without needing heavy deep networks.","feed_headline":"8 mm of tissue around a nodule lifts lung cancer AUC to 0.78","feed_subtitle":"Radiomics with peritumoral context beats deep-learning patch models on a public lung-screening CT benchmark.","key_machinery":"Peritumoral expansion is the mechanism: a 3D nodule mask, produced here by K-nearest-neighbors voxel clustering, is dilated outward by fixed distances of 2, 4, 6, 8, 10, and 12 mm; radiomics intensity, texture, and shape features are re-extracted from each dilated mask; and a logistic regression classifier converts them into a cancer score. The expansion isolates the contribution of the surrounding shell because the segmentation and feature set stay fixed while only the shell thickness changes, while the KNN segmentation is what defines the boundary the shells grow from. The 8 mm shell is the paper's preferred operating point.","core_discovery":"The central discovery is that radiomics features extracted from a nodule plus a surrounding peritumoral shell classify lung cancer better than features from the nodule alone, and the gain grows with shell thickness up to 8 mm before reversing. In the validation set, nodule-only KNN segmentation with logistic regression gives AUC 0.73 (95% CI 0.63–0.82); expanding by 4, 6, and 8 mm yields 0.75 (0.64–0.85), 0.77 (0.68–0.86), and 0.78 (0.69–0.87), while 12 mm drops to 0.70 (0.58–0.81). The best radiomics configuration outperforms the deep-learning image baselines, whose top AUC is 0.71, and the author attributes the improvement to tumor-microenvironment characteristics captured in the 8 mm shell, with further expansion introducing noise. This is presented as evidence that contextual tissue information, systematically harvested, is more useful for this classification task than the patch-level deep features that dominated the comparison baselines.","pith_inferences":["Editorial inference: the 8 mm optimum is likely dataset- and nodule-size-dependent; a fixed shell may under- or over-sample for very small or very large nodules, so an adaptive shell size is a plausible next experiment.","Editorial inference: because the reported confidence intervals for nodule-only and 8 mm AUC overlap, and the best distance was chosen from six candidates, a permutation or multiple-comparison correction could show the 8 mm advantage is smaller than it appears.","Editorial inference: a natural testable extension is to concatenate peritumoral radiomics features with deep-learning embeddings from a foundation-model feature extractor; if the two signal sources are complementary, the combined model should exceed both.","Editorial inference: the decline beyond 8 mm hints that peritumoral signal is concentrated in a limited shell; mapping where the information lives (e.g., by 1 mm increments or by anatomic compartment) could sharpen the finding and guide feature selection."],"forward_implications":["Radiomics pipelines for lung nodule classification should include a peritumoral shell, with 8 mm as the default thickness.","The expansion distance behaves like a signal-to-noise dial: performance rises to 8 mm and falls at 12 mm, so distance should be tuned rather than assumed.","A simple, interpretable radiomics model can match or beat deep-learning patch classifiers on the same benchmark, making it a viable baseline and a candidate for clinical deployment where interpretability matters.","Segmentation quality matters: KNN segmentation paired with logistic regression gave the highest validation AUC (0.89) before expansion, so segmentation choice is a driver of radiomics performance."],"supporting_citations":[{"why":"supplies the open-access lung-screening CT dataset, the deep-learning baseline results, and the training/validation split used throughout.","marker":"[4]"},{"why":"provides the foundation-model feature extractor whose AUC of 0.71 is the key deep-learning baseline to beat.","marker":"[6]"},{"why":"is the source text for the four segmentation algorithms (Otsu, Fuzzy C-Means, GMM, KNN) applied to the nodules.","marker":"[8]"},{"why":"is the radiomics feature extraction library that produces the intensity, texture, and shape features from each segmentation mask.","marker":"[9]"},{"why":"packages the ROC analysis used to compute and compare AUC values.","marker":"[12]"}],"fun_headline_variants":["Peritumoral shell boosts lung nodule AUC to 0.78","Radiomics with 8mm margin beats deep nets in lung cancer","8mm peritumoral zone best for lung cancer radiomics","Lung nodule radiomics: 8mm expansion best, 12mm loses","Tumor edge context lifts lung cancer AUC to 0.78"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the improvement from adding an 8 mm shell (AUC 0.78 vs 0.73 for the nodule alone) is a genuine biological signal rather than the result of testing six expansion distances and reporting the best one, since the 95% confidence intervals overlap and no significance test or multiple-comparison correction is provided.","fun_headline_variants_meta":{"raw":{"variants":["Peritumoral shell boosts lung nodule AUC to 0.78","Radiomics with 8mm margin beats deep nets in lung cancer","8mm peritumoral zone best for lung cancer radiomics","Lung nodule radiomics: 8mm expansion best, 12mm loses","Tumor edge context lifts lung cancer AUC to 0.78"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1556,"prompt_tokens":1069,"completion_tokens":487,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":685,"completion_tokens_details":{"reasoning_tokens":391}},"tokens_in":685,"tokens_out":487,"duration_ms":4147,"temperature":1.0,"reasoning_tokens":391,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:38:14.033323+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same KNN-segmentation and logistic-regression pipeline on an independent external lung CT dataset with the expansion fixed at 8 mm, and compare against nodule-only features with a permutation test over expansion distances; if the 8 mm AUC gain is within the chance distribution of best-of-six distances, or the confidence intervals still overlap, the central claim would be unsupported.","supporting_citations":[{"cited_title":"Foundation model for cancer imaging biomarkers,","cited_arxiv_id":null,"evidence_quote":"provides the foundation-model feature extractor whose AUC of 0.71 is the key deep-learning baseline to beat."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the source text for the four segmentation algorithms (Otsu, Fuzzy C-Means, GMM, KNN) applied to the nodules."},{"cited_title":"Computational radiomics system to decode the radiographic phenotype,","cited_arxiv_id":null,"evidence_quote":"is the radiomics feature extraction library that produces the intensity, texture, and shape features from each segmentation mask."}],"review_version":1}