{"id":"ad31d5f5-5a41-4305-8671-2bdc69c6e7a8","arxiv_id":"2412.16758","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"For CT radiomics of lung nodules, harmonizing benign and malignant nodules separately or with a biological covariate beats collective harmonization at preserving features and predicting malignancy.","lead":"This preprint compares three ways of applying ComBat harmonization to CT radiomics of benign and malignant lung nodules. It concludes that treating these nodule types separately, or with a biological covariate, preserves more acquisition-independent features and improves cancer prediction compared with standard collective harmonization.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 90.9% acquisition-independence rate for separate harmonization is computed on the same training data used to fit ComBat, so the central claim lacks a valid out-of-sample demonstration.","rationale":"The reader's weakest assumption was the future-derived LCS labels, which is a serious problem for the predictive comparison. The reader also mentioned in the rationale that Kruskal-Wallis was evaluated on the same training data used to fit separate harmonization, but did not make this the central attack. I regard the in-sample Kruskal-Wallis metric as the most load-bearing issue because it is the direct evidence for the paper's mechanistic claim that benign and malignant nodules need different corrective transformations. The abstract/body mismatch is a presentation defect, and the label-leakage issue affects the secondary predictive result, but if the Kruskal-Wallis percentages are not out-of-sample, the first pillar of the central claim collapses even before considering labels. The proposed test is inexpensive because the train/test splits already exist and only requires applying the fitted ComBat estimators to test folds and rerunning the same statistical test. If separate harmonization does not retain its high acquisition-independence rate on held-out scans, the paper should be revised to make the (separately problematic) LCS predictive comparison primary rather than the Kruskal-Wallis percentages.","tokens_in":21827,"tokens_out":5479,"duration_ms":52962,"concrete_test":"Using the existing 49 trial train/test splits, apply the ComBat estimators fit on each training fold to the corresponding test fold, then run the same Kruskal-Wallis test (p ≤ 0.05, performed separately on benign and malignant subgroups) on the harmonized test features for collective, covariate, and separate harmonization. Report the three acquisition-independence percentages. If separate harmonization's percentage falls to the level of covariate or collective harmonization once evaluated out-of-sample, the reported 90.9% is an in-sample artifact and the paper's central claim would need to be re-evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative evidence that benign and malignant nodules require different ComBat corrections is the Kruskal-Wallis acquisition-independence rate: 2.1% collective, 27.3% covariate, and 90.9% separate. This rate is computed by applying Kruskal-Wallis to the same training set on which the ComBat estimators were fit (Methods C: 'We performed Kruskal-Wallis on the ... training set'). For separate harmonization, ComBat fits subgroup-specific location and scale corrections for each acquisition parameter instance; testing those corrections on the very data used to fit them does not measure whether the harmonized features are acquisition-independent in unseen scans. The 90.9% figure can be inflated by in-sample fitting even if separate harmonization fails to generalize. The only out-of-sample evidence for the central claim is the LCS predictive comparison, whose labels are future-derived and may be correlated with acquisition (e.g., contrast-enhanced scans are ordered when there is concern). Thus the central claim currently lacks clean out-of-sample support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether radiomic features of benign and malignant pulmonary nodules require different ComBat harmonization transformations when correcting for CT acquisition variability, and compares three strategies: collective harmonization, covariate harmonization that preserves subgroup distinctions, and separate subgroup harmonization. Using 567 chest CT scans partitioned into benign, malignant, and lung-cancer-screening (LCS) subgroups, the authors measure post-harmonization acquisition dependence with Kruskal-Wallis tests and train LASSO-SVM classifiers on acquisition-independent features. They report that separate harmonization leaves 90.9% of features acquisition-independent versus 27.3% for covariate and 2.1% for collective harmonization, and that covariate and separate models outperform collective models on LCS test scans. The Discussion acknowledges that LCS labels are future-derived and that the authors cannot know whether a later-malignant nodule was malignant at scan time.","tokens_in":21954,"tokens_out":4195,"duration_ms":41692,"significance":"If the central claim is correct, the paper would fill a real gap: standard ComBat practice for diagnostic radiomics typically ignores possible differences in acquisition effects between benign and malignant tissue, and the paper's proposed biology-aware harmonization could improve multi-site diagnostic model transfer. The study has several commendable design choices: identical training/test folds across harmonization methods, 49 repeated cross-validation trials, Holm-Bonferroni adjustment for multiple comparisons, detailed feature-level tables (Tables S2-S3), and a control experiment using uniform-acquisition groups to sanity-check the Kruskal-Wallis procedure. These features make the paper's framework reproducible in principle. However, the two load-bearing pieces of evidence for the central claim are the Kruskal-Wallis acquisition-independence rate, computed on the same training data used to fit ComBat, and the LCS predictive comparison, whose labels are future-derived and plausibly correlated with contrast enhancement. Both need out-of-sample or sensitivity support before the conclusions can be accepted.","major_comments":[{"comment":"The abstract reports a training-set augmentation experiment and specific ROC-AUC values (0.74 [0.69-0.79] for covariate and 0.71 [0.66-0.77] for separate harmonization) that do not appear anywhere in the full text. The Results section reports only DeLong-test significance percentages for ROC-AUC comparisons and gives accuracy/sensitivity/specificity in Table 3, not the AUC values or their confidence intervals. Since the predictive comparison is the only out-of-sample evidence in the paper, the primary numeric result must be present in the Results and reproducible from Figure 3.","section":"Abstract and Results"},{"comment":"The 90.9% acquisition-independence rate for separate harmonization is computed by applying the Kruskal-Wallis test to the same training set on which the ComBat estimators were fit: Methods C states 'We performed Kruskal-Wallis on the ... training set.' This is a self-consistency check, not a measure of whether the harmonized features are acquisition-independent in unseen scans. The in-sample fit can inflate the reported rate even if separate harmonization fails to generalize, because ComBat's subgroup-specific location and scale corrections are optimized on those exact observations. The authors should apply the fitted harmonization estimators to held-out test folds and perform Kruskal-Wallis there, or validate on an independent acquisition-uniform cohort, before claiming that separate harmonization recovers acquisition-independent distributions.","section":"Methods C and Results"},{"comment":"The LCS predictive comparison is the only out-of-sample support for the central claim, but its labels are future-derived and may be correlated with acquisition parameters. The Discussion acknowledges that 'we could not know whether these were already malignant at scan time,' and Table S1 shows a strong imbalance in contrast enhancement: 132/186 malignant scans are contrast-enhanced versus 13/323 LCS scans and 32/58 benign scans. Since the paper itself notes that contrast is typically administered only when there is already concern (Introduction), the future malignancy labels on LCS scans may track the acquisition protocol. If so, the reported advantage of covariate and separate harmonization over collective harmonization could reflect preservation of an acquisition-label artifact rather than true biological signal. A sensitivity analysis restricting LCS evaluation to non-contrast scans, or using a ground-truth cohort with contemporaneous diagnosis, is needed.","section":"Discussion and Table S1"}],"minor_comments":[{"comment":"The Discussion states that separate harmonization models had 'lower specificity for test samples in the LCS subgroup,' but Results and Table 3 show that separate harmonization had higher specificity (96.1%) than covariate harmonization (90.3%) and lower sensitivity. The wording should be corrected to 'lower sensitivity.'","section":"Discussion"},{"comment":"The abstract mentions augmentation with 'later-development benign and malignant PNs (n=225),' but the full text never defines this number or explicitly reports a model trained only on early-development PNs. Please reconcile the abstract with the Methods and Results, or provide the augmentation experiment in the full text.","section":"Methods A and Table 1"},{"comment":"The text says '10 trials of 5-fold cross validation' but then excludes one trial and reports 49 trials. The exclusion criteria are described, but the relationship between the nominal 50 trials and the final 49 should be stated more directly, and Table 3 should clarify that collective-harmonization metrics are reported over only the 38 trials in which a model could be trained.","section":"Methods C"},{"comment":"The Kruskal-Wallis test is applied to 107 features and four acquisition parameters without any multiple-comparison adjustment. This affects all three harmonization methods similarly, but the 95% confidence intervals around the reported percentages would be more interpretable if the authors stated whether any multiplicity control was applied or why it was omitted.","section":"Methods C and Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant and well-posed methodological question, and the experimental scaffolding is mostly careful. My main concern is evidentiary rather than conceptual: the two load-bearing results are an in-sample harmonization check and a test set whose labels are future-derived and potentially confounded with contrast enhancement. These can be addressed with a revised analysis on held-out folds or a sensitivity analysis, so I do not see grounds for outright rejection, but the current version does not yet substantiate the central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Greg — quick read of arXiv:2412.16758. The thing to know: the abstract you were sent describes a training-set augmentation experiment with ROC-AUCs around 0.71–0.74, and the full text is a different study on 567 scans comparing three ComBat strategies. No augmentation experiment appears in the full text. That is not a cosmetic issue; it means the submitted claims and the manuscript do not match, and the paper cannot be evaluated as-is.\n\nWhat is genuinely new and worth keeping: applying Orlhac's tissue-type guidance to benign vs malignant forms of the same tissue is a real question, and this is a reasonable first attempt to test it on clinical CT. The design (nested ComBat over four acquisition parameters, three harmonization schemes, identical folds across methods) is sensible, and the sanity check that Kruskal-Wallis rarely rejects uniform-acquisition groups (94% non-significant) is a nice control. The finding that CE is the hardest parameter to harmonize also rings true.\n\nThe soft spots are serious. The stress-test note is correct: the 90.9% vs 27.3% vs 2.1% acquisition-independence rates come from Kruskal-Wallis on the same training set that produced the ComBat estimators. Separate harmonization fits subgroup-specific corrections, so testing them on their own training data mostly checks self-consistency, not generalization. The only genuinely out-of-sample evidence is the LCS predictive comparison, and there the labels are future-derived: 'benign' means no later diagnosis, 'malignant' means later diagnosis, and the authors admit they cannot know malignancy at scan time. That alone would be survivable, but the concern is that contrast-enhancement is ordered when there is suspicion, so the label-acquisition correlation can masquerade as biological signal. Also, unharmonized models achieve similar LCS metrics (p = 0.22, 0.33, 0.08), and the authors dismiss that as not the target — fair, but it weakens the predictive argument further. Reporting collective metrics on 38 of 49 trials is a smaller issue but should be fixed.\n\nCitation pattern is fine; no self-citation problem, and the key Orlhac/ComBat refs are there. The writing is careful about what it claims, which makes the abstract mismatch all the more puzzling.\n\nWho is it for: radiomics methodologists interested in ComBat preprocessing. A clinician should not act on it. I'd send it to peer review rather than desk reject, but with the clear expectation of major revision: align abstract and body, re-run the acquisition-independence evaluation on held-out data, and reframe the LCS result as proof-of-principle.","headline":"Worth asking the biology-aware harmonization question, but this version of the paper does not support its headline claims: the abstract and full text describe different experiments, and the central 90.9% acquisition-independence figure is measured on the same data used to fit ComBat.","tokens_in":22552,"tokens_out":2889,"would_cite":false,"duration_ms":27278,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Radiomic features of benign and malignant pulmonary nodules need different ComBat harmonization corrections, because a single collective correction destroys most predictive signal for lung cancer diagnosis.","keywords":["radiomics","ComBat harmonization","pulmonary nodules","lung cancer screening","CT imaging","acquisition variability","machine learning","benign versus malignant"],"falsifier":"Take a cohort in which every pulmonary nodule has a pathology-confirmed diagnosis at the moment of the CT scan rather than from future follow-up, split it into benign and malignant groups, and repeat the three harmonization pipelines. If collective harmonization then keeps most features acquisition-independent or matches the predictive performance of separate harmonization, the paper's central claim would be falsified. A cheaper probe is to stratify by contrast enhancement: if the advantage of separate harmonization disappears when only non-contrast scans are used, the effect may have been driven by label correlation with contrast rather than by a general benign/malignant difference in acquisition effects.","tokens_in":21574,"feed_emoji":"🫁","tokens_out":6457,"duration_ms":49688,"temperature":0.7,"pith_summary":"This paper asks a methodological question that must be answered before radiomics can be used to diagnose lung cancer: when a model has to combine CT scans of benign and malignant pulmonary nodules, should the standard correction for acquisition differences treat both groups as one pool? The authors show, on 567 chest CT scans, that it should not. They compare three ways of applying nested ComBat harmonization to remove dependence on contrast enhancement, scanner manufacturer, acquisition voltage, and focal spot size: one correction for all data, a correction that preserves a benign-versus-malignant covariate, and fully separate corrections per subgroup. Separate harmonization leaves 90.9% of features acquisition-independent, covariate harmonization 27.3%, and collective harmonization only 2.1%; classifiers built on separately or covariately harmonized features predict malignancy on screening scans with ROC-AUC around 0.71 to 0.74, while classifiers on collectively harmonized features collapse to calling everything benign. The conclusion is that acquisition effects themselves differ between benign and malignant nodules, so harmonization must be biology-aware. If this is right, it changes how radiomic lung-cancer models should be built and how earlier results that used collective harmonization should be read.","feed_headline":"Benign and malignant nodules need separate scan harmonization","feed_subtitle":"Standard correction leaves 2.1% of features usable; separate correction keeps 90.9% and rescues cancer prediction.","key_machinery":"The central machinery is Optimized Permutation Nested ComBat (OPNCB), a sequential application of ComBat that removes one acquisition parameter at a time in the order that frees the most features. It is paired with a Kruskal-Wallis test on the benign and malignant subgroups separately to decide whether a harmonized feature is acquisition-independent, and a LASSO feature selector plus linear SVM to turn the surviving features into a malignancy classifier. The setup carries the argument: by partitioning known-benign and known-malignant scans before harmonizing, the authors can measure how much predictive signal each harmonization strategy preserves.","core_discovery":"We show that radiomic features of benign and malignant pulmonary nodules require different corrective transformations to recover acquisition-independent distributions. Standard collective ComBat, which applies a common correction, removes the signal: only 2.1% of features remain acquisition-independent and the resulting LASSO-SVM model cannot separate the classes. Harmonizing the benign and malignant subgroups separately recovers 90.9% of features as acquisition-independent, and harmonizing with a covariate that distinguishes the subgroups recovers 27.3%; both allow predictive models that generalize to lung cancer screening scans, with no conclusive winner between the two. Contrast enhancement is the hardest protocol to correct and is the main source of residual dependence, which is consistent with iodine enhancement behaving differently in malignant tissue.","pith_inferences":["The same biology-aware harmonization logic may apply to other paired benign and malignant tissue contexts, such as breast or thyroid nodules, so the paper's conclusion could generalize beyond lung; this is unstated in the paper.","The benign-versus-malignant difference in acquisition effects is likely smallest for non-contrast screening scans; a testable consequence is that on low-dose non-contrast screening CT, collective harmonization may be adequate, and the paper's recommended methods matter most where contrast is used.","Because contrast enhancement is ordered when there is clinical concern, the LCS labels may correlate with contrast use; if that correlation drives the results, the true biological claim would be weaker. The paper itself acknowledges the label-timing issue, so this is an inference worth testing rather than a settled criticism."],"forward_implications":["Separately or covariately harmonized radiomic features, rather than collectively harmonized ones, should be used as inputs when training lung-cancer diagnostic models on mixed benign and malignant CT scans.","With a covariate, the model keeps sensitivity (60.8%) at the cost of specificity (90.3%); separate harmonization flips this (45.6% sensitivity, 96.1% specificity), so on screening populations the choice of method depends on which error is more costly.","Contrast enhancement is the acquisition parameter most likely to frustrate harmonization; studies that ignore contrast or harmonize benign and malignant scans together may underestimate or overestimate the portability of their radiomic models to other sites.","Successfully harmonized features can still be non-predictive; acquisition-independence is a constraint on feature choice, not a guarantee of diagnostic value."],"supporting_citations":[{"why":"Supplies the PyRadiomics feature extractor that produces the 107 radiomic features used throughout the study.","marker":"[9]"},{"why":"Introduces the ComBat batch-effect model that all harmonization methods build on.","marker":"[11]"},{"why":"Provides the Optimized Permutation Nested ComBat algorithm used to correct acquisition parameters sequentially.","marker":"[15]"},{"why":"The guide to ComBat harmonization that defines when a covariate versus separate harmonization is appropriate and is the methodological baseline the paper tests.","marker":"[16]"},{"why":"Demonstrates that distinct tissue types require distinct ComBat transformations, the precedent the paper extends to benign versus malignant forms of the same tissue.","marker":"[17]"},{"why":"A prior lung-cancer diagnosis radiomic model combining benign and malignant nodules without distinguishing them, the baseline the paper argues against.","marker":"[21]"}],"fun_headline_variants":["Separate nodule harmonization rescues lung cancer prediction","Biology-aware harmonization beats standard combat for nodules","Harmonize benign and malignant nodules separately for CT prediction","Standard scan correction fails, biology-aware harmonization wins","Nodule type matters in CT harmonization for cancer prediction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The clinical test set (the LCS subgroup) is labeled as benign or malignant using future information: a scan labeled malignant was later diagnosed as cancer, and a scan labeled benign merely had no later cancer diagnosis, so the label is not the nodule's true state at scan time.","fun_headline_variants_meta":{"raw":{"variants":["Separate nodule harmonization rescues lung cancer prediction","Biology-aware harmonization beats standard combat for nodules","Harmonize benign and malignant nodules separately for CT prediction","Standard scan correction fails, biology-aware harmonization wins","Nodule type matters in CT harmonization for cancer prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000568,"raw_usage":{"total_tokens":2719,"prompt_tokens":1008,"completion_tokens":1711,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":1632}},"tokens_in":624,"tokens_out":1711,"duration_ms":11522,"temperature":1.0,"reasoning_tokens":1632,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:15:38.841457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a cohort in which every pulmonary nodule has a pathology-confirmed diagnosis at the moment of the CT scan rather than from future follow-up, split it into benign and malignant groups, and repeat the three harmonization pipelines. If collective harmonization then keeps most features acquisition-independent or matches the predictive performance of separate harmonization, the paper's central claim would be falsified. A cheaper probe is to stratify by contrast enhancement: if the advantage of separate harmonization disappears when only non-contrast scans are used, the effect may have been driven by label correlation with contrast rather than by a general benign/malignant difference in acquisition effects.","supporting_citations":[{"cited_title":"Improved generalized ComBat methods for harmonization of radiomic features","cited_arxiv_id":null,"evidence_quote":"Provides the Optimized Permutation Nested ComBat algorithm used to correct acquisition parameters sequentially."},{"cited_title":"External validation of radiomics-based predictive models in low-dose CT screening for early lung cancer diagnosis","cited_arxiv_id":null,"evidence_quote":"A prior lung-cancer diagnosis radiomic model combining benign and malignant nodules without distinguishing them, the baseline the paper argues against."}],"review_version":1}