{"id":"1be8fe68-7ee5-49ed-85d9-66931569851b","arxiv_id":"2508.18612","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Combined training on public and private PET-CT data gives the most balanced tumor segmentation across oesophageal, lung, and AutoPET test sets, supporting data diversity over architecture.","lead":"This paper tests whether training a standard tumor-segmentation AI on combined public and private PET-CT datasets makes it generalize across cancer types better than training on a single dataset. The authors introduce two new expert-annotated cohorts (oesophageal, Australia; lung, India) and report that combined training gives the most balanced performance across all test sets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The cross-cancer conclusion is confounded: cancer type is entangled with scanner, protocol, population, and annotation differences between cohorts, so the observed transfer failures do not isolate cancer-type overfitting; the diversity-over-novelty framing additionally lacks an architecture…","rationale":"The reader's weakest assumption identifies exactly the confound I consider most load-bearing: the oesophageal and lung cohorts differ not only in cancer type but in site, protocol, population, and annotation practice. My read of the paper is that the combined-training observation is plausible and directionally consistent with prior AutoPET findings, so I do not object to the empirical benchmark itself. The problem is interpretive: the abstract and conclusion claim that multi-demographic, multi-center, and multi-cancer integration is the key driver of robust generalization, but the experimental contrast cannot separate 'multi-cancer' from 'multi-center/protocol' effects. The proposed AutoPET-only cross-cancer test is feasible with public data and would settle whether cancer type alone produces the observed transfer failure. The lack of confidence intervals is a secondary but real concern, especially because the combined model's lung improvement over AutoPET-only is only 1.3 DSC points. The absence of an alternative architecture also means the 'outweighs architectural novelty' phrase is not directly tested. None of this changes the reader's CONDITIONAL verdict; it reinforces it, because the core balanced-training result is credible but the generalization claim requires the additional controls and uncertainty quantification already requested.","tokens_in":8193,"tokens_out":6396,"duration_ms":68226,"concrete_test":"Run a cross-cancer transfer experiment entirely within the AutoPET dataset, which shares scanner and annotation protocols: train a 3D nnU-Net with the same settings as Section 2.2 on non-lung AutoPET cases (lymphoma and melanoma) and test on held-out AutoPET lung cases; conversely, train on lung-only AutoPET cases and test on held-out non-lung cases, using the official AutoPET challenge split. If the non-lung-trained model retains high DSC on lung cases (within about 10 points of the lung-trained model), then cancer-type shift alone is weak and the OC-only collapse in Table 2 must be attributed to center/protocol/population confounds; if the drop is large, the cross-cancer interpretation is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that multi-cancer dataset diversity, rather than architectural novelty, drives robust generalization. The evidence for the 'multi-cancer' part rests on contrasting the Australian oesophageal cohort (n=279, annotated by three surgical oncology fellows) with the Indian lung cohort (n=54, annotated by a nuclear medicine physician), as described in Section 2.1. Table 2 shows the OC-only model collapsing on the lung cohort (DSC=1.25) despite strong in-domain performance (DSC=57.8), and Section 3.1 attributes this to 'cancer-type and site-specific overfitting.' However, cancer type is perfectly confounded with scanner hardware, acquisition protocol, reconstruction settings, patient demographics, and annotation label definitions. No scanner metadata, per-center stratification, or statistical adjustment is reported, so the OC-only collapse and the combined model's improvement could reflect general domain shift or a data-volume effect rather than cancer-type specificity. This is load-bearing because the headline conclusion about cross-cancer generalization depends on that attribution. A related weakness is that the combined model's lung DSC (52.9) is only 1.3 points above the AutoPET-only model (51.6), a gap well within plausible uncertainty given that no confidence intervals or significance tests are provided. The 'most balanced' conclusion thus rests on aggregate point estimates. Finally, no alternative architecture is evaluated, so the claim that dataset diversity 'outweighs architectural novelty' is an extrapolation from AutoPET challenge findings rather than a result of this study.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces two private PET-CT datasets (Australian oesophageal cancer, n=279; Indian lung cancer, n=54) and uses 3D nnU-Net v2 to compare three training configurations: target-only (oesophageal), public-only (AutoPET), and combined training. The models are evaluated on the oesophageal test set, the AutoPET test set, and the Indian lung cohort using DSC, precision, recall, and HD95. The reported results show that the oesophageal-only model is best in-domain but collapses on external cohorts, the AutoPET-only model generalizes moderately but underperforms on oesophageal cases, and the combined model gives the most balanced performance across cohorts. The paper concludes that dataset diversity, rather than architectural novelty, is the primary driver of robust generalization.","tokens_in":8402,"tokens_out":6562,"duration_ms":60016,"significance":"If the results hold, the two expert-annotated datasets and the three-way training comparison are a useful contribution to benchmarking PET-CT lesion segmentation under domain shift. The combined-training result provides a reproducible baseline for future work on oesophageal and lung cancer segmentation. However, the headline claim that dataset diversity 'outweighs architectural novelty' is not tested by the experimental design, and the cross-cancer interpretation is confounded by site and protocol differences. The paper would benefit from uncertainty quantification and more careful scoping of its claims.","major_comments":[{"comment":"The inference that the OC-only model's collapse on the lung cohort (DSC 1.25) is due to 'cancer-type and site-specific overfitting' is confounded. Cancer type is entangled with scanner hardware, acquisition protocol, reconstruction settings, patient demographics, and annotation practice between the Australian oesophageal cohort and the Indian lung cohort. The manuscript reports no scanner metadata, no per-center stratification, and no statistical adjustment, so the observed failure could be a generic domain-shift or data-composition effect rather than evidence about cross-cancer specificity. Please provide acquisition metadata and per-center analyses that isolate cancer type, or explicitly narrow the conclusion to cross-cohort generalization.","section":"Section 2.1 and Section 3.1 (Table 2)"},{"comment":"The claim that 'dataset diversity ... outweighs architectural novelty' is not supported by the experimental design. No alternative architecture is evaluated; all models use the same 3D nnU-Net and differ only in training-data composition. The results can at most support the conclusion that, within nnU-Net, training-data composition affects cross-domain generalization. An architectural comparison or a revised, accurately scoped claim is required.","section":"Abstract and Section 4 (Discussion)"},{"comment":"The 'most balanced' conclusion rests on aggregate point estimates without confidence intervals, significance tests, or per-cohort variance summaries. For example, the combined model's lung DSC (52.9) is only 1.3 points above the AutoPET-only model (51.6), a difference likely within sampling variability given the small external cohort (n=54). Please report bootstrap confidence intervals and paired tests where appropriate, or explicitly state that the observed differences are not statistically assessed.","section":"Section 3.1, Table 2"},{"comment":"The data quantities are internally inconsistent. A 70:30 split of 279 oesophageal cases gives approximately 195 training and 84 test cases, but Table 1 lists 210 train and 69 test (while its caption says 200 train). Similarly, the introduction describes AutoPET as 1,014 studies, but Table 1 uses 324 training and 139 test cases, whose sum (463) is not 70% of 1,014. Please clarify the actual inclusion criteria, selection procedure, and exact split for both the private and public datasets, and explain any discrepancy with the cited AutoPET corpus.","section":"Table 1 and Section 2.1"}],"minor_comments":[{"comment":"The phrase 'mean DSC = lung (52.9); oesophageal (40.7); AutoPET (60.9)' is awkward; please reword as 'lung: 52.9; oesophageal: 40.7; AutoPET: 60.9'.","section":"Abstract"},{"comment":"There are several typos and grammar issues: 'commitee', 'annoataions', 'sybtypes', and the sentence 'we have used two RTX A6000 NVIDIA GPU card with 24 GB memory each for training' is a fragment.","section":"Section 2.1 and Section 2.2"},{"comment":"The text refers to 'Figure 2 describes the number of images used for training and testing', but Figure 2 is captioned 'Representation of the datasets' and the split counts appear in Table 1; please correct the cross-reference. Also, the demographic and radiomics panels described in Section 2.1 are not clearly labeled in the figure caption.","section":"Section 2.1 and Figure 2"},{"comment":"The abstract reports the OC-only model's external performance as 'mean DSC ≤ 3.4', but Table 2 gives 1.25 for the lung cohort and 3.4 for AutoPET; please use exact values or clarify the bound.","section":"Abstract and Table 2"},{"comment":"Quantitative example values in the qualitative analysis are given as proportions (e.g., DSC = 0.89) while Table 2 reports DSC as percentages (e.g., 57.8); please use a consistent scale throughout.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The two new datasets and the clear three-way comparison are potentially valuable for the community, but the manuscript's central claims need reworking: the 'diversity outweighs architecture' conclusion is not tested without an architectural comparison, and the cross-cancer interpretation is confounded by site-specific factors. The internal inconsistencies in the reported dataset sizes also need to be resolved before the paper can be considered further."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The first thing you should know: this paper's real contribution is the data. The oesophageal cancer PET-CT cohort (n=279, Australian, expert-annotated) appears to be the first of its kind for segmentation research, and the Indian lung cohort (n=54) is also new. These are not public, but they are still the kind of resource the field needs, and the authors should get credit for curating them.\n\nWhat the paper actually shows is straightforward and credible: a 3D nnU-Net trained on oesophageal data alone collapses on the lung cohort (DSC ~1.3), an AutoPET-only model generalizes but does poorly on oesophageal (DSC 26.7), and the combined model gives the most balanced results across all three test sets. That pattern is consistent with the existing literature, and the authors do not over-egg the numbers themselves. The error analysis in Section 3.4 is also honest, noting small lesions, misregistered scans, and annotation pitfalls rather than blaming the model alone.\n\nThe soft spots are real but manageable. The abstract and discussion claim that \"dataset diversity outweighs architectural novelty,\" but no alternative architecture is evaluated anywhere in the paper. That claim is an extrapolation from the AutoPET challenge results, not a finding of this study, and it should be reworded. More importantly, the \"cross-cancer\" attribution is not actually isolated: the oesophageal and lung cohorts differ not just in cancer type but also in country, scanner hardware, acquisition protocol, and annotation staff. The OC-only collapse on the lung cohort could be domain shift in general, not cancer-type-specific overfitting. Given that the headline depends on this attribution, the authors need either to add per-center controls or soften the claim. The lack of confidence intervals or significance tests is also a problem, especially since the combined model's lung DSC is only 1.3 points higher than AutoPET-only (52.9 vs 51.6) — that difference is plausibly noise. And the private datasets are not released, which limits reproducibility, and the AutoPET split is not reconciled with the official challenge protocol.\n\nNone of this kills the paper. The practical takeaway — train on diverse public data plus domain-specific data if you want a robust PET-CT segmenter — is decent and supported by the aggregate numbers. But the framing goes beyond the evidence, and the confounds need to be addressed, either analytically or in the claims.\n\nI would send this to peer review. A serious referee could help the authors tighten the language and add uncertainty quantification. It is not a paper I would cite for the diversity-versus-architecture claim, but it is worth discussing as an example of a benchmark study with genuinely new data. For a reading group, I would probably include it.","headline":"A genuinely useful pair of new PET-CT datasets and a clean, honest benchmark, but the headline claim about diversity versus architecture is not actually tested and the cross-cancer interpretation is confounded by site and population.","tokens_in":9013,"tokens_out":1918,"would_cite":false,"duration_ms":19827,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Diverse training data, not architectural novelty, is the main driver of robust PET-CT tumor segmentation.","keywords":["PET-CT","tumor segmentation","nnU-Net","cross-cancer generalization","oesophageal cancer","lung cancer","multi-center data","AutoPET"],"falsifier":"A direct test would be to train the oesophageal-only model on additional oesophageal scans acquired with the same scanner and protocol used for the Indian lung cohort; if the model still scores near zero on the lung scans, cancer type is implicated, whereas a large improvement would show that site and acquisition differences, not cancer type, drove the failure.","tokens_in":7958,"feed_emoji":"🩻","tokens_out":5894,"duration_ms":54413,"temperature":0.7,"pith_summary":"This paper tries to establish that, for 3D PET-CT tumor segmentation, dataset diversity matters more than architectural novelty. It introduces two new expert-annotated whole-body datasets, 279 oesophageal cancer scans from an Australian cohort and 54 lung cancer scans from an Indian cohort, and uses them with the public AutoPET dataset to stress-test a 3D nnU-Net under three training regimes. The oesophageal-only model scores high on its own domain (mean DSC 57.8) but collapses on the external lung cohort (mean DSC below 3.4); the AutoPET-only model generalizes better but underperforms on oesophageal cases (mean DSC 26.7). The combined model gives the most balanced result (mean DSC 52.9 lung, 40.7 oesophageal, 60.9 AutoPET), with better boundary metrics. If the claim holds, clinical deployment should prioritize assembling diverse multi-center, multi-cancer training data over designing bespoke architectures.","feed_headline":"Diverse data, not fancier models, lift PET-CT segmentation","feed_subtitle":"Combined oesophageal+AutoPET training gives balanced Dice: 52.9 lung, 40.7 oesophageal, 60.9 AutoPET.","key_machinery":"The central object is the 3D nnU-Net, a self-configuring 3D U-Net for biomedical image segmentation that automatically determines patch size, architecture, and training schedule. The carrying argument is the three-way training comparison: a target-only model trained on oesophageal cancer, a public-only model trained on AutoPET, and a combined model trained on both, with all three evaluated on the same oesophageal test set, the independent lung cohort, and the reserved AutoPET test set. This setup isolates what each training-data composition contributes to in-domain accuracy and cross-domain transfer.","core_discovery":"The central claim is that a 3D nnU-Net trained on a combined dataset covering multiple cancers, centers, and patient populations generalizes more reliably across PET-CT domains than either a target-domain-only model or a public-only model. The oesophageal-only model shows severe cancer-type and site-specific overfitting: it reaches a mean DSC of 57.8 on its own test set but falls to 3.4 or below on the independent Indian lung cohort. The public-only model generalizes to 51.6 on the lung cohort but drops to 26.7 on oesophageal cases. The combined OC+AutoPET model delivers the most balanced outcome, with mean DSC 52.9 on lung, 40.7 on oesophageal, and 60.9 on AutoPET, along with lower Hausdorff boundary errors. The paper interprets this as quantitative evidence that multi-center, multi-cancer data diversity, rather than model complexity, is the key driver of clinically robust generalization.","pith_inferences":["Inference: because the two new cohorts differ in center, scanner, protocol, and annotation style as well as cancer type, the cleanest reading is that training-data diversity of any kind helps; separating cancer type from domain shift would require a dataset with all four combinations of cancer type and acquisition site.","Inference: a testable extension is to add oesophageal scans from the Indian site and lung scans from the Australian site; if combined training still balances performance, the conclusion is about diversity per se rather than the specific cancers.","Inference: the same three-way training protocol could be applied to other PET-avid cancers such as head-and-neck, prostate, or lymphoma to see whether one combined model can hold a portfolio of cancers above a clinical threshold.","Inference: the balanced model could serve as automated triage or region proposal, with radiologist review reserved for the small or low-contrast lesions identified in the error analysis."],"forward_implications":["A single-domain model should not be deployed outside its own cohort; the oesophageal-only model's near-zero lung performance is a concrete failure mode.","Combined training is the only tested configuration that keeps mean DSC above 40 on all three cohorts, making it a safer default for multi-cancer PET-CT workflows.","Architectural novelty is not the primary lever for robust generalization; effort spent on collecting and curating diverse data is likely to yield larger gains.","Small, low-contrast lesions remain a failure mode in every configuration, so human-in-the-loop review remains necessary for clinical use.","Improved boundary metrics under combined training matter for tasks such as radiotherapy planning, where edge accuracy determines dose delivery."],"supporting_citations":[{"why":"This citation supplies the nnU-Net method used for all three training configurations in the study.","marker":"[9]"},{"why":"This citation provides the public whole-body FDG-PET/CT dataset used as the public-only training set.","marker":"[7]"},{"why":"This citation establishes the AutoPET challenge conclusion that data quantity and quality matter more than architecture, the claim this paper extends to cross-cancer evaluation.","marker":"[6]"},{"why":"This citation defines the AutoPET challenge task and dataset used for the public and combined training arms.","marker":"[5]"},{"why":"This citation supports the claim that a well-configured nnU-Net can match more complex architectures in 3D segmentation.","marker":"[10]"},{"why":"This citation provides prior nnU-Net performance benchmarks on lung and head-and-neck PET/CT that motivate the baseline expectations.","marker":"[14]"}],"fun_headline_variants":["Data diversity, not model tweaks, drives PET-CT segmentation robustness","Multi-cancer PET-CT data yields most robust tumor segmentation","For PET-CT AI, dataset diversity outweighs model complexity","Combined training data improves PET-CT tumor segmentation across cohorts","Data mix beats architecture: PET-CT tumor segmentation improves"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the oesophageal-only model's collapse on the lung cohort is caused by cancer type and anatomical site, not by the scanner, protocol, population, and annotation differences that also separate the two cohorts.","fun_headline_variants_meta":{"raw":{"variants":["Data diversity, not model tweaks, drives PET-CT segmentation robustness","Multi-cancer PET-CT data yields most robust tumor segmentation","For PET-CT AI, dataset diversity outweighs model complexity","Combined training data improves PET-CT tumor segmentation across cohorts","Data mix beats architecture: PET-CT tumor segmentation improves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000803,"raw_usage":{"total_tokens":3591,"prompt_tokens":1068,"completion_tokens":2523,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":684,"completion_tokens_details":{"reasoning_tokens":2442}},"tokens_in":684,"tokens_out":2523,"duration_ms":16814,"temperature":1.0,"reasoning_tokens":2442,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:54:41.088412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to train the oesophageal-only model on additional oesophageal scans acquired with the same scanner and protocol used for the Indian lung cohort; if the model still scores near zero on the lung scans, cancer type is implicated, whereas a large improvement would show that site and acquisition differences, not cancer type, drove the failure.","supporting_citations":[{"cited_title":"nnu-net: a self-configuring method for deep learning-based biomedical image segmentation","cited_arxiv_id":null,"evidence_quote":"This citation supplies the nnU-Net method used for all three training configurations in the study."},{"cited_title":"A whole-body fdg-pet/ct dataset with manually anno- tated tumor lesions","cited_arxiv_id":null,"evidence_quote":"This citation provides the public whole-body FDG-PET/CT dataset used as the public-only training set."},{"cited_title":"Results from the autopet challenge on fully automated lesion segmentation in oncologic pet/ct imaging","cited_arxiv_id":null,"evidence_quote":"This citation establishes the AutoPET challenge conclusion that data quantity and quality matter more than architecture, the claim this paper extends to cross-cancer evaluation."},{"cited_title":"The autopet challenge: towards fully automated lesion segmentation in oncologic pet/ct imag- ing","cited_arxiv_id":null,"evidence_quote":"This citation defines the AutoPET challenge task and dataset used for the public and combined training arms."},{"cited_title":"nnu-net revisited: A call for rigorous vali- dation in 3d medical image segmentation","cited_arxiv_id":null,"evidence_quote":"This citation supports the claim that a well-configured nnU-Net can match more complex architectures in 3D segmentation."}],"review_version":2}