{"id":"bc5282aa-7597-4595-a8e5-fbbdfc43390c","arxiv_id":"2507.22017","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A multi-center MRI benchmark and federated learning framework for IPMN malignancy-risk stratification, with an internal AUC of 0.85 on T2-weighted MRI.","lead":"Cyst-X is a new public dataset of 1,461 MRI scans from 764 patients across seven international medical centers, plus a deep-learning pipeline that sorts pancreatic cysts into high-risk and low/no-risk groups. If the reported performance holds, it gives researchers a common benchmark to build and test AI tools that could reduce unnecessary pancreas surgery and catch early cancer.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline AUC compares high-risk IPMN against a negative class that includes ~30% patients with no pancreatic cyst; the clinically relevant high-risk-vs-low-risk IPMN discrimination is never reported, so the 'malignancy-risk stratification' claim may be inflated by cyst detection.","rationale":"The reader's label-stability concern is legitimate, but it concerns mislabeling within the negative class and would most likely bias against the model; it does not explain an inflated AUC. The inclusion of no-cyst controls, by contrast, can only make the classification task easier and is directly testable from the released data. The paper deserves credit for the public dataset, the per-center external evaluation, and the honest federated-learning setup; however, the central 'malignancy-risk stratification' claim depends on a subgroup comparison that is not reported. A conditional acceptance requiring the high-risk-versus-low-risk analysis (with and without FedProx) is the appropriate outcome. The 87.8% sensitivity cited in the Discussion also does not match Table 3's 56.83% internal sensitivity, which further weakens the radiologist-comparison wording and should be corrected.","tokens_in":42807,"tokens_out":8793,"duration_ms":109874,"concrete_test":"Recompute the binary DenseNet-121 T2W result excluding all no-risk controls (normal imaging or benign non-IPMN cysts), leaving only low-risk IPMN (N≈344) versus high-risk IPMN (N≈152), using the same per-center stratified four-fold cross-validation and the same threshold-calibration policy; also recompute FedProx(µ=0.1) on the same restricted cohort. If the high-risk-versus-low-risk AUC drops more than ~0.05 from the reported 0.85, or the FedProx gap widens beyond the reported ~0.7 points, the headline stratification claim is overstated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.1 defines the 'No risk / Control' category as 'individuals with normal pancreatic imaging, no cystic lesions or benign cysts not associated with IPMN,' and Table A1 shows that on T2W the no-risk group contributes 159 of the 503 no/low-risk negatives in the binary analysis. Roughly 30% of the negative class therefore consists of patients without any IPMN, so a classifier can achieve a high AUC by detecting the presence of a cyst rather than by grading malignant potential. The paper's central numeric claims (T2W AUC 0.85, FedProx 0.8458) and the radiologist comparison are all computed on this mixed negative class. No analysis isolates high-risk IPMN from low-risk IPMN among patients with known IPMN, which is the decision that 'malignancy-risk stratification' actually refers to. If performance on that subgroup is materially lower, the abstract's first two claims are not supported by the reported experiments.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Cyst-X assembles a multi-center MRI cohort of 1,461 abdominal scans from 764 patients at seven institutions, with expert pancreas segmentations and three-tier IPMN risk labels anchored in histopathology or three-year imaging follow-up. The paper benchmarks a PanSegNet-based segmentation pipeline and a 3D DenseNet-121 classifier, reports internal cross-validated T2-weighted AUC of 85.28%, a leave-one-center-out external AUC of 81.43%, and FedProx internal AUC of 84.58%, and compares the classifier against three blinded radiologists on a 629-case subset. The authors frame the work as a public benchmark, a federated-learning capability demonstration, and a step toward MRI-based IPMN malignancy-risk stratification. The dataset release and large-scale multi-center scope are genuine strengths, but several load-bearing empirical claims depend on the composition of the negative class, the threshold calibration for the reader comparison, and the choice of internal rather than external evaluation for the federated claim.","tokens_in":42998,"tokens_out":5699,"duration_ms":64430,"significance":"If the central claims hold, Cyst-X would be a substantial community resource: it is the largest public multi-center MRI dataset for IPMN risk stratification, with per-center external evaluation, MRQy-based heterogeneity analysis, and a reproducible benchmark including radiomics, multiple 3D CNN baselines, and federated training. The segmentation results (PanSegNet Dice 89.62% on T2W) and the reported internal-to-external AUC drop are honestly quantified, and the observation that federation hurts dense segmentation more than whole-region classification is an interesting empirical finding. However, the headline 'malignancy-risk stratification' claim is not yet supported by the reported experiments, because the binary task is high-risk versus a negative class that includes roughly one-third patients without any pancreatic cyst, and the clinically relevant high-risk-versus-low-risk IPMN discrimination is never isolated. The radiologist comparison also depends on per-center thresholds tuned on the training partition, which makes the 'matched conditions' claim fragile.","major_comments":[{"comment":"The headline binary task is high-risk versus no/low-risk, but the negative class pools 344 low-risk IPMNs with 159 T2W 'No Risk/Control' patients who have no pancreatic cyst, roughly 31.6% of the 503 negatives in Table A1. A classifier can therefore achieve a high AUC by detecting the presence of a cyst rather than by grading malignant potential. The paper never reports the clinically relevant high-risk versus low-risk IPMN-only discrimination, so the 'malignancy-risk stratification' claim is not directly supported by the reported experiments. Please add a subgroup analysis restricted to patients with a known IPMN (low-risk versus high-risk), ideally per center and under the same four-fold protocol, and report whether the T2W AUC of 0.85 is preserved in that subgroup. If the dataset cannot support such an analysis, the title and abstract should be reworded to describe the task as high-risk versus no/low-risk discrimination without claiming IPMN malignancy-risk stratification as the demonstrated endpoint.","section":"§4.1, Table A1"},{"comment":"The radiologist comparison is not matched in the sense claimed. The classifier's 'clinically optimized' sensitivity/specificity in Table 3 are obtained from per-center thresholds selected in Section A.6 by maximizing accuracy subject to sensitivity>35% and specificity>85% on the training partition, whereas the three radiologists operate at naturally chosen, reader-specific operating points. At the standard 50% probability threshold, the T2W DenseNet-121 has sensitivity 38.85% at specificity 96.84%, which is below the mean reader sensitivity of 46.01% at specificity 93.91%. The statement that the classifier 'matched or exceeded sensitivity at comparable specificity' therefore depends on an operating point tuned per center for accuracy, not matched to reader specificity. Please report a fixed global threshold or, preferably, compute sensitivity at each reader's specificity (or specificity at each reader's sensitivity) with confidence intervals for the sensitivity difference, and clearly state that the optimized-threshold numbers are not a head-to-head matched comparison.","section":"§A.6, Table 3"},{"comment":"The abstract and Section 2.3 state that FedProx preserves the T2W discrimination of the centralized model within 0.7 AUC points, but this result comes from the internal four-fold cross-validation only. Under the leave-one-center-out external protocol in Table A10, the global T2W AUC for centralized DenseNet-121 is 81.43%, but it drops to 74.71% for FedAvg and 69.70% for FedProx(mu=0.1), an 11.7-point gap relative to the centralized external model. The paper's claim that federation 'preserves discrimination' is therefore not supported by the external evaluation, and the large external federated degradation is not discussed in Sections 2.3 or 2.4. Please report and discuss the external federated results, qualify the abstract's federation claim to internal cross-validation, or modify the federation protocol so that the external claim is actually tested.","section":"§2.3, §2.4, Table A10"},{"comment":"The low-risk label includes presumed-IPMN lesions without histology that remained stable on at least three years of imaging follow-up (growth under 2.5 mm and no worrisome features). This assumption is load-bearing because the negative class in the binary analysis is dominated by low-risk cases (344 of 503 on T2W, Table A1). A meaningful fraction of such 'stable' lesions could later progress, which would bias the high-risk versus no/low-risk AUC and the radiologist comparison. Please add a sensitivity analysis using only histologically confirmed low-risk IPMNs (versus pathologically confirmed high-risk) to quantify the effect of follow-up-based labels, and report the fraction of low-risk labels assigned by follow-up stability versus histology within the binary cohort.","section":"§4.1"}],"minor_comments":[{"comment":"The Discussion states that the model has 'higher sensitivity for detecting malignant IPMNs (87.8% vs. 64.1%) while preserving specificity,' but 87.8% does not appear in Table 3 for any reported operating point (the T2W DenseNet values are 56.83% and 38.85% at the two thresholds). Please either cite the exact threshold and table that yields 87.8% or correct the sentence.","section":"§3, Table 3"},{"comment":"The text says 'center AHN contains only a single case,' but Table A1 lists 16 to 18 cases for AHN depending on modality. The intended statement appears to be that AHN contains only a single no-risk case; please correct the wording.","section":"§4.4, Table A1"},{"comment":"The Calinski-Harabasz index is cited to reference [32], which is the Davies-Bouldin paper; please use the original Calinski-Harabasz reference.","section":"Figure 3 caption"},{"comment":"There are minor typographical errors, including 'mdoels' in Section 4.4 and 'Proceeedings' in reference [50]; a light proofreading pass would improve the manuscript.","section":"§4.4, Reference [50]"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a serious dataset-and-benchmark contribution with public code and data, and the internal versus external AUC reporting is unusually transparent. My main concern is framing: the negative-class composition, the per-center threshold calibration, and the unexamined external federated results each touch a central claim. I do not see evidence of any attempt to mislead; all three issues are addressable with additional analyses or appropriately qualified claims. If the authors can show that the high-risk versus low-risk IPMN-only AUC remains materially above chance and that the radiologist comparison holds at a fixed operating point, the paper would be a strong candidate for acceptance. Otherwise, the title and abstract should be revised to reflect what the experiments actually establish."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Cyst-X is a genuinely useful dataset: 1,461 MRI scans, 764 patients, seven centers, with segmentation masks and three-tier labels. That is the real contribution, and if the release works as advertised it fills a structural gap. The paper also reports a clean empirical asymmetry—FedProx preserves within 0.7 AUC points for whole-pancreas classification but costs 7–8 Dice points for segmentation—which is worth publishing on its own.\n\nThe soft spots are real, and one is load-bearing. The headline AUC 0.85 is for high-risk versus no/low-risk, where 'no' means no cyst at all (Table A1: 159 of 503 negatives on T2W). That means the model can score points by detecting the presence of a cyst, not by grading malignant potential. The clinically relevant comparison—high-risk IPMN versus low-risk IPMN among patients who actually have an IPMN—is never reported. The three-class AUC (82.55) is not a substitute; it pools no-cyst with low-risk as the 'low' category. Until the authors report pairwise high-vs-low, the abstract's \"malignancy-risk stratification\" claim is not supported.\n\nSecond, the FedProx preservation claim is internal-CV only. In the leave-one-center-out table (A10), centralized DenseNet drops to 81.43 on T2W, but FedProx drops to 69.70. The abstract does not mention this. If federation is being sold as a deployment-ready answer, that gap matters.\n\nThird, the radiologist comparison uses per-center thresholds tuned on the training partition, and reader performance is reported without confidence intervals. The paper's own text is honest that the model does not beat the most sensitive reader (64.08 vs 56.83), yet the abstract says 'matched or exceeded.' That is an overstatement.\n\nThe low-risk labels from three-year imaging stability are a plausible but noisy proxy; the paper acknowledges it, so I treat it as a minor caveat, not a flaw.\n\nBottom line: this is a dataset paper and should be reviewed as one. The benchmark, the federation asymmetry, and the honest external-evaluation tables are all worth refereeing. But the authors need to fix the negative-class analysis, tone down the abstract, and report FedProx's external numbers before I'd trust the risk-stratification claims. Send it to peer review, with a request for major revision.","headline":"Valuable dataset, but the headline risk-stratification AUC mixes no-cyst controls into the negative class; the high-vs-low IPMN comparison is missing.","tokens_in":43712,"tokens_out":3226,"would_cite":true,"duration_ms":37251,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"With 1,461 MRI scans from 764 patients across seven centers, a 3D DenseNet-121 separates high-risk pancreatic cystic neoplasms at a mean AUC of 0.85 on T2-weighted MRI, and federated training preserves that discrimination without sharing…","keywords":["IPMN","pancreatic cystic neoplasm","malignancy risk stratification","MRI","federated learning","deep learning","pancreas segmentation","benchmark dataset"],"falsifier":"Track the presumed low-risk IPMN patients beyond three years, for example five to ten years of imaging or surgical pathology, and count how many previously stable lesions develop high-grade dysplasia or invasive carcinoma; if a meaningful fraction progress, the high-risk versus no/low-risk discrimination reported here is biased. A complementary check is a temporally held-out test partition using all scans from 2022 onward, or a pseudo-prospective study where the locked classifier makes predictions before radiologist scores are collected.","tokens_in":42614,"feed_emoji":"🩻","tokens_out":7190,"duration_ms":76809,"temperature":0.7,"pith_summary":"Pancreatic cystic neoplasms, especially IPMNs, present a clinical dilemma: most are indolent but some progress to cancer, and current guideline features both over-treat and under-detect. This paper claims that a large, multi-center MRI benchmark, 1,461 scans from 764 patients at seven centers, lets a 3D DenseNet-121 classifier separate high-risk from no/low-risk IPMNs with a mean AUC of 0.85 on T2-weighted MRI, raising average precision from a 0.23 prevalence baseline to 0.64. It further claims that federated training with FedProx preserves this discrimination within 0.7 AUC points without exchanging raw images, and that on a 629-case reader subset the classifier matches or exceeds the radiologist average sensitivity at comparable specificity. If these claims hold, imaging-based risk stratification could reduce both unnecessary surgeries and missed high-risk lesions, and federated training would let institutions collaborate without centralizing patient data.","feed_headline":"MRI model hits 0.85 AUC for high-risk pancreatic cysts","feed_subtitle":"A 764-patient, seven-center benchmark also beats the average blinded radiologist's sensitivity at matched specificity.","key_machinery":"The central object is the Cyst-X benchmark itself: 1,461 MRI scans (723 T1-weighted, 738 T2-weighted) from 764 patients at seven centers, labeled high-risk, low-risk, or no-risk by histopathology or at least three years of imaging follow-up, with expert pancreas segmentation masks. The analytical mechanism is a pipeline that couples PanSegNet, an nnU-Net-based pancreas segmenter with linear self-attention, to a 3D DenseNet-121 classifier on the segmented pancreas region, with a parallel radiomics pipeline, 1,409 hand-crafted features with mRMR selection and a random forest, as a classical baseline. For distributed training, FedProx adds a proximal term to each client's loss to counter client drift, and the paper compares FedAvg and FedProx across seven institutional silos. The key comparison that carries the argument is the near-parity of centralized and federated classification AUC, 85.28% versus 84.58% on T2-weighted MRI, alongside a large federated segmentation penalty, which the paper reads as evidence that dense per-voxel prediction is more sensitive to inter-site heterogeneity than whole-region classification.","core_discovery":"On its own terms, the paper establishes that IPMN malignancy-risk stratification from MRI is feasible at multi-center scale. A 3D DenseNet-121, trained on pancreas regions segmented by PanSegNet and evaluated on T2-weighted MRI, discriminates high-risk from no/low-risk IPMNs with mean AUC 85.28% (95% CI 84.48 to 86.08%), average precision 0.64 versus a 0.23 prevalence baseline, and sensitivity 56.83% at specificity 93.05% under a clinically calibrated threshold, while the three blinded radiologists scored 46.01% mean sensitivity at 93.91% mean specificity on the same 629-case imaging-only subset. Federated training with FedProx ($\\mu=0.1$) reaches AUC 84.58% on T2-weighted MRI, 0.7 points below the centralized baseline, whereas federated pancreas segmentation with Swin-UNETR loses 7.10 to 7.83 Dice points. The paper identifies the segmentation-versus-classification asymmetry under federation as its key empirical finding, and it releases the dataset, masks, and models publicly.","pith_inferences":["A testable extension would be to re-label the no/low-risk group with longer follow-up, for example five to ten years, and check whether presumed-IPMN lesions that remained stable for three years later progress; if they do, the reported AUC is optimistic for true malignancy risk.","Because the paper combines imaging with neither cyst fluid genomics nor clinical variables, a natural next step is a multimodal risk profile, and the imaging-only framing suggests the classifier's contribution is complementary rather than a stand-alone surgical trigger.","The four-point internal-to-external AUC drop and center-level variance predict that deploying the locked classifier at a new site with a different scanner mix will require site-specific calibration or federated fine-tuning, and that small centers will dominate variance.","If FedProx's proximal term regularizes a noisy optimization landscape, then the T1-weighted federated AUC exceeding the centralized baseline, 81.20% versus 78.60%, is likely a regularization artifact rather than a signal that federation improves discrimination; a direct test would compare FedProx against centralized training with an equivalent proximal penalty."],"forward_implications":["A T2-weighted MRI-based classifier can be trained to flag high-risk IPMNs with sensitivity above the average blinded radiologist at comparable specificity, offering a decision-support tool for imaging-only assessment.","Federated training with FedProx allows multiple institutions to build an IPMN risk classifier without pooling raw images, with an AUC loss of about one point relative to centralized training on T2-weighted MRI.","The segmentation-versus-classification asymmetry implies that pancreas segmentation models should be trained centrally and distributed, while classification heads can be federated across sites.","Leave-one-center-out evaluation gives an external AUC of 81.43% on T2-weighted MRI, about four points below internal cross-validation, indicating that site heterogeneity is a real cost that fine-tuning or federation against new participants would need to address.","Public release of the dataset, segmentation masks, and trained models gives subsequent work fixed splits and baselines for MRI-based IPMN risk stratification."],"supporting_citations":[{"why":"Provides PanSegNet, the segmentation model that defines the pancreas ROI and establishes the upper bound for downstream classification.","marker":"[17]"},{"why":"Supplies FedProx, the federated optimizer whose proximal term preserves classification AUC within 0.7 points of centralized training.","marker":"[37]"},{"why":"Defines the Kyoto criteria used both for low-risk label assignment, stability and no worrisome features, and for the radiologist reader scoring benchmark.","marker":"[4]"},{"why":"Is the nnU-Net backbone underlying PanSegNet, and its centralized dataset fingerprinting is why PanSegNet cannot be federated.","marker":"[30]"},{"why":"Is the DenseNet-121 architecture used as the principal deep-learning classifier.","marker":"[57]"},{"why":"Swin-UNETR is the federatable transformer segmentation baseline whose seven-to-eight-point Dice penalty contrasts with classification parity.","marker":"[31]"},{"why":"FedAvg is the baseline federated averaging algorithm against which FedProx is compared.","marker":"[36]"},{"why":"MRQy supplies the 21 image-quality indicators used to characterize inter-center heterogeneity and slice-thickness clustering.","marker":"[49]"}],"fun_headline_variants":["Federated MRI model keeps 0.85 AUC for cyst risk","AI beats radiologists on pancreatic cyst sensitivity","Seven-center MRI dataset fuels cyst risk AI to 0.85","Pancreatic cyst AI: 0.85 AUC, tops radiologists"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that lesions judged to be low-risk IPMNs without a tissue diagnosis, which stayed stable on imaging for at least three years, really are low-risk; if a meaningful share of those stable lesions would later turn out to be high-risk, the reported separation between high-risk and no/low-risk groups is optimistic.","fun_headline_variants_meta":{"raw":{"variants":["Federated MRI model keeps 0.85 AUC for cyst risk","AI beats radiologists on pancreatic cyst sensitivity","Seven-center MRI dataset fuels cyst risk AI to 0.85","Pancreatic cyst AI: 0.85 AUC, tops radiologists"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1513,"prompt_tokens":1110,"completion_tokens":403,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":726,"completion_tokens_details":{"reasoning_tokens":330}},"tokens_in":726,"tokens_out":403,"duration_ms":4546,"temperature":1.0,"reasoning_tokens":330,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:06:25.567213+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Track the presumed low-risk IPMN patients beyond three years, for example five to ten years of imaging or surgical pathology, and count how many previously stable lesions develop high-grade dysplasia or invasive carcinoma; if a meaningful fraction progress, the high-risk versus no/low-risk discrimination reported here is biased. A complementary check is a temporally held-out test partition using all scans from 2022 onward, or a pseudo-prospective study where the locked classifier makes predictions before radiologist scores are collected.","supporting_citations":[{"cited_title":"Medical image analysis99, 103382 (2025)","cited_arxiv_id":null,"evidence_quote":"Provides PanSegNet, the segmentation model that defines the pancreas ROI and establishes the upper bound for downstream classification."},{"cited_title":"Pancreatology24(2), 255–270 (2024)","cited_arxiv_id":null,"evidence_quote":"Defines the Kyoto criteria used both for low-risk label assignment, stability and no worrisome features, and for the radiologist reader scoring benchmark."},{"cited_title":"In: Proceedings of the International MICCAI Brainlesion Workshop, pp","cited_arxiv_id":null,"evidence_quote":"Swin-UNETR is the federatable transformer segmentation baseline whose seven-to-eight-point Dice penalty contrasts with classification parity."},{"cited_title":"Medical physics47(12), 6029–6038 (2020)","cited_arxiv_id":null,"evidence_quote":"MRQy supplies the 21 image-quality indicators used to characterize inter-center heterogeneity and slice-thickness clustering."}],"review_version":1}