{"id":"be2add12-7650-42c6-98bb-d17794035388","arxiv_id":"2507.01539","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A multi-center CT phantom dataset with 1378 image series plus baseline harmonization evaluation is made publicly available.","lead":"This paper releases a public CT dataset of a 3D-printed anthropomorphic phantom scanned across 13 scanners, 4 manufacturers, and several dose levels to study how scanner settings change images and features. It also provides baseline metrics for image similarity, feature stability, and liver tissue classification so future harmonization methods can be compared.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark's primary liver-tissue classification task is near-saturated (Table 7: 10-fold accuracy 0.997-1.000) and no harmonization method is evaluated, so the sensitivity of the proposed metrics to harmonization is unverified and the central benchmarking claim is not yet supported.","rationale":"The reader's weakest assumption concerns the phantom's tissue-mimicking limits. I agree this is a limitation, but it is acknowledged in Methods and largely intentional: fixing anatomy is the point of a phantom benchmark, and the paper explicitly restricts claims to controlled scanner and dose shifts. The more load-bearing gap is that the benchmark's only task-level evaluation is saturated and no harmonization method is run. Table 7 shows near-perfect classification even in LOSO with 12 training scanners, and the authors state that the task is 'limited in terms of complexity.' If a harmonization method cannot change a 1.000 accuracy, then the headline evaluation does not measure what the benchmark promises. The image-level and ICC metrics may have headroom, but sensitivity to harmonization is not demonstrated because no harmonization baseline is included. This is a missing validation, not an internal inconsistency, and a simple experiment with ComBat or histogram matching would settle it. The reader's conditional verdict already requires addressing these points, so I recommend keeping the verdict unchanged.","tokens_in":20470,"tokens_out":10272,"duration_ms":117219,"concrete_test":"Apply a simple, established harmonization method (e.g., ComBat on radiomics features, or histogram matching of the CT volumes) to the 10 mGy series and recompute Table 7 (ICC, LOSO accuracy with 1 and 12 training scanners, 10-fold accuracy) and Tables 3-5 (RMSE/PSNR/SSIM). If classification accuracy remains at the ceiling (within one standard deviation of the reported 0.997-1.000 values) while image-level metrics change, the classification component cannot rank harmonization methods. If no proposed metric changes in the expected direction, the evaluation is insensitive and the benchmark's central claim is unsupported. As a secondary check, repeat the LOSO classification on the 1 mGy dose subset, where accuracy might drop; if it is still at ceiling, the task is saturated even under dose shift.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the dataset can serve as a benchmark for developing and comparing AI harmonization methods. For this to hold, the proposed evaluation must be able to detect differences between harmonization methods. This condition is currently unverified and partly contradicted. In the Technical Validation section, the authors report that 'all the models performed almost perfectly' for liver tissue classification; Table 7 gives 10-fold CV accuracies of 0.997±0.001 (radiomics), 1.000±0.000 (shallow CNN), and 0.998±0.002 (SwinUNETR), with LOSO accuracies of 0.920-0.949 for one training scanner and 0.985-0.998 for twelve. A task at or near ceiling cannot show improvements from harmonization: a harmonized model cannot exceed 1.0, so the classification metric has no dynamic range for ranking methods. The image-level metrics (RMSE/PSNR/SSIM, Tables 3-5) and feature-level ICC (Table 7) might be sensitive, but the paper applies no harmonization method at all, so there is no evidence that these metrics respond to harmonization in a meaningful or monotonic way. The paper itself acknowledges the task is too easy and suggests adding other tissue classes, but this is not part of the released benchmark. Thus the load-bearing condition that the benchmark can discriminate harmonization methods is missing support. The phantom's limited attenuation range is a real but explicitly acknowledged limitation that does not by itself invalidate the controlled-shift goal; the metric-sensitivity gap is more central.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Amirian et al. present a publicly available benchmark dataset for evaluating AI harmonization methods in CT imaging. The dataset consists of 1,378 CT image series acquired from a single 3D-printed iodine-ink anthropomorphic phantom on 13 scanners from four manufacturers at eight Swiss institutions, using a harmonized acquisition protocol and five dose levels. The paper describes the phantom, the acquisition protocol, scanner-specific parameters, the data organization, and an open-source code repository. Baseline evaluations without harmonization are reported at three levels: image-level similarity (RMSE, PSNR, SSIM), feature-level stability (ICC), and accuracy of a four-class liver tissue classification task using radiomics and two deep-learning feature extractors under 10-fold and leave-one-scanner-out cross-validation.","tokens_in":20731,"tokens_out":10556,"duration_ms":118931,"significance":"If the benchmark is validated, it will be a valuable community resource: the use of a fixed physical phantom eliminates anatomical and physiological confounds, allowing controlled study of scanner- and dose-induced distribution shifts; the dataset is large (1,378 series), multi-vendor, and public on TCIA; the acquisition protocol is described in detail; and the series counts are internally consistent (Table 2 sums to 1,378). The open-source code and predefined splits lower the barrier for future comparisons. The main weakness is that the evaluation protocol's sensitivity to harmonization is not yet demonstrated.","major_comments":[{"comment":"The central claim of the paper is that the dataset can be used as a benchmark for developing and comparing AI harmonization methods. A benchmark requires that the proposed evaluation metrics have dynamic range to rank methods. This is not demonstrated. In Table 7, the liver tissue classification task is essentially saturated: 10-fold CV accuracy is 0.997±0.001 for radiomics, 1.000±0.000 for shallow CNN, and 0.998±0.002 for SwinUNETR, and LOSO with 12 training scanners reaches 0.985–0.998. A harmonization method cannot improve accuracy beyond 1.0, so this metric cannot discriminate between harmonization methods. The authors acknowledge this ('all the models performed almost perfectly') and propose to add tissue classes, but the released benchmark does not include such a task. To support the central claim, the paper should either add a more challenging task to the released benchmark or include at least one harmonization baseline (e.g., histogram-based image harmonization or ComBat) that demonstrates that the proposed metrics respond in the expected direction.","section":"Technical Validation, Table 7 and surrounding text"},{"comment":"The image-level metrics RMSE, PSNR, and SSIM are proposed as benchmark measures, yet the paper provides no evidence that they are sensitive to harmonization. Table 5 reports SSIM values mostly above 0.95, and the text states that 'global similarity measures do not seem to well capture inter-scanner differences.' Since no harmonization method is evaluated, it is unknown whether these metrics change monotonically or meaningfully when a harmonization method is applied. Please compute the metrics on the liver region of interest (rather than the whole phantom volume) and/or demonstrate with a simple harmonization step that the metrics improve.","section":"Technical Validation, final paragraph and Tables 3–5"}],"minor_comments":[{"comment":"The manuscript contains an unrelated passage beginning '1008 D. Groheux et al. Figure 4. Pre-surgery imaging...' inserted between Figure 1 and Figure 2; this appears to be text from a different publication and must be removed.","section":"After Figure 1"},{"comment":"The PSNR definition uses max(Ir, Is), the maximum of the two specific image series, whereas PSNR is normally defined with the system's dynamic range; the later statement that L=3000 HU should be used in the PSNR definition for consistency.","section":"Methods, Eq. (2)"},{"comment":"The Figure 5 caption says the UMAP embedding was 'optimized over 100 epochs' while the main text and Figure 6 caption say 1000 epochs; please correct the inconsistency.","section":"Figure 5 caption"},{"comment":"The phrase 'which purpose is' should be 'whose purpose is'.","section":"Abstract"},{"comment":"The abstract and usage notes should state explicitly that the benchmark is scoped to liver/soft-tissue harmonization, since the phantom's minimum attenuation is that of paper and does not cover air-containing structures such as lung; the Methods section does mention this, but the broader wording in the abstract may mislead readers.","section":"Methods, Anthropomorphic Phantom and Usage Notes"},{"comment":"Please clarify the unit of analysis for the ICC: are the six ROIs used as targets and the 13 scanners as raters, and are ICCs averaged across ROIs or across features? This affects the interpretation of the standard deviations in Table 7.","section":"Technical Validation, Eq. (4)"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision. The dataset is a valuable resource and the acquisition protocol is carefully documented, but the benchmark's sensitivity to harmonization must be demonstrated before the central benchmarking claim is supported. The extraneous passage after Figure 1 also needs removal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful dataset contribution, and the acquisition protocol is the most careful part of the paper. The 3D-printed iodine-ink phantom, scanned on 13 scanners across 4 manufacturers and 8 institutions at 5 dose levels, gives the field a controlled testbed it did not have. The numbers check out—1378 series from 649 scans, internally consistent—and the data, masks, and code are public on TCIA. That alone makes the paper worth taking seriously. The baselines (RMSE/PSNR/SSIM, ICC, liver tissue classification) are standard, and the UMAP visualizations show clear manufacturer clustering, supporting the claim that the dataset captures scanner-induced shifts.\n\nThe soft spots are real but not fatal. First, Figure 1's caption contains an unrelated block of text from a different paper about FDG PET-CT; that must be fixed before publication. Second, and more substantively, no harmonization method is applied. The paper provides baselines without harmonization, so there is no evidence yet that the proposed metrics respond to harmonization in a meaningful way. The classification task is near-saturated—10-fold accuracies of 0.997-1.000, LOSO around 0.92-0.95 for one training scanner—so that task cannot rank methods. The authors acknowledge this and suggest adding other tissue classes, but it remains a gap. The image-level metrics might be sensitive, but without a single harmonization run, we don't know. Third, the phantom's attenuation range is limited to paper, so no lung or air; the authors acknowledge this, and it does not undermine the controlled-shift goal.\n\nThe citation pattern looks appropriate, and the paper is honest about its limitations, including the note that resampling already acts as a basic harmonization step. It deserves a serious referee and, in my view, conditional acceptance once the caption is fixed and ideally with either a harmonization baseline or a harder classification task to demonstrate metric sensitivity. For a reader working on CT harmonization, this is a potentially valuable resource. I'd engage with it.","headline":"A solid, carefully documented public CT phantom dataset for harmonization benchmarking, with a real caveat: metric sensitivity to harmonization is unverified and one figure caption is contaminated.","tokens_in":21439,"tokens_out":4350,"would_cite":true,"duration_ms":45480,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper builds a public benchmark of 1,378 CT image series of the same 3D-printed phantom scanned on 13 clinical CT scanners, so that AI harmonization methods can be developed and compared where patient variation is held fixed.","keywords":["CT harmonization","anthropomorphic phantom","benchmark dataset","domain shift","radiomics feature stability","tissue classification","multi-centre CT","image harmonization"],"falsifier":"Take a harmonization method that removes per-manufacturer clustering on the phantom and apply it to a small cohort of real patients scanned on two of the same scanners with the same protocol: if the method succeeds on the phantom but leaves patient feature clusters intact, the phantom does not capture the scanner effects the benchmark claims to model. A cheaper check settles the coverage question by inspection: measure the phantom's attenuation range and compare it with the soft-tissue and low-density values, such as lung, that clinical harmonization must handle.","tokens_in":20218,"feed_emoji":"🩻","tokens_out":19510,"duration_ms":177057,"temperature":0.7,"pith_summary":"This paper's claim is that scanner-induced domain shift in CT can be turned into a controlled, measurable benchmark by scanning one fixed object many times. The object is a 3D-printed anthropomorphic phantom made of iodine-inked paper that mimics human liver tissue, and it was scanned on 13 clinical CT scanners from four manufacturers at five radiation-dose levels, producing 649 CT scans and 1,378 reconstructed image series in which anatomy never varies. Because the phantom is identical in every acquisition, all remaining differences between image series must come from the scanner and reconstruction pipeline — exactly the variation that AI harmonization aims to remove. The paper contributes the public dataset, an evaluation protocol (image-similarity metrics, feature stability via the intraclass correlation coefficient, and four-class liver-tissue classification with leave-one-scanner-out cross-validation), and baseline results computed without any harmonization, to serve as reference numbers. If the dataset works as intended, it gives the many competing harmonization techniques a common yardstick for what counts as improvement.","feed_headline":"1,378 phantom image series from 13 scanners test AI harmonization","feed_subtitle":"Same phantom, fixed anatomy: scanner-driven variation becomes the benchmark's only variable.","key_machinery":"The central object is the anthropomorphic phantom: a 3D print of a real human CT scan in which iodine ink injected into paper raises its attenuation to match liver tissue, accompanied by a thoracic segment and synthetic test patterns, and carrying six annotated liver ROIs from four tissue classes (two cysts, a hemangioma, a metastasis, and two normal regions). Around the phantom sits the harmonized acquisition protocol — acquisition and reconstruction parameters averaged from a survey of clinical thoracoabdominal CT protocols — which fixes tube voltage, pitch, rotation time, collimation, dose, and reconstruction settings as closely as vendor limits allow, so that the scanner becomes the main free variable. The evaluation machinery has three parts: image-level similarity via root mean square error, peak signal-to-noise ratio, and structural similarity between rigidly registered scans; feature-level stability via the intraclass correlation coefficient (ICC(3,1)), which measures feature variation across scanners relative to variation within a scanner; and four-class liver tissue classification on fixed image patches under two protocols — leave-one-scanner-out cross-validation, which tests generalization to scanners never seen in training, and 10-fold cross-validation, where all scanners appear in the training set.","core_discovery":"The paper establishes a test-retest benchmark: the same physical phantom, scanned repeatedly under controlled settings, so that scanner-related variation is separated from all patient-related variation. The authors show that the benchmark captures real domain shift by demonstrating that features from the phantom's six liver ROIs — handcrafted radiomics, a shallow CNN trained on organ recognition, and a transformer pre-trained on 3D CT volumes — cluster by scanner manufacturer in low-dimensional projections, with reconstruction technique adding a second source of spread. Their baselines indicate the shift exists but does not yet break the provided task: liver-tissue classification accuracy is high even when the test scanner was absent from training, and global structural similarity across scanners is high. The paper reads these observations as showing that the four-class classification task is too easy to expose the benefits of harmonization, and that harmonization effects should be evaluated on target tissues rather than on whole volumes. The contribution is the dataset, the evaluation methodology, and the reference numbers — not a new harmonization algorithm.","pith_inferences":["The most informative outcome for this benchmark would be a negative one: if a harmonization method that removes manufacturer clustering on the phantom fails to do so on real multi-centre patient data, the failure would pinpoint which scanner effects the phantom cannot reproduce, most likely in low-density structures.","A natural extension is to convert the phantom's thoracic segment and synthetic test patterns into additional classification targets, since the four-class liver task is nearly saturated and cannot distinguish between competing harmonization methods.","Because the authors find high global SSIM but strong per-manufacturer feature clustering, harmonization evaluation on this dataset is best restricted to the liver ROIs, where the scanner effects actually appear, rather than to whole volumes.","The phantom's paper-density floor limits its attenuation range, so the benchmark says nothing about scanner effects on lung and other low-density tissue; a direct probe would be to insert a low-density calibration insert of known attenuation and re-scan a subset of the 13 scanners."],"forward_implications":["Harmonization methods can be benchmarked against the published baselines: improvement means better image-level similarity, higher feature ICC, and higher leave-one-scanner-out classification accuracy.","Because anatomy, physiology, and disease are fixed across all acquisitions, a reduction in cross-scanner differences achieved on the phantom is attributable to genuine scanner effects rather than to patient variation.","The five dose levels and three reconstruction families (filtered backprojection, iterative reconstruction, and deep-learning reconstruction) allow harmonization across dose and reconstruction to be assessed alongside manufacturer effects.","The near-saturated baseline classification implies that the four-class liver task alone will not separate good from bad harmonization methods; the authors suggest adding further tissue classes from the phantom, such as organs or bone.","The authors caution that the dataset is not recommended for developing segmentation models, despite providing masks, because of the phantom's missing anatomical diversity and the task's simplicity."],"supporting_citations":[{"why":"Defines and evaluates the 3D-printed iodine-ink paper phantom that serves as the dataset's fixed imaging object.","marker":"[6]"},{"why":"Supplies the radiopaque 3D-printing technique used to manufacture the anthropomorphic phantom from a real patient CT scan.","marker":"[16]"},{"why":"Provides the phantom's six liver ROIs with four tissue classes and the earlier task-based stability analysis this dataset builds on.","marker":"[17]"},{"why":"Source of the shallow CNN pre-trained on CT organ classification whose latent features form the first deep baseline.","marker":"[18]"},{"why":"Defines the handcrafted radiomics feature extraction used for the radiomics baseline.","marker":"[35]"},{"why":"Provides the self-supervised SwinUNETR whose bottleneck features form the second deep baseline.","marker":"[32]"},{"why":"Defines the ICC(3,1) formula used to score feature stability across scanners.","marker":"[30]"},{"why":"Defines the structural similarity (SSIM) metric used in the image-level similarity baselines.","marker":"[38]"},{"why":"Frames the harmonization problem and the scarcity of benchmark datasets that this work addresses.","marker":"[24]"}],"fun_headline_variants":["Phantom CT benchmark: 1,378 series, 13 scanners to test AI harmonization","Same phantom, many scanners: new dataset isolates CT harmonization targets","Benchmark shows scanner shift but liver task too easy for harmonization","1,378 phantom scans, 13 scanners: new AI harmonization benchmark dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise, stated by the authors in the Methods section, is that the iodine-ink paper phantom reproduces the scanner effects that matter in real patients: its attenuation cannot go below that of paper, so low-density structures such as lung are absent, and it has no anatomical or pathological variability, so harmonization methods tuned on the phantom could still fail on patient data.","fun_headline_variants_meta":{"raw":{"variants":["Phantom CT benchmark: 1,378 series, 13 scanners to test AI harmonization","Same phantom, many scanners: new dataset isolates CT harmonization targets","Benchmark shows scanner shift but liver task too easy for harmonization","1,378 phantom scans, 13 scanners: new AI harmonization benchmark dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000287,"raw_usage":{"total_tokens":1673,"prompt_tokens":920,"completion_tokens":753,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":669}},"tokens_in":536,"tokens_out":753,"duration_ms":7885,"temperature":1.0,"reasoning_tokens":669,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:48:08.140473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a harmonization method that removes per-manufacturer clustering on the phantom and apply it to a small cohort of real patients scanned on two of the same scanners with the same protocol: if the method succeeds on the phantom but leaves patient feature clusters intact, the phantom does not capture the scanner effects the benchmark claims to model. A cheaper check settles the coverage question by inspection: measure the phantom's attenuation range and compare it with the soft-tissue and low-density values, such as lung, that clinical harmonization must handle.","supporting_citations":[{"cited_title":"3D-printed iodine-ink CT phan- tom for radiomics feature extraction-advantages and challenges.Medical Physics, 50(9):5682–5697, 2023","cited_arxiv_id":null,"evidence_quote":"Defines and evaluates the 3D-printed iodine-ink paper phantom that serves as the dataset's fixed imaging object."},{"cited_title":"Radiopaque three- dimensional printing: a method to create realistic CT phantoms.Radiology, 282(2):569–575, 2017","cited_arxiv_id":null,"evidence_quote":"Supplies the radiopaque 3D-printing technique used to manufacture the anthropomorphic phantom from a real patient CT scan."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the phantom's six liver ROIs with four tissue classes and the earlier task-based stability analysis this dataset builds on."},{"cited_title":"Obmann, André Anjos, Henning Müller, and Adrien Depeursinge","cited_arxiv_id":null,"evidence_quote":"Source of the shallow CNN pre-trained on CT organ classification whose latent features form the first deep baseline."},{"cited_title":"Compu- tational radiomics system to decode the radiographic phenotype.Cancer Research, 77:e104–e107, 2017","cited_arxiv_id":null,"evidence_quote":"Defines the handcrafted radiomics feature extraction used for the radiomics baseline."},{"cited_title":"Roth, Bennett Land- man, Daguang Xu, Vishwesh Nath, and Ali Hatamizadeh","cited_arxiv_id":null,"evidence_quote":"Provides the self-supervised SwinUNETR whose bottleneck features form the second deep baseline."},{"cited_title":"Intraclass correlations: uses in assessing rater reliability.Psychological bulletin, 86(2):420, 1979","cited_arxiv_id":null,"evidence_quote":"Defines the ICC(3,1) formula used to score feature stability across scanners."},{"cited_title":"Making radiomics more reproducible across scanner and imaging protocol variations: a review of harmonization methods","cited_arxiv_id":null,"evidence_quote":"Frames the harmonization problem and the scarcity of benchmark datasets that this work addresses."}],"review_version":1}