{"id":"b2cfd13e-164f-49c1-921c-ea521ac46432","arxiv_id":"2608.06158","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In realistic calcified coronary phantoms, deep silicon photon-counting CT reduced percent area stenosis and vessel area measurement errors relative to conventional energy-integrating CT, using micro-CT as ground truth.","lead":"This study scanned realistic coronary artery phantoms with three CT systems and used micro-CT as a reference to compare how accurately each system measured vessel narrowing from calcium deposits. The silicon-based photon-counting CT produced smaller measurement errors than the conventional detector, especially for heavily calcified vessels.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Segmentation threshold sensitivity is the key unvalidated link: the reported dSi-PCCT advantage could be an artifact of modality-dependent multi-Otsu boundaries; a threshold-perturbation test is needed.","rationale":"The reader's conditional verdict is well-placed. The central claim is a difference in MAE between two imaging systems, and that difference is produced by an automated segmentation whose thresholds are not shown to be modality-invariant. The authors' own third limitation concedes this exact point, so the concern is not manufactured; it is the weakest link in the evidence chain. A threshold-perturbation experiment is the natural, low-cost way to test it. I also noted the Eq.2/Table 2 inconsistency as a secondary but real reproducibility issue; it does not by itself overturn the claim because the tabulated values suggest the intended metric is calcification-area fraction, but it makes the paper's specification incomplete. The paper's strengths—an external Micro-CT reference, a secondary ellipse-derived analysis that agrees, and candid discussion of static, resolution-optimized conditions—support keeping the verdict CONDITIONAL rather than moving to REJECT or UNVERDICTED. The requested sensitivity analysis and metric clarification are exactly what conditional acceptance should require.","tokens_in":12161,"tokens_out":11364,"duration_ms":121286,"concrete_test":"Rerun the whole primary analysis on the existing registered volumes with the multi-Otsu thresholds perturbed by ±20 HU and ±50 HU (or equivalently by ±5% and ±10% of the inter-threshold interval) for each of the 12 sections, and recompute the whole-profile MAE for Micro-CT-referenced percent area stenosis and vessel area. If the dSi-PCCT advantage (1.62% vs 3.10%, 0.31 vs 0.55 mm²) disappears or reverses under any perturbation, the reported improvement is a segmentation artifact. In parallel, verify the numerator in Eq.2 by checking whether Table 2's Micro-CT entries equal A_calc/A_ellipse rather than A_lumen/A_ellipse; if the printed equation is the one actually used, the direction of stenosis errors is opposite to what is reported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the three-class multi-Otsu segmentation (Section 2.3.1) produces modality-invariant lumen and calcification boundaries after rigid registration. All reported MAE values are computed from these masks, so any modality-dependent behavior of the thresholding would directly change the central comparison. The two systems differ not only in resolution but also in HU response (10 mg/mL iodine reads 276 HU on EID-CT vs 337 HU on dSi-PCCT, Section 4) and in noise, reconstruction kernel, and voxel size; the Otsu thresholds are estimated independently per volume, so the physical boundary selected for the lumen/calcification transition can differ between modalities. In EID-CT, a larger partial-volume fraction at the calcification edge makes the threshold choice especially consequential: small shifts can systematically inflate or deflate the calcified area, changing percent-area-stenosis error. The paper itself acknowledges this risk in Section 4's third limitation, stating the pipeline 'does not reproduce the clinical workflow' and that reader studies are needed to separate true boundary fidelity from segmentation behavior. An additional internal inconsistency compounds this: Eq.2 defines the primary stenosis metric as 100*A_lumen/A_vessel, yet Table 2's Micro-CT values increase with calcification extent (Type I 26.96% vs Type IV 41.37% at 10 mg/mL), which matches A_calc/A_vessel instead. If the wrong numerator was used, all reported MAE numbers would be ambiguous. The central claim therefore rests on an unvalidated, potentially threshold-sensitive segmentation and an incompletely specified metric.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares a conventional energy-integrating detector CT (EID-CT) scanner with a deep silicon photon-counting CT (dSi-PCCT) system for coronary stenosis quantification in realistic calcified coronary artery phantoms, using Micro-CT as the ground-truth reference. Twelve vessel sections spanning four calcification geometries and three iodine concentrations were scanned with all three systems, rigidly registered to Micro-CT, and segmented with an automated multi-Otsu pipeline. The primary analysis reports lower whole-profile mean absolute error in Micro-CT-referenced percent area stenosis for dSi-PCCT (1.62% vs 3.10%, p=0.027) and lower segmented vessel-area error (0.31 vs 0.55 mm², p=0.001); the secondary ellipse-derived analysis is concordant. The authors conclude that the high-resolution dSi-PCCT protocol improves task-based coronary stenosis quantification under static, resolution-optimized conditions.","tokens_in":12443,"tokens_out":9129,"duration_ms":95182,"significance":"If the results hold, this is a valuable controlled comparison for a newly cleared silicon-based photon-counting CT system. The study uses an external Micro-CT reference, matched kVp and tube current (dose), explicit pre-registered analysis, and a realistic phantom set with clinical calcification morphologies. The authors are transparent about limitations, especially the non-clinical automated segmentation and static imaging conditions. The main value is as a quantitative task-based baseline for future dynamic and clinical studies. However, the central estimate depends on the modality invariance of the multi-Otsu segmentation, which is not independently validated, and on statistical assumptions that may overstate significance because the 12 profiles come from only six physical phantoms.","major_comments":[{"comment":"The automated three-class multi-Otsu thresholding (Section 2.3.1) is the sole source of the lumen and calcification masks from which every reported metric is computed. The thresholds are estimated independently per volume, and the two modalities differ in HU response (276 vs 337 HU for 10 mg/mL iodine, Section 4), noise, voxel size, and reconstruction kernel. A systematic shift in the selected lumen/calcification boundary between EID-CT and dSi-PCCT would directly change the primary and secondary outcomes. The third limitation in Section 4 concedes that the pipeline does not reproduce the clinical workflow, but the manuscript provides no quantitative test of segmentation robustness. A threshold-perturbation sensitivity analysis, an independent reader-based validation on a subset, or a comparison against the known physical dimensions of the phantom would be needed to establish that the reported improvement reflects boundary fidelity rather than thresholding behavior. Because every MAE value depends on these masks, this is load-bearing for the central claim.","section":"2.3.1 / 4"},{"comment":"The statistical comparison treats the 12 vessel-section profiles as independent matched blocks, but these profiles are derived from only six physical phantoms (three vessel geometries × two iodine concentrations), with multiple calcification regions per phantom. Profiles from the same physical phantom share manufacturing tolerances, registration transformations, and noise realizations, so the effective sample size is smaller than 12. The paired block permutation test in Section 2.3.3 does not account for this clustering. For the primary percent-area-stenosis MAE, the reported p=0.027 may not survive a phantom-level clustered analysis. Please report per-phantom or per-geometry MAE values and provide a clustered permutation test or a mixed-effects analysis, and consider the two primary endpoints together when interpreting significance.","section":"2.3.3 / 8"}],"minor_comments":[{"comment":"In Table I, the dSi-PCCT and Micro-CT rows omit the kVp value (120 kVp is given in the text). The table should list all acquisition parameters explicitly and consistently for each system.","section":"2.2 / Table I"},{"comment":"The two-step multi-Otsu segmentation strategy 'for vessel sections containing the highest iodine concentration' is mentioned but not described; please specify how the second step is initialized and why this adaptation is needed only at the highest concentration.","section":"2.3.1"},{"comment":"The conclusion attributes the improvement to the detector's higher native resolution, but the comparison also involves different reconstruction algorithms (DL-High UHD Ultra vs ASiR-V 50% Bone) and voxel sizes. The manuscript mostly phrases this as a system-level protocol comparison, but the abstract and conclusion should consistently avoid implying a pure detector-physics attribution.","section":"4"},{"comment":"With 12 paired profiles, there are only 2^12 = 4096 distinct sign permutations; if the 10,000-permutation procedure uses resampling rather than exact enumeration, the discreteness of the permutation distribution should be acknowledged when reporting p-values.","section":"2.3.3"},{"comment":"For the primary analysis, profiles are aligned by assigning each modality's own slice of maximum calcification to z=0. If blooming shifts the calcification peak location in EID-CT, this alignment could shift the EID-CT profile relative to the Micro-CT reference and inflate the MAE. Please clarify whether the Micro-CT reference z-grid is used as the common coordinate for all modalities.","section":"2.3.3"},{"comment":"The Micro-CT values in Table 2 are ellipse-derived percent area stenosis (Eq. 3) at the maximum-calcification cross-section, not values computed with Eq. 2. The observation that these values increase with calcification extent is consistent with Eq. 2 as well, because A_lumen/A_vessel decreases as calcification increases; there is no internal inconsistency between the two formulations.","section":"Table 2 / Eq. 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of physics in medicine and biology. The main technical risk is the unvalidated modality invariance of the multi-Otsu segmentation, which is acknowledged in the limitations but not addressed with data. A sensitivity analysis would substantially strengthen the paper. The statistical clustering of profiles within phantoms is also worth addressing before publication. The involvement of GE HealthCare co-authors and the use of a phantom set from a preprint by the same group should be carefully acknowledged, though the study design is a legitimate controlled comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read. This is the first quantitative stenosis-quantification comparison for the deep-silicon photon-counting detector against conventional EID-CT using realistic calcified coronary phantoms and an external Micro-CT reference. The study design is genuinely well-controlled: matched kVp and dose, system-specific resolution-optimized reconstructions, rigid registration to Micro-CT, longitudinal profile metrics, and a secondary ellipse-based analysis that agrees with the primary one. The reported reductions in MAE (3.10% to 1.62% for percent area stenosis, 0.55 to 0.31 mm² for vessel area) are plausible and the direction is consistent with prior CdTe PCCT work.\n\nThe main soft spot is exactly the one the paper flags in its own limitation section: the automated three-class multi-Otsu segmentation is load-bearing for every reported number, and modality-invariance of the thresholds is not demonstrated. Given the two systems differ in HU response, noise, kernel, and voxel size, an independent threshold-perturbation or reader-based sensitivity check would materially strengthen the central claim. The reader's concern about the 12 profiles coming from six phantoms is real but minor; the permutation test correctly treats each profile as a matched block, so the p-values are not egregiously liberal.\n\nThere is also a potentially important error in the manuscript text: Eq. 2 defines percent area stenosis as 100*A_lumen/A_vessel, but the Micro-CT values in Table 2 increase with calcification extent, which matches A_calc/A_vessel (or 100*(1 - A_lumen/A_vessel)). If the implementation used the printed equation, the reported numbers would be nonsense; more likely the equation has a typo. The authors should fix this before publication, because as written the metric is ambiguous.\n\nThe paper does not provide code or data, which limits independent replication, and the conclusions are appropriately scoped to static, resolution-optimized conditions. I do think the central argument holds up: the improvement is task-based, measured against an external reference, and the secondary analysis and the visible contours in the figures corroborate the primary result. I'd send this to peer review, asking for a threshold sensitivity analysis and a corrected Eq. 2. For someone working on CT performance evaluation or PCCT clinical translation, this is a useful data point.","headline":"A well-controlled phantom study with a plausible dSi-PCCT advantage, but the threshold-based segmentation and an apparent equation typo are the soft spots a referee should probe.","tokens_in":13009,"tokens_out":2755,"would_cite":true,"duration_ms":28278,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["87.57.Q-"],"model":"deepseek-v4-flash","headline":"Deep-silicon photon-counting CT improves coronary stenosis quantification in realistic calcified phantoms, cutting Micro-CT-referenced percent-area-stenosis error from 3.10% to 1.62%.","keywords":["photon-counting CT","deep silicon detector","coronary stenosis quantification","calcium blooming","coronary artery phantom","Micro-CT ground truth","spatial resolution","energy-integrating detector CT"],"falsifier":"A reader-based contour study on the same registered phantom images that finds no meaningful difference in Micro-CT-referenced percent area stenosis MAE between dSi-PCCT and EID-CT (for example, a gap below one percentage point) would indicate the headline improvement is an artifact of the multi-Otsu thresholding rather than a true resolution benefit; equally, a dynamic phantom experiment in which simulated cardiac motion erases the gap would bound the clinical relevance.","tokens_in":11965,"feed_emoji":"🩻","tokens_out":8435,"duration_ms":73843,"temperature":0.7,"pith_summary":"This paper sets out to establish that a deep-silicon photon-counting detector CT system (dSi-PCCT), by virtue of its higher native spatial resolution, quantifies coronary stenosis in calcified vessels more accurately than a conventional energy-integrating detector CT (EID-CT). The evidence comes from twelve realistic coronary phantom sections spanning four calcification geometries and three iodine concentrations, scanned statically and registered to Micro-CT as ground truth; the paper reports that Micro-CT-referenced percent area stenosis error fell from 3.10% to 1.62% mean absolute error (p=0.027) and segmented vessel-area error from 0.55 to 0.31 mm² (p=0.001). If correct, the result implies that the resolution limits of current detectors, not just motion or reconstruction choices, are a measurable cause of calcium-blooming stenosis overestimation, and it gives a quantitative baseline for judging future dynamic-phantom and clinical photon-counting CT studies.","feed_headline":"Photon-counting CT halves stenosis error in calcified arteries","feed_subtitle":"Against Micro-CT ground truth, mean stenosis error dropped from 3.10% to 1.62% in realistic phantoms.","key_machinery":"The argument is carried by a three-way scanner comparison on the same physical objects: a Micro-CT reference at 0.02 mm isotropic resolution, a conventional EID-CT scanner, and the dSi-PCCT scanner, with the two clinical systems run at dose-matched parameters but each at its resolution-optimized reconstruction (0.39×0.39×0.63 mm³ voxels with a Bone kernel for EID-CT versus 0.12×0.12×0.41 mm³ voxels with UHD Ultra and deep-learning reconstruction for dSi-PCCT). After rigid registration to Micro-CT, an automated three-class multi-Otsu thresholding step labels every voxel as lumen, calcification, or background; the union of the lumen and calcification masks defines the vessel area at each longitudinal position. Using the Micro-CT vessel area as the common denominator for both modalities converts the comparison into an isolation of segmentation and boundary-fidelity differences, which is what lets the paper attribute the improvement to reduced partial-volume and calcium-blooming effects from higher spatial resolution.","core_discovery":"On its own terms, the paper's central discovery is that, under matched static acquisition conditions, the dSi-PCCT system reproduces Micro-CT-referenced lumen and calcification geometry along calcified coronary phantom sections more faithfully than EID-CT, and that this boundary fidelity translates directly into lower percent-area-stenosis error. Across all 12 vessel sections the whole-profile mean absolute error in Micro-CT-referenced percent area stenosis drops from 3.10% (EID-CT) to 1.62% (dSi-PCCT), and segmented vessel-area MAE drops from 0.55 to 0.31 mm²; in the secondary ellipse-derived analysis the absolute deviation from Micro-CT ranges from 0.1% to 6.5% for dSi-PCCT versus 0.8% to 24.2% for EID-CT. The advantage is largest exactly where EID-CT struggles most: Type IV (circumferential) calcification and the lowest iodine concentration. The authors deliberately frame the primary metric as a task-based measure of imaging accuracy, distinct from conventional clinical percent stenosis, and they present the result as a baseline for dynamic and clinical follow-up.","pith_inferences":["My inference: because the phantoms were static, the 1.48-percentage-point MAE gap is probably an upper bound for clinical practice; cardiac motion will add error to both systems, and motion correction or compensation will be needed to retain the resolution benefit in vivo.","A testable extension the authors did not run: add controlled motion to the same phantom set, for example with a moving platform or gated acquisitions at different heart rates, to measure how much of the dSi-PCCT advantage survives realistic blur.","My inference: the automated multi-Otsu pipeline could systematically favor the sharper modality; a reader-contour study on the same registered images would confirm whether the improvement is true boundary fidelity instead of thresholding behavior.","Because the paper did not use dSi-PCCT's spectral capabilities, the combined effect of high resolution plus material decomposition (iodine/calcium separation) is an open question and could be larger than the resolution-only gain reported here."],"forward_implications":["If the result holds, high-resolution dSi-PCCT roughly halves the mean absolute error of percent area stenosis measurement in static calcified coronary phantoms, from 3.10% to 1.62%.","The largest improvements occur for circumferential (Type IV) calcification and at 10 mg/mL iodine, implying the technology helps exactly the high-risk calcified cases where conventional CCTA is least reliable.","The secondary ellipse-derived analysis shows the advantage also appears when the reference lumen is estimated from the image itself rather than from Micro-CT, suggesting the benefit is not an artifact of the ground-truth comparison.","The reported values provide a baseline for future dynamic-phantom and clinical CCTA studies to judge how much of the resolution gain survives cardiac motion and reader variability."],"supporting_citations":[{"why":"Supplies the realistic coronary artery phantom designs and fabrication methods that all three scanners image.","marker":"Pack et al. 2024"},{"why":"Defines the Type I-IV quadrant-based calcification grading used to create the plaque geometries.","marker":"Qi et al. 2016"},{"why":"Supplies SimpleITK, the registration tool that aligns EID-CT and dSi-PCCT volumes to the Micro-CT reference.","marker":"Lowekamp et al. 2013"},{"why":"Prior evidence that photon-counting detector CT reduces calcium blooming, which the phantom experiment tests with an independent ground truth.","marker":"Koons et al. 2024"},{"why":"Ex vivo finding that ring-shaped calcifications enlarge measurement error, used to interpret the Type III and Type IV results.","marker":"Sandstedt et al. 2021"},{"why":"Defines quantitative CCTA reference-lumen standards, which the paper distinguishes from its task-based Micro-CT-referenced metric.","marker":"Nieman et al. 2024"}],"fun_headline_variants":["Silicon photon-counting CT nearly halves stenosis error","Deep-silicon PCCT cuts stenosis error by half vs EID-CT","Stenosis error drops 48% with silicon photon-counting CT","PCCT improves coronary stenosis accuracy in calcified vessels","dSi-PCCT shows major stenosis error reduction over EID-CT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automatic thresholding that separates lumen, calcium, and background is equally fair to both scanner types after registration; if the blurrier EID-CT images are systematically harder to threshold, the reported error gap could come from the segmentation tool rather than from the detector.","fun_headline_variants_meta":{"raw":{"variants":["Silicon photon-counting CT nearly halves stenosis error","Deep-silicon PCCT cuts stenosis error by half vs EID-CT","Stenosis error drops 48% with silicon photon-counting CT","PCCT improves coronary stenosis accuracy in calcified vessels","dSi-PCCT shows major stenosis error reduction over EID-CT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00039,"raw_usage":{"total_tokens":2158,"prompt_tokens":1153,"completion_tokens":1005,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":769,"completion_tokens_details":{"reasoning_tokens":916}},"tokens_in":769,"tokens_out":1005,"duration_ms":10257,"temperature":1.0,"reasoning_tokens":916,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:45:01.741268+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader-based contour study on the same registered phantom images that finds no meaningful difference in Micro-CT-referenced percent area stenosis MAE between dSi-PCCT and EID-CT (for example, a gap below one percentage point) would indicate the headline improvement is an artifact of the multi-Otsu thresholding rather than a true resolution benefit; equally, a dynamic phantom experiment in which simulated cardiac motion erases the gap would bound the clinical relevance.","supporting_citations":[],"review_version":1}