{"id":"59f154c5-e5cb-446b-9348-7aa6befd6b41","arxiv_id":"2411.16992","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using CycleGAN synthetic CT and U-Net auto-segmented contours in a hybrid point-to-distance metric improves prostate and rectum DIR on pelvic CBCT, with prostate DSC rising from 0.61 to 0.82 for auto and 0.89 for expert contours.","lead":"This paper combines an AI-generated synthetic CT and automatic organ outlines with a structure-aware registration metric to improve how planning CT and cone-beam CT images are aligned in prostate radiotherapy. It reports better organ overlap and smaller fiducial marker shifts than intensity-only alignment, pointing toward faster, more accurate adaptive treatment workflows.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DSC/HD gains for CycleGAN PD are scored with the same U-Net sCT contours that drive the PD metric; bladder/rectum improvements lack independent validation, and bladder HD worsens.","rationale":"The reader's verdict of CONDITIONAL is appropriate, and my concern does not change it; it sharpens the condition. The reader's weakest assumption was that CycleGAN preserves true anatomy in sCT. My concern is closely related but more specific: even if the sCT is anatomically accurate, the evaluation pipeline conflates registration accuracy with segmentation accuracy because the moving contours used for scoring CycleGAN PD are the same U-Net contours used to drive the PD metric. The prostate fiducial result is a genuine independent confirmation for that organ, so the paper's central mechanism is not invalidated. However, the bladder and rectum claims are not independently supported, and the bladder HD result moves in the wrong direction. Therefore the paper should be accepted only conditionally, with the scope explicitly limited to prostate fiducial-validated improvements until the scoring-contour confound is resolved. My proposed test directly addresses the confound by holding the scoring contours fixed across all three arms. This does not require retraining any model, only re-evaluating existing deformation fields with expert contours on the moving image, so it is feasible for the authors to perform.","tokens_in":12051,"tokens_out":4986,"duration_ms":52210,"concrete_test":"For the 7 DIR validation patients, recompute post-registration DSC and 95% HD for all three workflows using identical expert-delineated contours: fixed CT expert contours and moving CBCT expert contours warped by the final DVF from each workflow. Do not use U-Net sCT contours in the scoring of CycleGAN PD. If CycleGAN PD still significantly improves bladder and rectum DSC/HD over No PD, and bladder 95% HD no longer worsens, the concern is resolved; otherwise the reported organ-level gains are artifacts of the scoring contour source.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CycleGAN-based auto-segmentation improves DIR accuracy for all three organs rests on a scoring asymmetry. For the CycleGAN PD arm, the contours that define the PD metric are the U-Net predictions on sCT, and the reported DSC and 95% HD appear to be computed using those same predicted contours as the moving-image structures. For the No PD and Expert PD arms, the moving structures are necessarily expert contours on CBCT. Thus the CycleGAN PD DSC/HD partly measures how well the optimized DVF makes U-Net contours overlap fixed expert contours, not how well true anatomy is aligned. This is not merely a philosophical issue: the prostate fiducial results provide independent support for that organ (fiducial separation 4.07 mm, p = 0.014), but no fiducials exist for bladder or rectum. For the bladder, the reported 95% HD actually worsens under CycleGAN PD (21.48 mm vs 19.59 mm, p > 0.05), yet the abstract states improvements were observed for bladder and rectum. The manuscript also reveals that the U-Net's sCT labels were 'contours transferred from CBCT' (Section 2.1/2.4), so the segmentation accuracy on sCT is calibrated against CBCT-derived labels rather than independent anatomical truth. Without a common scoring contour set across all three arms, the claim that CycleGAN auto-segmentation improves DIR accuracy outside the prostate is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an improved deformable image registration (DIR) workflow for CT-to-CBCT registration in adaptive radiotherapy. A hybrid similarity metric combines intensity matching with a point-to-distance (PD) structural term. A CycleGAN generates synthetic CT (sCT) from CBCT to improve soft-tissue contrast, and a 3D U-Net auto-segments the prostate, bladder, and rectum on the sCT. Three workflows are compared: intensity-only DIR (No PD), DIR using U-Net contours on sCT (CycleGAN PD), and DIR using expert manual contours (Expert PD). On a validation cohort of 7 patients, the paper reports improved DSC and 95% HD for prostate and rectum, improved DSC for bladder, and reduced fiducial separation for the prostate, with p<0.05 for many pairwise comparisons.","tokens_in":12357,"tokens_out":5850,"duration_ms":51211,"significance":"The paper addresses a clinically relevant problem—improving DIR accuracy on low-contrast CBCT images—and combines image translation, auto-segmentation, and structure-guided registration in a plausible pipeline. The prostate fiducial results provide an independent and clinically meaningful metric, which is a notable strength. However, the central claim of accuracy improvement across all three organs is not fully supported: the bladder 95% HD worsens under CycleGAN PD in Table 2, and the evaluation protocol appears to use different contour sources across arms, confounding the DSC/HD comparisons. If the confound is resolved, the method could be valuable, but as presented the evidence for all-organ improvement is incomplete.","major_comments":[{"comment":"The manuscript does not specify which contours are used to compute DSC and 95% HD in each arm. In the CycleGAN PD arm, the only moving-image structures available are the U-Net contours on sCT, whereas in the No PD and Expert PD arms the moving structures are presumably expert contours on CBCT. If this is the case, the reported DSC/HD differences between CycleGAN PD and No PD are confounded by the contour source, not solely by the registration method. Please state the exact evaluation protocol, or recompute metrics using a common contour set (e.g., expert CBCT contours propagated by each DVF). This is load-bearing because the abstract's claim of improved accuracy for bladder and rectum rests on these DSC/HD numbers, and the bladder 95% HD in Table 2 actually worsens under CycleGAN PD (21.48 mm vs 19.59 mm, p > 0.05).","section":"§2.5, Table 2"},{"comment":"The weights λ and η_n in Eq. (1) are never reported. These hyperparameters control the balance between intensity similarity, regularization, and structural guidance, and the behavior of the hybrid metric depends critically on them. Without their values or a sensitivity analysis, the experiments are not reproducible. Please report the values used for all experiments and describe how they were selected or tuned.","section":"§2.2, Eq. (1)"},{"comment":"The CycleGAN is 2D and processes axial slices independently, yet no quantitative assessment of sCT quality or slice-to-slice consistency is provided; Figure 3 is only a visual example. Furthermore, Section 2.1 states that the sCT labels used to train the U-Net were 'contours transferred from CBCT,' meaning the segmentation accuracy on sCT is not validated against independent ground truth. Please provide quantitative sCT evaluation (e.g., mean absolute error against CT, and a 3D consistency metric) and, if possible, assess U-Net segmentation on sCT against expert contours drawn directly on CT.","section":"§2.3, §2.4"},{"comment":"The validation cohort consists of only 7 patients. With n=7, the reported p-values from ANOVA/Kruskal-Wallis and post-hoc tests are highly sensitive to outliers and have low statistical power. The paper should report effect sizes and confidence intervals, describe how the 7 validation patients were selected from the 14-case holdout, and discuss the limitation posed by the small validation cohort.","section":"§2.1, Table 2"}],"minor_comments":[{"comment":"The abstract states that 'Improvements were also observed for the bladder and rectum,' but Table 2 shows that the bladder 95% HD worsens under CycleGAN PD (21.48 mm vs 19.59 mm) and is not statistically significant. Please adjust the wording to reflect the mixed result.","section":"Abstract"},{"comment":"The sentence 'Overall, percentage improvements in DSC and 95% HD ranged from a 56.64% increase in DSC for the bladder to a 73.42% reduction in 95% HD' is unclear and appears inconsistent with Table 2. Please specify which comparisons yield these percentages and correct the values if needed.","section":"§3.3"},{"comment":"For the bladder 95% HD row, the p-value column lists both ANOVA (0.1414) and Kruskal-Wallis (0.005). This is confusing; specify which test was used and whether the pairwise comparison was adjusted for multiple testing.","section":"Table 2"},{"comment":"In the sentence following Eq. (1), the phrase 'π_n′ is the moving image' appears to be a typo; presumably π_n′ denotes a point set on the moving image boundary. Please correct.","section":"§2.2, Eq. (1)"},{"comment":"With n=7 per group, the Shapiro-Wilk normality test has very low power. Please consider reporting the raw distributions or using a more robust approach, and note this limitation in the statistical analysis.","section":"§2.7"}],"recommendation":"major_revision","confidential_remarks":"The work is within the journal's scope and addresses a practical problem. The main concern is the evaluation confound across arms; if the authors can provide a common-contour evaluation and address the bladder HD issue, the paper could be acceptable. The abstract overstates the bladder result and should be revised. The missing hyperparameters for Eq. (1) should also be supplied."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate integration-and-evaluation paper, not a breakthrough. The novelty is assembling CycleGAN sCT + U-Net contours as inputs to a point-to-distance term in B-spline DIR, and testing it on 7 pelvis cases. Prostate results are genuinely supported by independent fiducial markers: fiducial separation drops from 8.95 to 4.07 mm. That is the strongest evidence in the paper and it holds up.\n\nWhat it does well: the workflow is clearly described, components are standard and cited, the evaluation uses clinically meaningful metrics (DSC, 95% HD, fiducials), and the authors are honest about limitations (segmentation quality dependence, dataset size). The open-source Plastimatch context helps reproducibility in principle.\n\nWhere I'd push back: the abstract and discussion claim improvements 'across all structures' including bladder and rectum. The bladder 95% HD actually worsens in the CycleGAN PD arm (21.48 vs 19.59 mm, p > 0.05), and the authors themselves note it is not significant. That discrepancy with the abstract is a real accuracy problem in the manuscript.\n\nMore important, the stress-test concern is correct: for the CycleGAN PD arm, the contours used to compute DSC and 95% HD are the same U-Net sCT contours that drive the PD metric. So the registration is being scored partly on how well the DVF aligns the auto-contours to the fixed expert contours, not on true anatomical alignment. The prostate fiducial results provide independent support for that organ, but there are no fiducials for bladder or rectum. And because the U-Net's sCT labels were transferred from CBCT contours, the segmentation accuracy itself is calibrated against CBCT-derived truth, not independent anatomy. So the claim that auto-segmentation improves DIR accuracy outside the prostate is not established by the data as presented.\n\nOther soft spots: the PD weights (lambda and eta_n) are not reported, which makes the method hard to reproduce; there is no raw-CBCT auto-segmentation baseline, so we can't tell whether sCT is actually necessary; and the validation cohort is seven patients, so confidence intervals are wide. Code/data are not provided, which weakens the 'open' framing.\n\nWho this is for: people working on CBCT-guided adaptive pelvic radiotherapy, especially those interested in contour-guided DIR and synthetic CT. It's a useful data point, not a definitive result. I'd send it to review but expect heavy revision: the abstract needs to be reined in, the scoring asymmetry needs to be addressed or at least discussed, and the missing parameters and baseline need to be supplied.","headline":"A legitimate integration-and-evaluation study for CBCT-guided pelvic DIR, with prostate improvements supported by fiducials, but the all-organ claim is undermined by a scoring asymmetry and a bladder result that goes the wrong way.","tokens_in":12908,"tokens_out":1740,"would_cite":true,"duration_ms":16489,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Auto-segmented contours from CycleGAN-corrected synthetic CT markedly improve CT-to-CBCT deformable registration, nearly matching expert contours.","keywords":["Deformable image registration","Adaptive radiation therapy","CycleGAN","Auto-segmentation","Hybrid similarity metric","Point-to-distance metric","CT-CBCT registration","U-Net"],"falsifier":"Take a planning CT with expert contours, apply a known deformation to produce a synthetic CBCT, and run all three workflows; if the CycleGAN pipeline aligns true anatomy, its registration error against the known displacement should track the Expert PD result, and any large divergence would show the auto-contours are pulling the metric toward synthetic boundaries.","tokens_in":11862,"feed_emoji":"🩻","tokens_out":10852,"duration_ms":91646,"temperature":0.7,"pith_summary":"Deformable image registration between planning CT and daily cone-beam CT struggles in low-contrast soft tissue because pure intensity matching has no reliable landmarks there. This paper argues that adding a point-to-distance penalty on organ surfaces to the registration cost fixes most of that loss, and that the organ contours feeding that penalty can be produced automatically rather than by hand. The pipeline translates CBCT slices into synthetic CT with a CycleGAN, segments prostate, bladder, and rectum with a 3D U-Net, and uses those contours in a hybrid metric during B-spline registration. On a 14-case pelvic validation set the auto-contour workflow raised the prostate overlap score from 0.61 with intensity-only registration to 0.82, versus 0.89 with expert contours, and cut fiducial marker separation from 8.95 mm to 4.07 mm. If it holds, adaptive radiotherapy can get structure-guided registration without the manual-contouring step that currently slows on-line treatment adaptation.","feed_headline":"AI-drawn contours nearly match expert-guided CT-CBCT registration","feed_subtitle":"Adding auto-segmented organ boundaries to the registration metric lifts prostate Dice to 0.82 and halves fiducial error.","key_machinery":"The load-bearing object is the point-to-distance (PD) penalty added to the B-spline registration cost. For each structure the penalty sums, over contour points $\\pi_n$ on the fixed image, the value of an unsigned distance map at the transformed point in the moving image, weighted by $\\eta_n$: $$C = \\sum_{T(\\$\\theta$)\\in\\$\\Omega$}\\Psi(f,m) + \\$\\lambda$ S(\\nu) + \\sum_{n=1}^{N}\\eta_n \\sum_{\\pi_n\\in\\Pi_n}|d_{\\text{map}}(\\pi'_n)|.$$ Minimizing this cost with L-BFGS-B forces the moving contour onto the fixed contour even where intensities give no gradient. The paper's contribution is to feed this term automatically: a 2D CycleGAN generates synthetic CT with better soft-tissue contrast from CBCT, and a 3D U-Net segments the three pelvic organs on those sCT images, so the PD term no longer depends on manual contouring. The cycle-consistency and identity losses in the CycleGAN are what keep the translated images anatomically faithful enough for segmentation to be meaningful.","core_discovery":"The paper's central claim is that the hybrid similarity metric—intensity similarity plus a weighted point-to-distance (PD) term computed from organ boundaries—is what lifts CT-CBCT registration accuracy, and that the boundaries can be machine-drawn rather than expert-drawn. In the reported prostate results, adding CycleGAN-based auto-contours improved DSC from 0.61 ± 0.18 to 0.82 ± 0.13, reduced 95% Hausdorff distance from 11.75 mm to 4.86 mm, and reduced fiducial separation from 8.95 mm to 4.07 mm; expert manual contours yielded 0.89 ± 0.05, 3.27 mm, and 4.11 mm, respectively. The bladder and rectum also improved on DSC, though the bladder's 95% Hausdorff distance with auto-contours was not statistically different from intensity-only registration. The authors conclude that AI-based synthetic CT plus auto-segmentation can replace manual contours in the PD metric and that the hybrid metric is the main source of the accuracy gain.","pith_inferences":["Beyond the paper: the near-identical fiducial separation between CycleGAN PD (4.07 mm) and Expert PD (4.11 mm), despite a DSC/HD gap, suggests auto-contours are centered correctly but noisier in shape; a per-contour surface-error map would test this directly.","Beyond the paper: the bladder's non-significant 95% HD improvement hints that a single PD weight per structure may be too coarse for hollow, volume-changing organs; varying $\\eta_n$ per organ could recover the expert-level boundary accuracy.","Beyond the paper: deliberately corrupting auto-contours (erosion, dilation, or random slice replacement) and measuring the resulting registration error would map how much contour quality the PD metric tolerates before it loses to intensity-only registration.","Beyond the paper: since the architecture is modality-agnostic, retraining on MR-CBCT or PET-CBCT pairs is a natural next step, with the same unpaired-translation advantage."],"forward_implications":["Online adaptive radiotherapy could use the auto-contour pipeline to drive structure-guided registration without a manual contouring step.","The roughly 55% reduction in prostate fiducial separation (8.95 mm to 4.07 mm) supports the possibility of tighter planning margins and less healthy-tissue dose.","Because the PD term is structure-agnostic, the same hybrid cost should transfer to other organs and imaging modalities whenever a segmentation model is available.","Because the CycleGAN requires only unpaired CT and CBCT volumes, the sCT step can be reproduced at institutions that lack paired training data."],"supporting_citations":[{"why":"Introduces the point-to-distance hybrid similarity metric that this paper automates.","marker":"Shah et al 2021"},{"why":"Provides the CycleGAN architecture for unpaired CBCT-to-sCT translation with cycle-consistency loss.","marker":"Zhu et al 2017"},{"why":"Supplies the 3D U-Net architecture used to segment prostate, bladder, and rectum on CT and sCT.","marker":"Çiçek et al 2016"},{"why":"Defines the U-Net encoder-decoder with skip connections that the segmentation model builds on.","marker":"Ronneberger et al 2015"},{"why":"The Pelvic Reference Dataset is the main source of CT-CBCT pairs and expert contours for training and validation.","marker":"Yorke et al 2019"},{"why":"Provides the registration toolkit used to implement B-spline DIR with the hybrid metric.","marker":"Sharp et al 2010"},{"why":"Establishes the B-spline free-form deformation model that the registration cost regularizes.","marker":"Rueckert et al 1999"},{"why":"Documents the PD metric's dependence on manual contours, the bottleneck this paper removes with auto-segmentation.","marker":"Shah 2022"}],"fun_headline_variants":["Auto-segmented contours boost CT-CBCT registration accuracy","AI contours lift CT-CBCT Dice from 0.61 to 0.82","Hybrid metric plus AI contours trims registration error","Auto-segmentation aids low-contrast CBCT registration","AI-drawn contours rival manual ones in CT-CBCT DIR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the CycleGAN's synthetic CT images show the patient's real anatomy rather than plausible-looking artifacts, because the auto-contours are only useful if they trace true organ boundaries; if the sCT fabricates or blurs structures, the point-to-distance penalty will align to synthetic shapes and the reported accuracy gains would not reflect real anatomy.","fun_headline_variants_meta":{"raw":{"variants":["Auto-segmented contours boost CT-CBCT registration accuracy","AI contours lift CT-CBCT Dice from 0.61 to 0.82","Hybrid metric plus AI contours trims registration error","Auto-segmentation aids low-contrast CBCT registration","AI-drawn contours rival manual ones in CT-CBCT DIR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000642,"raw_usage":{"total_tokens":3062,"prompt_tokens":1165,"completion_tokens":1897,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":781,"completion_tokens_details":{"reasoning_tokens":1806}},"tokens_in":781,"tokens_out":1897,"duration_ms":12304,"temperature":1.0,"reasoning_tokens":1806,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:39:42.999144+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a planning CT with expert contours, apply a known deformation to produce a synthetic CBCT, and run all three workflows; if the CycleGAN pipeline aligns true anatomy, its registration error against the known displacement should track the Expert PD result, and any large divergence would show the auto-contours are pulling the metric toward synthetic boundaries.","supporting_citations":[],"review_version":1}