{"id":"f17c623b-5524-47d8-97d5-490386a4ab8b","arxiv_id":"2501.18109","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In 794 prostate cancer patients, 2D-Pix2Pix produced the best US-to-MRI synthetic scans by SSIM (0.855), but only 76 of 186 radiomics features survived translation intact.","lead":"The paper compares ten image-translation networks that turn prostate ultrasound scans into synthetic MRI, using 794 patients. It finds 2D-Pix2Pix gives the highest similarity scores, but doctors and radiomics checks show important diagnostic features are still lost.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark rests on an unverified US–MRI registration; without a registration-error analysis, SSIM and radiomics comparisons may reflect alignment error rather than network performance.","rationale":"The reader's weakest assumption—adequate US–MRI co-registration and valid transfer of MRI masks to synthetic images—is exactly the load-bearing concern. It underpins every voxel-level and feature-level comparison in the paper, including the headline SSIM ranking and the radiomics preservation counts. The paper provides no quantitative registration error assessment, and the phrase 'aligned by clinical collaborators' is not a reproducible or verifiable registration procedure. My proposed check would settle whether this concern actually lands: if registration error is small, the quantitative comparisons may stand; if it is large, the central claims would need to be recomputed under a validated registration. Because the reader already marked the verdict CONDITIONAL and this concern is fully consistent with that assessment, no adjustment to the reader's verdict is needed. I am not raising concerns about the qualitative or classification sections because they are downstream of the same alignment issue and because the classification comparisons, while imperfect, are not the primary point of failure identified here.","tokens_in":12950,"tokens_out":5147,"duration_ms":59542,"concrete_test":"On a random subset of 20–30 patients, independently register the preprocessed US and MRI volumes with a validated nonrigid algorithm (e.g., ANTs or Elastix with mutual information, initialized from the provided prostate segmentations). Report target registration error on manually placed landmarks (prostate apex, base, urethra, and capsule boundary) in mm before and after registration, and recompute 2D-Pix2Pix SSIM and the Group 2 radiomics correlations under this registration. If SSIM shifts by more than ~0.02 or feature group assignments change for more than 10% of features, the reported quantitative comparisons are alignment-dominated rather than network-dominated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims—2D-Pix2Pix's SSIM of 0.855, the preservation of 76/186 radiomics features, and the synthetic-vs-original MRI comparisons—all depend on voxel-level correspondence between US and MRI. Section 2.1 states only that 'US and MRI images were aligned by clinical collaborators, then cropped around the prostate center and resampled to a standardized size'; Section 2.3 adds that 'identical masks were used to extract these features from different images.' No registration method, transformation model, or error metric is reported. Because the US and MRI are separate acquisitions and the prostate deforms under the transrectal probe, even small residual misalignment can change SSIM substantially and make a radiomic feature in one volume correspond to different tissue in the other. The paper's own Discussion lists limitations but does not mention registration error. If misalignment is on the order of a few voxels, the reported feature correlations and network rankings are not trustworthy, independent of the translation models' true performance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript benchmarks ten 2D/3D image-to-image translation networks for ultrasound-to-MRI synthesis in 794 prostate cancer patients, using four quantitative metrics, radiomic feature correlations, qualitative assessment by seven physicians, and downstream classification. The main reported findings are that 2D-Pix2Pix achieves the highest SSIM (0.855±0.032), that 76 of 186 radiomic features are preserved after translation according to Spearman correlation, and that classifiers trained and tested on synthetic MRI reach accuracy/AUC around 0.93, outperforming classifiers based on US. The authors conclude that supervised paired translation currently outperforms diffusion-based models on global similarity but that low-level clinically relevant features remain imperfectly preserved.","tokens_in":13164,"tokens_out":7270,"duration_ms":66730,"significance":"If the quantitative claims were validated, the study would provide a useful comparative benchmark for I2I models in prostate US-to-MRI synthesis, combining global metrics with radiomics and clinician evaluation. The use of 794 patients, ten networks, publicly shared code, and standardized radiomics extraction (ViSERA/IBSI) are strengths. The radiomics comparison is performed between synthetic and real MRI, so it is not itself circular; the classification section, however, uses a train-and-test-on-synthetic protocol that measures internal consistency rather than clinical diagnostic value. The significance of the reported network ranking is conditional on an unverified US–MRI registration assumption, which has not been demonstrated in the manuscript.","major_comments":[{"comment":"The entire voxel-level comparison depends on the co-registration of US and MRI. Section 2.1 states only that images 'were aligned by clinical collaborators' and then resampled, and Section 2.3 notes that 'identical masks were used to extract these features from different images.' No registration method, transformation model, or residual-error metric is reported. Under transrectal-probe deformation, residual misalignment of even a few voxels can change SSIM and shift radiomic features to different tissue, so the reported network ranking (Fig. 1), the 76/186 preservation count, and the Group 1–3 radiomics stratification may partly reflect alignment error rather than network performance. The Discussion's limitation paragraph does not mention this. Please report a registration-error analysis or otherwise demonstrate that voxel and mask correspondences are sufficiently accurate for the claimed voxel-level and feature-level comparisons.","section":"§2.1, §2.3"},{"comment":"The claim that synthetic MRI improves classification over US is not supported by the current protocol. Experiments C3, C6, C11, and other synthetic-image combinations train and test the classifier on synthetic MRI produced by the same I2I networks; as the Discussion states, using 'consistent synthetic data for training and testing with PCA and RandF eliminated the domain gap.' This setup measures internal reproducibility of the synthetic domain, not diagnostic value on real clinical data, so the reported ~0.93 accuracy/AUC cannot be compared clinically with US-based classification. Please add external evaluation in which models trained on synthetic images are tested on held-out real MRI or real US, and models trained on real images are tested on synthetic images, with explicit confusion matrices.","section":"§2.4, §3.4, Discussion"},{"comment":"The statement that 2D-Pix2Pix 'significantly outperformed all other generative models' is based on pairwise paired t-tests across the ten networks without correction for multiple comparisons. With nine pairwise comparisons per metric, the reported P<0.01 should be adjusted, or a global test with post-hoc correction should be reported. This does not necessarily change the ranking, but it is required to support the 'substantially outperformed' claim in the abstract and conclusion.","section":"§3.1"}],"minor_comments":[{"comment":"The count of Group 2 features is internally inconsistent: the text says 76 radiomic features, but the listed subcategory counts (5 IS, 17 IH, 2 IVH, 26 GLCM, 6 NGLDM, 12 GLRLM, 3 GLSZM, 3 GLDZM, 1 NGTDM) sum to 75, and 18 + 76 + 93 = 187 rather than the stated total of 186. Please correct these numbers.","section":"§3.3"},{"comment":"The abstract says 2D-Pix2Pix outperformed 'the other 7 networks,' but ten networks were evaluated; this should read 'the other nine networks.'","section":"Abstract"},{"comment":"The heading '2.4. Classification Analysis' duplicates the numbering of the earlier '2.4. Qualitative Analysis' section, and the qualitative section refers to questions in 'Table 1, rows 2–9' when the questions appear in Table 2.","section":"§2.4"},{"comment":"The qualitative results text refers to Q7, Q8, and Q9, while Table 2 reports only eight questions (Q1–Q8), and the question labels are inconsistently mapped; please align the question numbering between text and table.","section":"§3.2"},{"comment":"The citation of Koo and Li [46] concerns intraclass correlation coefficients, not Spearman correlation coefficients; a source for the correlation cutoffs, or a direct justification of the 0.50 threshold, should be provided.","section":"§3.3"},{"comment":"The threshold 'SSIM > 0.85' is used to define high-performance networks, but 2D-Pix2Pix's mean SSIM is 0.855 with a standard deviation of 0.032, so some folds lie below the threshold; please clarify how the threshold was applied in the radiomics analysis.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript opens with a statement that it has already been published in IJCARS; if this submission is for a different venue, the editors may wish to confirm overlapping-publication policy. I have not treated this as a scientific flaw. The main technical risks are the unverified registration and the synthetic-only classification protocol, both of which are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a look despite its flaws. What's genuinely new is the scale and breadth: ten 2D/3D I2I networks on 794 prostate cancer patients, with IBSI-standardized radiomics, seven-physician qualitative review, and downstream classification. The finding that 2D-Pix2Pix wins on SSIM (0.855) and preserves 76 of 186 radiomic features, while half are lost, is a useful benchmark for the field. Code is shared, and the 5-fold CV at least supports the relative ranking.\n\nThe soft spots are real. The biggest is registration. Section 2.1 says only that US and MRI were 'aligned by clinical collaborators'; no transformation model, no error metric. Given that US and MRI are separate acquisitions and the prostate deforms under the probe, unquantified misalignment could drive the SSIM differences and radiomic correlations just as much as the networks' true performance. The stress-test note is on target: this is load-bearing, not a footnote.\n\nSecond, the classification section is partly circular. Training and testing classifiers on synthetic MRI data and then saying the domain gap is 'eliminated' measures internal consistency, not diagnostic value. The claim that synthetic MRI beats US for classification needs at least a synthetic-trained/real-MRI-tested arm. As written, it's overreach.\n\nThe thresholds (SSIM>0.85, Spearman>0.50) are arbitrary but minor; a sensitivity analysis would help. The qualitative part is thin but clearly labeled exploratory.\n\nBottom line: the benchmark contribution is solid enough to deserve referee time, but the central quantitative claims need a registration-error analysis and the classification claim needs non-circular validation. I'd send it to peer review with major-revision expectations.","headline":"A useful multi-network benchmark whose headline numbers depend on an unverified registration step and a circular classification test.","tokens_in":13664,"tokens_out":1991,"would_cite":true,"duration_ms":19971,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"US-to-MRI translation can power prostate cancer risk prediction, but only 76 of 186 radiomic features survive.","keywords":["image-to-image translation","prostate cancer","ultrasound to MRI synthesis","radiomics","synthetic MRI evaluation","Pix2Pix","diffusion models","outcome prediction"],"falsifier":"Compute point-to-point registration error on the dataset's tracked biopsy cores or other landmarks, then re-run the radiomic rank correlations and SSIM comparisons excluding the worst-registered cases; if the correlations shift dramatically or the 76-preserved/93-lost split changes, the central comparison would not stand. A simpler visual check is also available: the paper's Figure 2 shows a synthetic hyperintensity with no counterpart in the original MRI and a true lesion not preserved—if coregistering those two images with a different algorithm makes the bright spot coincide with a real structure, that discrepancy is registration-driven rather than a synthesis failure.","tokens_in":1844,"feed_emoji":"🩻","tokens_out":4377,"duration_ms":114373,"temperature":0.7,"pith_summary":"Image-to-image translation aims to turn cheap ultrasound into MRI-like images, and this paper asks whether the result is clinically trustworthy for prostate cancer. On 794 patients and ten networks, 2D-Pix2Pix produced the most similar synthetic MRI ($SSIM = 0.855 \\pm 0.032$), beating diffusion and unpaired GAN methods. Similarity, however, overstates usefulness: only 76 of 186 standardized radiomic features survived translation with acceptable correlation, and seven physicians could always tell synthetic from real MRI, often because true lesions disappeared or fake bright regions appeared. The most concrete gain is automated risk classification: classifiers trained and tested on synthetic MRI reached about 0.93 accuracy and AUC, beating ultrasound and approaching real-MRI performance of 0.95. The paper thus separates global image quality from diagnostic feature preservation.","feed_headline":"Synthetic MRI beats ultrasound for prostate risk scoring","feed_subtitle":"SSIM 0.855, yet only 76 of 186 radiomic features survive—still, synthetic MRI improves risk scoring over ultrasound.","key_machinery":"The load-bearing mechanism is a three-stage evaluation stack built around aligned US/MRI volumes. Ten networks—paired (Pix2Pix), unpaired GAN variants (CycleGAN, DiscoGAN, DualGAN, GcGAN), reconstruction-style 3D models (AutoEncoder, UNET), and diffusion models (ContourDiff, Med-DDPM)—translate each US volume into a synthetic MRI. Second, 186 standardized radiomic features are extracted from the segmented prostate using the same masks for original and synthetic MRI, and a Spearman rank correlation between original and synthetic feature values assigns each feature to one of three groups: preserved by most networks, preserved only by high-SSIM networks, or lost by all networks. Third, synthetic images are scored by seven physicians on eight questions and then fed to a principal-component-analysis-plus-Random-Forest radiomics classifier and a ResNet50 deep classifier. The rank-correlation grouping is what turns raw SSIM scores into a statement about which diagnostically relevant information survives translation.","core_discovery":"The central discovery is a dissociation between global similarity and clinically meaningful fidelity. 2D-Pix2Pix outperformed all nine other networks on MAE, MSE, SSIM, and PSNR ($P<0.01$), with average SSIM $0.855\\pm 0.032$, but radiomic-feature analysis showed the same network preserved only 76 of 186 standardized features at a correlation threshold of 0.50, while 93 features remained undetectable by any network. Seven experienced physicians, despite seeing images with SSIM above 0.85, consistently distinguished synthetic from original MRI and rated diagnosis as harder, citing artifacts; one representative case shows a false-positive hyperintensity introduced by synthesis and a true lesion omitted. Despite these limitations, radiomics-based classification of high- versus low-risk prostate cancer using synthetic MRI, with principal component analysis plus a Random Forest classifier, achieved average accuracy and AUC of about 0.93, exceeding the 0.88/0.87 achieved with ultrasound and approaching the 0.95/0.94 of real MRI. The paper concludes translation networks need improvement at lesion-level fidelity, but synthetic MRI already has measurable value over the source modality for downstream outcome prediction.","pith_inferences":["A testable extension is to use the same 186-feature rank-correlation grouping as a standardized 'radiomic preservation profile' for any future US-to-MRI network, so results across studies can be compared feature-by-feature rather than by SSIM alone.","Because the paper's alignment relies on clinical collaboration without a reported registration-error analysis, an independent check using the tracked biopsy-core coordinates in the dataset would clarify how much of the measured feature loss is due to translation rather than residual misalignment.","The false-positive and false-negative lesion discrepancies in the qualitative results suggest a direct clinical probe: asking radiologists to mark suspicious regions in synthetic versus original MRI and comparing those marks against biopsy-confirmed lesions would quantify how often synthesis alters the actionable finding.","If the synthetic-consistent training trick generalizes, a natural next step is to test whether classifiers trained on synthetic MRI transfer to real MRI at test time, or whether the benefit disappears when the domain changes."],"forward_implications":["If high SSIM can coexist with loss of half of radiomic features, SSIM and PSNR alone are insufficient acceptance criteria for synthetic images in clinical use; feature-preservation reporting should be part of any translation benchmark.","Networks that preserve Group-1 features even at lower SSIM, such as CycleGAN variants, may be worth using for specific tasks, so the best network depends on which features the downstream task actually needs.","Radiomics classifiers can be trained and tested entirely on synthetic MRI and still outperform the original ultrasound, suggesting a deployment path where translation acts as a preprocessing step to improve risk stratification.","The 93 Group-3 features that no network recovers define a concrete target: any future network that lifts those correlations above 0.50 would be a measurable advance in lesion-level fidelity.","Because the radiomics framework consistently beat ResNet50 on this dataset, the paper implies that for moderate-sized cohorts hand-crafted radiomics plus dimensionality reduction is currently more reliable than deep feature learning."],"supporting_citations":[{"why":"Supplies the 794-patient aligned ultrasound/MRI dataset with segmentation masks and biopsy-tracked risk scores; without it the translation and radiomics comparisons have no data.","marker":"[36]"},{"why":"Provides the review of GAN architectures that grounds the selection of the ten 2D/3D translation networks compared in the study.","marker":"[40]"},{"why":"Defines the conditional diffusion model (3D-Med-DDPM) that serves as one of the two diffusion baselines outranked by 2D-Pix2Pix.","marker":"[41]"},{"why":"Defines the contour-guided diffusion model (2D-ContourDiff) used as the second diffusion baseline for the low-level feature comparison.","marker":"[42]"},{"why":"Supplies the standardized radiomics feature generator used to extract the 186 features whose Spearman correlations form the paper's three-group feature-preservation analysis.","marker":"[43]"},{"why":"Supplies the ResNet50 architecture used as the deep-learning classifier that the radiomics framework outperforms.","marker":"[45]"},{"why":"Provides the correlation-coefficient thresholds (poor, moderate, good, excellent) used to set the 0.50 cutoff that defines the three radiomic feature groups.","marker":"[46]"}],"fun_headline_variants":["Synthetic MRI from ultrasound improves prostate risk scoring","Ultrasound-to-MRI translation boosts prostate cancer risk scoring","Synthetic MRI improves prostate risk prediction, but doctors can tell","AI-translated MRI better than ultrasound for prostate risk scoring","Synthetic MRI aids prostate risk scoring despite feature loss"],"cache_read_input_tokens":16000,"weakest_assumption_plain":"The quantitative comparison rests on the assumption that the ultrasound and MRI volumes are aligned well enough that a voxel or texture feature in one volume corresponds to the same tissue in the other. The paper reports that US and MRI were aligned by clinical collaborators and that identical masks were used for feature extraction, but it provides no registration-error analysis; if misalignment is substantial, the SSIM, radiomic correlations, and synthetic-versus-original comparisons all lose meaning.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic MRI from ultrasound improves prostate risk scoring","Ultrasound-to-MRI translation boosts prostate cancer risk scoring","Synthetic MRI improves prostate risk prediction, but doctors can tell","AI-translated MRI better than ultrasound for prostate risk scoring","Synthetic MRI aids prostate risk scoring despite feature loss"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000472,"raw_usage":{"total_tokens":2437,"prompt_tokens":1127,"completion_tokens":1310,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":743,"completion_tokens_details":{"reasoning_tokens":1230}},"tokens_in":743,"tokens_out":1310,"duration_ms":10109,"temperature":1.0,"reasoning_tokens":1230,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T00:37:21.417093+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute point-to-point registration error on the dataset's tracked biopsy cores or other landmarks, then re-run the radiomic rank correlations and SSIM comparisons excluding the worst-registered cases; if the correlations shift dramatically or the 76-preserved/93-lost split changes, the central comparison would not stand. A simpler visual check is also available: the paper's Figure 2 shows a synthetic hyperintensity with no counterpart in the original MRI and a true lesion not preserved—if coregistering those two images with a different algorithm makes the bright spot coincide with a real structure, that discrepancy is registration-driven rather than a synthesis failure.","supporting_citations":[{"cited_title":"Prostate MRI and Ultrasound With Pathology and Coordinates of Tracked Biopsy (Prostate-MRI-US-Biopsy) (version 2) [Data set],","cited_arxiv_id":null,"evidence_quote":"Supplies the 794-patient aligned ultrasound/MRI dataset with segmentation masks and biopsy-tracked risk scores; without it the translation and radiomics comparisons have no data."},{"cited_title":"A review on generative adversarial networks for image generation,","cited_arxiv_id":null,"evidence_quote":"Provides the review of GAN architectures that grounds the selection of the ten 2D/3D translation networks compared in the study."},{"cited_title":"Conditional diffusion models for semantic 3D brain MRI synthesis,","cited_arxiv_id":null,"evidence_quote":"Defines the conditional diffusion model (3D-Med-DDPM) that serves as one of the two diffusion baselines outranked by 2D-Pix2Pix."},{"cited_title":"Contourdiff: Unpaired image translation with contour -guided diffusion models,","cited_arxiv_id":null,"evidence_quote":"Defines the contour-guided diffusion model (2D-ContourDiff) used as the second diffusion baseline for the low-level feature comparison."},{"cited_title":"ViSERA: visualized & standardized environment for radiomics analysis -a shareable, executable, and reproducible workflow generator,","cited_arxiv_id":null,"evidence_quote":"Supplies the standardized radiomics feature generator used to extract the 186 features whose Spearman correlations form the paper's three-group feature-preservation analysis."},{"cited_title":"Improved prostate cancer diagnosis using a modified ResNet50-based deep learning architecture,","cited_arxiv_id":null,"evidence_quote":"Supplies the ResNet50 architecture used as the deep-learning classifier that the radiomics framework outperforms."},{"cited_title":"A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research,","cited_arxiv_id":null,"evidence_quote":"Provides the correlation-coefficient thresholds (poor, moderate, good, excellent) used to set the 0.50 cutoff that defines the three radiomic feature groups."}],"review_version":1}