{"id":"cd7e2f3b-c7e7-46da-9080-c9264d27af1e","arxiv_id":"1909.01182","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A multi-sequence U-Net trained with CycleGAN-synthesized LGE images and rotation-based scar augmentation achieves Dice scores of 0.90 (LV), 0.81 (MYO), and 0.87 (RV) on the MS-CMRSeg test set.","lead":"This paper describes a method to segment late gadolinium enhancement cardiac MRI by training a deep network with additional non-contrast MRI sequences, synthetic images, and scar-location augmentation. The authors report Dice scores above 0.80 on a 40-patient challenge test set while using only five labeled LGE volumes for training.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed improvement over LGE-only training rests on validation-set selection with a single labeled LGE patient and no test-set ablation.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but the most load-bearing concern is not primarily the anatomical fidelity of synthetic LGE images. Even if the synthetic images perfectly preserved cardiac boundaries, the paper would still not demonstrate the claimed improvement because no test-set ablation is reported and model selection is based on an extremely small validation set (one LGE patient). Conversely, the synthetic-fidelity issue, while real, is secondary: if the synthetic images were imperfect but the method still improved test Dice, the central claim would stand. The missing quantitative comparison directly undermines the central claim, so I focus on that. This aligns with part of the reader's rationale, but the reader's stated weakest assumption is different, hence partial agreement. A concrete test requiring all configurations to be evaluated on the test set would settle whether the improvement is genuine or an artifact of validation-set selection.","tokens_in":6291,"tokens_out":4027,"duration_ms":37181,"concrete_test":"Re-run all eight training configurations described in Section 3 and evaluate each on the same 40-patient test set; if test labels are unavailable, perform repeated cross-validation on the five labeled LGE volumes (e.g., leave-one-patient-out), reporting mean and per-patient Dice for LV, MYO, and RV for each configuration with paired bootstrap confidence intervals. If configuration 8 is not significantly better than configurations 1, 3, or 4 on unseen patients, the paper's central improvement claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that multi-sequence training with synthetic LGE and rotated-scar augmentation improves LGE-MRI segmentation over conventional training with real LGE data. The evidence in Section 3 consists of eight training configurations evaluated on a validation set; the authors state that configuration 8 'resulted in the best segmentations' and then evaluate only that model on the 40-patient test set (Table 3). No test-set Dice for configurations 1, 3, or 4 are reported, so the claimed improvement is not demonstrated on unseen data. Moreover, the validation set for LGE is 20% of the five labeled LGE volumes, i.e., a single patient; model selection on one volume cannot reliably rank configurations, and the apparent gains from synthetic images and augmentation may be noise. Table 2 compares a model trained on bSSFP versus one trained on synthetic LGE, but it is evaluated on the same five labeled LGE volumes, which are part of the training data and not an independent test. The paper itself acknowledges in the Conclusions that 'extensive validation will be performed to assess in detail the relative importance of the different steps,' implicitly confirming this gap. Thus the central improvement claim is unsupported as stated, even if the reported test Dice for configuration 8 are accurate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a deep-learning pipeline for segmenting the left ventricle, myocardium, and right ventricle in late gadolinium enhancement cardiac MRI (LGE-MRI), trained with only five labeled LGE volumes. To compensate for the small labeled dataset, the method augments the LGE training set in two ways: shape-guided rotations of the myocardium and LV to relocate scar tissue, and CycleGAN-based translation from bSSFP cine images into synthetic LGE images. A modified U-Net is then trained on eight configurations that combine real LGE, bSSFP, T2, synthetic LGE, and the scar-rotation augmentation. The authors select configuration 8 (all sequences plus synthetic LGE and rotations) based on validation performance and report its results on the 40-patient MS-CMRSeg test set, with average Dice scores of 0.898 (LV), 0.810 (MYO), and 0.866 (RV). The paper claims that multi-sequence training with synthetic images and augmentation improves LGE segmentation over conventional training with real LGE data only.","tokens_in":6527,"tokens_out":4137,"duration_ms":43573,"significance":"If the central claim holds, the paper would offer a practically valuable solution to the scarcity of annotated LGE-MRI data: it would show that non-contrast sequences and synthetic examples can substitute for large numbers of real LGE annotations. The use of an external challenge test set for the final model is a strength, as are the reported standard deviations and the comparison with a prior method trained on five times more LGE data. The work is also reproducible in style, as it identifies the public CycleGAN implementation and reports training details. However, the improvement claim is not currently established on unseen data, because test-set metrics are reported only for the selected configuration and the intermediate comparisons are evaluated on training volumes. The practical value therefore depends on additional experiments that the paper itself acknowledges are still needed.","major_comments":[{"comment":"The central claim that configuration 8 improves over conventional training is not supported by unseen-data evidence. Test-set Dice are reported only for the selected model (Table 3), and no test-set results are given for configurations 1, 2, 3, or 4, so the reader cannot verify the claimed improvement on the 40-patient test set. The only quantitative comparison between models trained with and without synthetic images (Table 2) is performed on the five labeled LGE volumes, which are part of the training data, not an independent test set. The paper's own conclusion in Section 4 states that 'extensive validation will be performed to assess in detail the relative importance of the different steps,' which confirms that the relative contribution of each component is not established in this manuscript. I would ask the authors to report test-set metrics for all eight configurations, or, if the challenge organizers cannot supply per-configuration test results, to provide a multi-fold cross-validation on the five labeled LGE patients and clearly separate training, validation, and test folds.","section":"Section 3, Table 3"},{"comment":"The model selection among the eight training configurations is based on a validation set consisting of 20% of the five labeled LGE volumes, i.e., a single patient. Ranking eight configurations on one volume cannot reliably distinguish genuine improvements from noise, especially when the selection considers three Dice metrics simultaneously. The reported gains from synthetic images and scar-rotation augmentation may therefore reflect overfitting to that single validation volume rather than a generalizable effect. Please provide either a leave-one-patient-out cross-validation over the five labeled LGE volumes or, preferably, test-set results for all eight configurations so that the selection is not made on a single patient.","section":"Section 2.3, Section 3"},{"comment":"The synthetic LGE images are evaluated only qualitatively (Figure 2) and indirectly through a segmentation experiment on the five labeled LGE volumes (Table 2). The paper assumes that the bSSFP ground-truth contours are valid supervision for the synthetic LGE images, but the anatomical fidelity of the synthetic images is not quantitatively verified. The reported Dice scores for the model trained with synthetic LGE (0.809 LV, 0.688 MYO, 0.820 RV) are not compared against an upper bound or against a model trained with real LGE images on a held-out set, so they do not establish that the synthetic images preserve cardiac boundaries well enough for LGE segmentation. Please add a quantitative measure of boundary fidelity (for example, boundary displacement between synthetic LGE and real LGE in matched slices, or a manual quality rating) and evaluate the synthetic-LGE ablation on unseen LGE data.","section":"Section 2.2, Table 2"}],"minor_comments":[{"comment":"The CycleGAN training description would benefit from specifying the image size, normalization, and preprocessing used for the generators, since these details affect the quality of the synthetic LGE images.","section":"Section 2.2"},{"comment":"The columns in Figure 4 are not labeled with the corresponding configuration numbers (3, 7, 8); adding explicit labels or extending the caption would make the qualitative comparison easier to follow.","section":"Figure 4"},{"comment":"Please clarify whether the scar-rotation augmentation is applied only to real LGE images or also to synthetic LGE images, since configuration 4 and configuration 8 differ in exactly which images receive the rotations.","section":"Section 2.2"},{"comment":"The sentence describing the reduction of filters after upsampling is a little ambiguous; it would be clearer to state that the number of filters in the final upsampling block is reduced to match the number of segmentation labels, as in the cited reference [16].","section":"Section 2.3"},{"comment":"The comparison with the results of Yue et al. [3] is informal; please indicate whether the Dice scores quoted for that method come from the same MS-CMRSeg test set and whether any statistical significance testing was performed.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a concise challenge contribution, and the missing ablation on the test set is likely due to space constraints, but it is load-bearing for the central claim. I would encourage the editor to request the additional experiments; if the challenge organizers can provide per-configuration test metrics, that would directly resolve the main concern. The informal comparison with prior work [3] should also be tightened before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a plausible MS-CMRSeg challenge entry, and the reported test Dice (0.898 LV, 0.810 MYO, 0.866 RV) are solid given only five labeled LGE volumes for training. But the paper's central claim—that adding synthetic LGE and rotated-scar augmentation improves segmentation—is not actually demonstrated on unseen data. The authors pick configuration 8 because it scores best on a validation set drawn from a single LGE patient, then report test results only for that model. That is a real gap, and the paper's own conclusion says extensive validation is still to come.\n\nWhat is new: the specific assembly of CycleGAN-based bSSFP-to-LGE synthesis plus landmark-guided rotation of the myocardium to redistribute scar locations. Each ingredient is established, but the combination for LGE-MRI segmentation is new, and the method description is clear enough to reproduce with effort.\n\nThe soft spots are in the evaluation. Table 2 compares a model trained on bSSFP versus one trained on synthetic LGE, but both are evaluated on the same five labeled LGE volumes that are part of the training data—so that comparison does not measure generalization. The eight-configuration ranking is done on 20% of those five volumes, essentially one patient, so the apparent gains from configurations 4 and 8 could easily be noise. No test-set Dice for configurations 1, 3, or 4 are reported, so we cannot tell whether the synthetic or augmentation steps actually generalize. The synthetic LGE fidelity is also only checked indirectly; the assumption that bSSFP contours are valid supervision for synthetic LGE is not verified quantitatively. These are not fatal to the pipeline, but they do make the improvement claim conditional.\n\nThe citation pattern is clean: standard references, no self-citation concern. The writing is clear and honest about limitations.\n\nThis paper is for readers working on cardiac MRI segmentation with limited annotated LGE data. They will find a useful template and a cautionary example of how easy it is to over-read validation-set rankings. It deserves a serious referee. I would accept it for peer review, but the revision should include test-set results for at least the LGE-only and the all-sequences-without-synthesis/augmentation variants, and model selection should use a larger validation set or cross-validation.","headline":"A credible challenge write-up with strong test Dice, but the central claim that synthetic and augmented training helps is selected on a one-patient validation set and never shown on the test set.","tokens_in":7035,"tokens_out":1672,"would_cite":false,"duration_ms":17341,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A U-Net trained on real multi-sequence MRI plus CycleGAN-synthesized LGE and rotated scar images reaches average Dice scores of 0.898 (LV), 0.810 (MYO), and 0.866 (RV) on the 40-patient test set, rivaling a network trained on five times…","keywords":["late gadolinium enhancement MRI","cardiac MRI segmentation","multi-sequence MRI","image synthesis","CycleGAN","data augmentation","U-Net","scar tissue"],"falsifier":"Measure, on the same patients, the contour displacement between synthetic LGE images and their real LGE counterparts at end-diastole; if the mean displacement approaches or exceeds the LGE slice thickness (5 mm), or if a model trained only on synthetic LGE fails to beat a model trained only on real LGE on a common held-out test set, the claim that synthetic LGE provides valid boundary information is refuted.","tokens_in":6139,"feed_emoji":"❤️","tokens_out":8688,"duration_ms":78485,"temperature":0.7,"pith_summary":"The paper asks whether scarce labeled late gadolinium enhancement (LGE) cardiac MRI can be stretched far enough to train a robust segmentation model. It answers yes, by feeding the segmenter not only LGE images but also the same patients' non-contrast bSSFP and T2 sequences, by synthesizing extra LGE images from bSSFP with a CycleGAN, and by rotating the myocardial wall within the LGE images so scar tissue appears at many locations. The best combined configuration outperforms single-sequence training and, with only five labeled LGE volumes, reaches average Dice scores of 0.898 (LV), 0.810 (MYO), and 0.866 (RV) on the 40-patient test set. The practical payoff would be accurate scar quantification in settings where LGE labels are scarce but cine and other routine sequences are abundant.","feed_headline":"Five labeled LGE volumes train a segmenter to Dice 0.90","feed_subtitle":"CycleGAN-made LGE images and rotated scar tissue let one network match models trained on five times more data.","key_machinery":"The argument rests on three components. The first is CycleGAN, an unpaired image-to-image translation network that converts bSSFP cine images into synthetic LGE images while preserving the underlying cardiac geometry, so the bSSFP ground-truth contours can supervise LGE segmentation. The second is shape-guided scar augmentation: 50 landmarks placed around the epicardium and endocardium allow the myocardial wall—and the scar within it—to be rotated in 20 steps of 7.2 degrees, spreading scar locations across up to 144 degrees. The third is a modified U-Net with deep supervision in the upsampling path and a reduced number of filters after each upsampling to match the label count; each sequence is fed as a separate single-channel input so one set of weights learns across modalities. These components together let the model exploit multi-sequence information and enlarged training sets without any registration step.","core_discovery":"The central claim is that LGE-MRI segmentation accuracy is improved by complementary information from non-contrast MRI sequences, provided through a single network that consumes each sequence as one input channel and does not require inter-sequence registration. The paper's best model—trained on all three sequences (LGE, bSSFP, T2), with synthetic LGE images generated by CycleGAN from bSSFP, and with shape-guided rotations of the myocardium that move scar locations around the wall—achieves average Dice scores of 0.898 (LV), 0.810 (MYO), and 0.866 (RV) on the 40-case challenge test set. In the authors' comparison, this is on par with a recent deep learning method trained on 25 labeled LGE volumes, five times more than the five used here. The authors also report that a model trained on synthetic LGE alone (Dice 0.809 for LV on the five labeled LGE volumes) far exceeds a model trained on bSSFP alone (0.503), which they take as evidence that the synthesized images carry useful information for LGE segmentation.","pith_inferences":["If synthetic LGE inherits bSSFP geometry, the same unpaired-translation recipe could extend to other contrast-enhanced modalities with scarce labels, such as delayed-enhancement CT, whenever a non-contrast sequence from the same anatomy is available.","The authors leave open whether the gain comes from faithful transfer of boundaries or from extra texture variation; a quantitative synthetic-to-real contour displacement measurement would settle which mechanism is doing the work.","A direct comparison of the rigid scar rotation used here against the proposed elastic deformations would test whether scar-location diversity, rather than global shape change, is the active ingredient of the augmentation gain."],"forward_implications":["A segmentation network trained on all three sequences plus synthetic LGE and scar-rotation augmentation outperforms, on the challenge validation set, every other configuration tested, including models trained on real LGE alone or on real multi-sequence data without synthesis.","With only five labeled LGE volumes, the method attains Dice scores similar to a recent deep learning LGE segmentation system trained on 25 labeled volumes, suggesting that synthesis and augmentation can substitute for a substantial amount of manual annotation.","Synthetic LGE images generated from bSSFP carry enough boundary information that a model trained on them alone segments the five labeled LGE volumes with a LV Dice of 0.809, well above the 0.503 obtained by training on the original bSSFP images.","Because the proposed pipeline does not register the sequences, it can be applied to multi-sequence cardiac data with differently aligned slices and consistently different boundary shapes.","Rotating the myocardial wall and scar within the LGE images reduces the risk of overfitting to a small number of scar locations, and in the qualitative examples it corrects segmentation errors that remain after adding synthetic images alone."],"supporting_citations":[{"why":"Provides the CycleGAN method for unpaired image-to-image translation used to generate synthetic LGE images from bSSFP.","marker":"[13]"},{"why":"Supplies the U-Net architecture that serves as the base of the CNN segmentation model.","marker":"[14]"},{"why":"Reports a recent deep learning LGE segmentation method trained on 25 labeled LGE volumes, the comparison point for the paper's Dice results.","marker":"[3]"},{"why":"Proposes the deep supervision term in the upsampling path that conditions the final segmentation predictions.","marker":"[15]"},{"why":"Proposes reducing the number of filters after each upsampling operation to match the number of labels, a design choice adopted in the network.","marker":"[16]"},{"why":"Describes an earlier multi-sequence atlas-based segmentation approach combining bSSFP, LGE, and T2, motivating the use of complementary sequences.","marker":"[6]"}],"fun_headline_variants":["Multi-sequence training with synthetic LGE matches models trained on 5x data","Synthetic LGE images from CycleGAN improve cardiac segmentation","Five labeled LGE volumes plus synthetic data match 25-volume training","Multi-sequence model with synthetic LGE rivals 5x more training data","LGE segmentation improved by synthetic images from cine MRI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the CycleGAN-generated synthetic LGE images keep the cardiac boundaries in the same place as the bSSFP images they were made from, so the bSSFP ground-truth contours are valid supervision for LGE segmentation; the paper supports this with qualitative examples and an indirect segmentation test rather than a direct measurement of boundary fidelity.","fun_headline_variants_meta":{"raw":{"variants":["Multi-sequence training with synthetic LGE matches models trained on 5x data","Synthetic LGE images from CycleGAN improve cardiac segmentation","Five labeled LGE volumes plus synthetic data match 25-volume training","Multi-sequence model with synthetic LGE rivals 5x more training data","LGE segmentation improved by synthetic images from cine MRI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000609,"raw_usage":{"total_tokens":2857,"prompt_tokens":987,"completion_tokens":1870,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":1781}},"tokens_in":603,"tokens_out":1870,"duration_ms":14369,"temperature":1.0,"reasoning_tokens":1781,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:24:48.472840+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, on the same patients, the contour displacement between synthetic LGE images and their real LGE counterparts at end-diastole; if the mean displacement approaches or exceeds the LGE slice thickness (5 mm), or if a model trained only on synthetic LGE fails to beat a model trained only on real LGE on a common held-out test set, the claim that synthetic LGE provides valid boundary information is refuted.","supporting_citations":[{"cited_title":"In: Proceedings of the IEEE interna- tional conference on computer vision","cited_arxiv_id":null,"evidence_quote":"Provides the CycleGAN method for unpaired image-to-image translation used to generate synthetic LGE images from bSSFP."},{"cited_title":"Cardiac Segmentation from LGE MRI Using Deep Neural Network Incorporating Shape and Spatial Priors","cited_arxiv_id":"1906.07347","evidence_quote":"Reports a recent deep learning LGE segmentation method trained on 25 labeled LGE volumes, the comparison point for the paper's Dice results."},{"cited_title":"In: International workshop on statistical atlases and computational models of the heart","cited_arxiv_id":null,"evidence_quote":"Proposes the deep supervision term in the upsampling path that conditions the final segmentation predictions."},{"cited_title":"In: Inter- national Workshop on Statistical Atlases and Computational Models of the Heart","cited_arxiv_id":null,"evidence_quote":"Proposes reducing the number of filters after each upsampling operation to match the number of labels, a design choice adopted in the network."},{"cited_title":"IEEE transactions on pattern analysis and machine intelli- gence (2018)","cited_arxiv_id":null,"evidence_quote":"Describes an earlier multi-sequence atlas-based segmentation approach combining bSSFP, LGE, and T2, motivating the use of complementary sequences."}],"review_version":1}