{"id":"1e6f8ed1-56b0-4a19-9fc8-10044ac3c5ad","arxiv_id":"1908.07726","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fine-tuning a CNN pretrained on T2 and bSSFP cardiac MRI on only five LGE-MR subjects yields 84.5% average Dice on LGE segmentation, far better than training from scratch.","lead":"A deep learning model trained on T2-weighted and bSSFP cardiac MRI, then fine-tuned on just a handful of LGE-MR images, segments heart chambers with about 85% Dice accuracy on a 40-patient test set. The result suggests that transfer learning can cut the need for manual annotations when a new MRI protocol is introduced.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation-set model selection and a 5-subject estimate of the from-scratch baseline weaken the reported comparison in Table 1; the headline contrast needs error bars or a fixed-protocol reproduction.","rationale":"The reader's weakest-assumption identifies feature transfer as the premise on which the claimed accuracy depends. I agree that transfer learning from T2/bSSFP to LGE is the conceptual crux, and that the four-to-five-sample fine-tuning regime is where that premise is least protected. The manuscript's own internal evidence is the main support for the claim, so the correctness risk is concentrated in the validation-set comparison of Section 3 and Table 1. My stress-test adds a more concrete procedural framing: the ambiguity between four and five LGE training samples, the unspecified fold-selection for the reported validation numbers, and the absence of error bars or significance testing for the from-scratch baseline mean that the 'significantly outperformed' statement is not yet quantitatively established. The 40-subject test results in Table 2 are a real strength: they show stable performance with standard deviations and multiple metrics, and they support the central finding of feasible LGE segmentation with very few annotated samples, provided the training protocol is clarified. The lack of a prior-art baseline is a fair criticism but not a soundness flaw; it is addressable. I do not see an internal inconsistency that would force a reject. The appropriate verdict is therefore conditional acceptance, with the concrete checks being (a) a fixed-protocol reproduction with per-fold error bars, and (b) clarification of whether Table 1 reports averaged folds or a selected fold. This matches the reader's conditional verdict and agrees that the weakest assumption concerns the transfer premise in the low-sample regime.","tokens_in":6858,"tokens_out":1689,"duration_ms":14465,"concrete_test":"Re-run the experiment with a fixed protocol: train from-scratch and adapted models on the same four LGE training subjects, select both models by the same rule (e.g., best epoch on a held-out validation subject), repeat across all five folds, and report mean plus per-fold standard deviation for every structure in Table 1. If the adapted-minus-from-scratch gap remains above 10 Dice points in all folds, the central claim survives the concern; if the gap collapses or reverses in some folds, the headline comparison is not robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's strongest claim is that fine-tuning from T2+bSSFP with four or five LGE samples significantly outperforms training from scratch on the same LGE samples, with Table 1 reporting 80.7% vs 66.9% average Dice on the validation set. This comparison rests on two fragile procedural details. First, Section 2.4 says LGE fine-tuning used five subjects with 5-fold cross-validation, but the abstract and Section 3 say the adapted network was trained with just four LGE samples, and the exact split or fold used for Table 1 is not specified; if Table 1's adapted row came from the fold whose validation subject was easiest, or from one selected fold rather than averaged folds, the 13.8-point gap could be inflated. Second, the from-scratch baseline is reported as a single average over five subjects with no per-structure standard deviation or significance test; the conclusion 'significantly outperformed' in the abstract is not supported by any statistical evidence. A further gap is that no prior multi-sequence method (e.g., the MvMM baseline cited as [6]) is compared on the same test set, so 84.4% average Dice on 40 test subjects is not contextualized against the state of the art. The central mechanism—feature transfer from T2/bSSFP to LGE—is plausible, but the quantitative evidence for its advantage over from-scratch training on four or five samples is not yet secured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a supervised domain adaptation approach for late gadolinium-enhanced (LGE) cardiac MRI segmentation. An encoder-decoder network is first trained on T2-weighted and bSSFP sequences with pixel-level annotations, then fine-tuned on a small number of LGE-MR subjects. The authors evaluate on the MS-CMRSeg 2019 challenge dataset and report validation Dice scores of 80.7% for the adapted network with augmentation versus 66.9% for a from-scratch baseline and 31.3% for a model trained only on T2/bSSFP. On the 40-subject test set they report an average Dice of 84.4%. The central claim is that feature transfer from T2/bSSFP to LGE with only a few target samples significantly outperforms training from scratch on the same LGE data.","tokens_in":7139,"tokens_out":3450,"duration_ms":34637,"significance":"If the reported results are robust, the method is practically relevant because it addresses the high cost of annotating LGE-MR images: the proposed transfer-learning recipe is simple, uses public challenge data, and could reduce the number of required target-domain annotations. The paper also clearly demonstrates the domain-shift problem by showing that a T2/bSSFP-only model performs poorly on LGE images. However, the significance is limited by the lack of statistical validation of the headline comparison, an unresolved inconsistency in the number of LGE training samples, and the absence of any comparison with prior multi-sequence segmentation methods on the same test set.","major_comments":[{"comment":"The abstract states that the domain-adapted network was trained with just four LGE-MR training samples, while Section 2.4 says the model was re-trained with five LGE subjects using 5-fold cross-validation. This is a load-bearing discrepancy because the paper's contribution is specifically that very few target-domain samples are sufficient. The authors must state exactly how many LGE subjects were used, how the folds were constructed, and which fold or aggregation produced the numbers in Table 1.","section":"Abstract and Section 2.4"},{"comment":"The validation comparison in Table 1 reports single average Dice values over five subjects (0.669 for the from-scratch baseline and 0.807 for the proposed method) with no per-subject results, no standard deviation, and no significance test. The abstract's claim that the proposed method 'significantly outperformed' the baseline is therefore not supported by the evidence presented. The authors should report per-fold results, confidence intervals, and a paired statistical test (for example Wilcoxon signed-rank or a bootstrap interval) on the five validation subjects.","section":"Table 1"},{"comment":"The 40-subject test evaluation in Table 2 reports results only for the domain-adapted method. Without a from-scratch baseline or a published multi-sequence method (e.g., MvMM, reference [6]) evaluated on the same test set, the central claim that domain adaptation outperforms no adaptation cannot be assessed on the test data. The 84.4% Dice value is useful, but it does not by itself establish the advantage of the adaptation mechanism. The authors should provide a test-set comparison, or clearly state the challenge ranking and report all compared methods under the same protocol.","section":"Section 3, Table 2"}],"minor_comments":[{"comment":"The keyword 'Myocardial Infraction' should be 'Myocardial Infarction'.","section":"Keywords"},{"comment":"The phrase 'the test set comprises of 40 LGE-MR subjects' is grammatically incorrect; 'comprises' should be used without 'of'.","section":"Abstract"},{"comment":"The definition 'domain Dtb = {S, P(X)}' is incomplete: the notation is introduced informally and the label space Y is mentioned before the domain is properly defined. Please tighten the notation for clarity.","section":"Section 2.1"},{"comment":"The text says 'Since T1 and bSSFP images have very few slices', but the paper only uses T2, bSSFP, and LGE sequences; this should read 'T2 and bSSFP'.","section":"Section 2.3"},{"comment":"The text says 'the average Dice improved to 80.9%' while Table 1 reports 0.807 for the same row; use a consistent rounding convention.","section":"Section 3 and Table 1"},{"comment":"The caption contains the typo 'withe color', which should be 'white color'.","section":"Figure 3 caption"},{"comment":"The soft Dice loss in Eq. (1) omits the conventional factor of 2 in the numerator, so a perfectly overlapping prediction gives a loss of 0.5 rather than 0. Please clarify whether this is intentional or a typographical error.","section":"Equation (1)"}],"recommendation":"major_revision","confidential_remarks":"The paper's main weakness is reporting rigor rather than novelty: the 4-versus-5 LGE sample inconsistency and the missing significance testing are fixable in revision, but they are central to the claimed contribution. I would not recommend rejection if the authors can clarify the protocol and substantiate the 'significantly outperformed' claim with error bars or a paired test."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent but incremental application of vanilla fine-tuning to LGE cardiac MRI segmentation, and the headline comparison is weaker than the prose suggests. The genuinely useful piece is Table 2, the 40-subject test-set performance of the domain-adapted model (avg Dice 0.844 ± 0.063). The novel part is not the method—supervised domain adaptation by fine-tuning is familiar and the authors cite the relevant literature—but the specific multi-sequence setting and the observation that T2+bSSFP pretraining gives a strong initialization for LGE with few labels. That is a legitimate contribution for the MS-CMRSeg challenge track and likely of practical interest to people working on CMR segmentation.\n\nThe paper does several things well. The architecture is clearly described, and the preprocessing and augmentation details are sufficient to approximate the pipeline. Reporting test-set Dice with standard deviations is good practice, and the comparison with a source-only model (31.3% Dice) demonstrates a real domain shift. The claim that fine-tuning is better than from-scratch training with few samples is plausible.\n\nWhere it gets soft: the abstract says 'trained with just four LGE-MR training samples,' while Section 2.4 says five LGE subjects with 5-fold cross-validation and Section 3 repeats five. That inconsistency matters because the validation comparison in Table 1 is the only evidence for the 'significantly outperformed' claim. The from-scratch baseline is a single average over five subjects with no per-structure variance; given that the model is selected on validation, fold choice could materially change the 13.8-point gap. The test-set numbers in Table 2 are more robust, but without a from-scratch baseline on the same 40 subjects or a comparison with an existing multi-sequence method (e.g., the MvMM baseline cited as [6]), the advantage over alternatives is not established. 'Significantly outperformed' is not supported by any statistical test. These are addressable, and I don't see a load-bearing flaw in the idea itself—fine-tuning from a related source domain is a reasonable way to handle scarce LGE labels.\n\nWho this is for: researchers working on cardiac MRI segmentation, especially in low-label settings. It deserves referee time, but a serious referee should ask for a fixed-protocol from-scratch baseline with error bars and a comparison against prior multi-sequence methods. I would not desk-reject it; I would send it to review with expectations for revision.","headline":"A competent, incremental fine-tuning paper whose 40-subject test-set result is probably real, but whose headline 'significantly outperformed' claim rests on a five-subject validation comparison with no error bars and an inconsistency on training sample count.","tokens_in":776,"tokens_out":772,"would_cite":true,"duration_ms":32469,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning on just four late-gadolinium-enhanced cardiac MR scans lets a segmentation network pre-trained on T2 and bSSFP images reach 84.5% average Dice on 40 test subjects.","keywords":["cardiac MRI segmentation","supervised domain adaptation","transfer learning","late gadolinium enhancement","left ventricle segmentation","fully convolutional network","multi-sequence MRI","Dice loss"],"falsifier":"Run the same fine-tuning procedure on an external LGE-MR dataset acquired with a different scanner or field strength, using the same T2+bSSFP pre-trained weights and four annotated subjects for adaptation; if average Dice falls to the source-only level of about 31% rather than remaining near 85%, the adaptation does not generalize to new LGE domains.","tokens_in":6658,"feed_emoji":"🫀","tokens_out":9197,"duration_ms":82081,"temperature":0.7,"pith_summary":"The paper tries to show that a cardiac MRI segmentation network can be adapted to a new imaging sequence with very few annotated examples. It first trains an encoder-decoder CNN on T2-weighted and balanced steady-state free precession images, then fine-tunes the same network on only four or five late gadolinium-enhanced (LGE) MR images. On a 40-subject LGE test set, the adapted network reaches an average Dice score of about 0.85, and on the validation set it clearly beats the same network trained from scratch on the few LGE samples (about 81% versus 67% average Dice). The point is that pixel-level annotations in one MR sequence can be leveraged to segment another sequence, reducing the annotation burden for LGE images.","feed_headline":"Four heart scans adapt a cardiac MRI segmenter to 85% overlap","feed_subtitle":"Pre-training on T2 and bSSFP images lets a CNN segment late gadolinium-enhanced MRI with almost no target labels.","key_machinery":"The load-bearing mechanism is two-stage transfer learning, a supervised domain adaptation scheme in which a small number of labeled target-domain (LGE) samples guide adaptation of a model trained on source-domain (T2+bSSFP) images. First, an encoder-decoder fully convolutional network with residual connections, skip connections, and a dilated-convolution bottleneck is trained on T2+bSSFP images using a multi-class soft Dice loss. Second, the same architecture is initialized with those learned weights and fine-tuned, with identical hyperparameters, on LGE-MR images. Data augmentation during fine-tuning further lifts validation Dice from 76.6% to 80.7%.","core_discovery":"The central claim is that supervised domain adaptation by weight transfer across cardiac MR sequences makes LGE-MR segmentation feasible with very scarce target labels. A fully convolutional encoder-decoder trained on T2+bSSFP images with pixel-wise labels is initialized with those weights and then fine-tuned on just a handful of LGE-MR subjects. On a 40-subject LGE test set, the fine-tuned model obtains an average Dice of 0.844 ± 0.063 (myocardium 0.788, LV 0.912, RV 0.832). In contrast, a model trained from scratch on the same few LGE samples achieves 66.9% average Dice on the validation set, while the un-adapted source model essentially fails on LGE images with 31.3% average Dice. The paper attributes the improvement to better weight initialization and domain-invariant features learned from the source sequences.","pith_inferences":["Beyond the paper, the same fine-tuning recipe could be applied to other scarce cardiac sequences, such as T1 mapping or edema-weighted images, whenever an abundant annotated sequence exists; the paper does not test this.","A T2-only or bSSFP-only pre-training ablation, which the paper does not report, would reveal whether combining both source sequences is necessary or whether any single annotated sequence suffices.","Because the network is 2D and slice thickness differs across sequences (5 mm for LGE versus 8–20 mm for source sequences), the transfer may rely on in-plane texture rather than volumetric anatomy; a slice-consistency or 3D evaluation would test this.","The fine-tuning set contains only five subjects with 5-fold cross-validation, so the paper does not measure how much performance varies with the choice of adaptation subjects; repeated sampling from a larger label pool would quantify that variance."],"forward_implications":["A practical path to LGE-MR segmentation with a handful of labeled subjects is to pre-train on annotated T2/bSSFP images and then fine-tune, reaching 0.844 average Dice on 40 held-out subjects.","Without fine-tuning, source-sequence knowledge does not transfer by itself: the un-adapted model scores only 0.313 average Dice on LGE images.","Training from scratch on the same few LGE labels is much weaker, with 0.669 validation Dice, so the method's value lies in weight initialization from multi-sequence data rather than in the architecture alone.","Data augmentation during fine-tuning adds roughly four Dice points on the validation set, from 0.766 to 0.807, indicating that aggressive augmentation helps in the small-target-label setting."],"supporting_citations":[{"why":"Supplies the multi-sequence CMR dataset and the clinical multi-sequence segmentation task that the method is validated on.","marker":"[6]"},{"why":"Defines the supervised domain adaptation setting that the two-stage training procedure instantiates.","marker":"[10]"},{"why":"Shows transfer learning by fine-tuning for MRI domain adaptation, the direct precedent for adapting features to LGE images.","marker":"[12]"},{"why":"Provides the encoder-decoder architecture with dilated convolutions that the paper adapts for cardiac segmentation.","marker":"[13]"},{"why":"Supplies the residual connection design used in each encoder block to improve gradient flow.","marker":"[14]"},{"why":"Supplies the dilated-convolution context aggregation used in the bottleneck to enlarge the receptive field.","marker":"[15]"},{"why":"Provides the multi-class soft Dice loss used to train the network under class imbalance.","marker":"[16]"}],"fun_headline_variants":["Just 4 LGE scans adapt cardiac MRI segmentation to 85% Dice","Transfer learning from T2+bSSFP makes LGE segmentation work with only 4 labels","Four target labels unlock 85% Dice for cardiac MRI segmentation","Supervised domain adaptation: 4 LGE images yield 85% Dice","T2+bSSFP pre-training: 4 LGE scans beat from-scratch by 18 points"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the assumption that anatomical features learned from T2 and bSSFP images transfer to LGE images through fine-tuning with only four or five annotated LGE subjects, despite large differences in contrast, resolution, slice thickness, and acquisition protocol.","fun_headline_variants_meta":{"raw":{"variants":["Just 4 LGE scans adapt cardiac MRI segmentation to 85% Dice","Transfer learning from T2+bSSFP makes LGE segmentation work with only 4 labels","Four target labels unlock 85% Dice for cardiac MRI segmentation","Supervised domain adaptation: 4 LGE images yield 85% Dice","T2+bSSFP pre-training: 4 LGE scans beat from-scratch by 18 points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000839,"raw_usage":{"total_tokens":3681,"prompt_tokens":991,"completion_tokens":2690,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":2582}},"tokens_in":607,"tokens_out":2690,"duration_ms":113669,"temperature":1.0,"reasoning_tokens":2582,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:57:36.214751+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same fine-tuning procedure on an external LGE-MR dataset acquired with a different scanner or field strength, using the same T2+bSSFP pre-trained weights and four annotated subjects for adaptation; if average Dice falls to the source-only level of about 31% rather than remaining near 85%, the adaptation does not generalize to new LGE domains.","supporting_citations":[{"cited_title":"In: MICCAI","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-sequence CMR dataset and the clinical multi-sequence segmentation task that the method is validated on."},{"cited_title":"In: 2017 IEEE International Conference on Computer Vision (ICCV)","cited_arxiv_id":null,"evidence_quote":"Defines the supervised domain adaptation setting that the two-stage training procedure instantiates."},{"cited_title":"In: MICCAI 2017","cited_arxiv_id":null,"evidence_quote":"Shows transfer learning by fine-tuning for MRI domain adaptation, the direct precedent for adapting features to LGE images."},{"cited_title":"In: STACOM","cited_arxiv_id":null,"evidence_quote":"Provides the encoder-decoder architecture with dilated convolutions that the paper adapts for cardiac segmentation."},{"cited_title":"In: CVPR","cited_arxiv_id":null,"evidence_quote":"Supplies the residual connection design used in each encoder block to improve gradient flow."},{"cited_title":"In: ICLR","cited_arxiv_id":null,"evidence_quote":"Supplies the dilated-convolution context aggregation used in the bottleneck to enlarge the receptive field."},{"cited_title":"In: 2016 Fourth International Conference on 3D Vision (3DV)","cited_arxiv_id":null,"evidence_quote":"Provides the multi-class soft Dice loss used to train the network under class imbalance."}],"review_version":1}