{"id":"9771aa51-ed43-41a4-92b9-78ebe6505009","arxiv_id":"2505.06527","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Adding SAM to uniGradICON training yields small Dice gains on five medical registration test sets, but the main baseline comparison is not numerically reported.","lead":"This paper applies Sharpness-Aware Minimization (SAM), a training trick that seeks flat loss minima, to a medical image registration foundation model called uniGradICON. The modified model reports slightly higher organ overlap scores on several test datasets, but the key comparison against the plain foundation model is only shown as a boxplot without numbers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central generalization claim rests on a boxplot with no reported numbers, no significance test, and no stated evidence that the uniGradICON baseline was retrained under identical conditions with SAM as the only change.","rationale":"The reader's conditional verdict is well aligned with the evidence. The central claim is empirical: SAM improves cross-dataset registration accuracy of uniGradICON. For that claim to hold, the comparison in Fig. 5 must be a controlled experiment in which the only difference is the optimizer. The paper does not describe the baseline's training provenance, and the direct comparison is presented only as a boxplot. The main quantitative table omits uniGradICON entirely, and the margins over the included state-of-the-art methods are small relative to the reported standard deviations. No significance testing is reported anywhere. The loss-landscape figure is suggestive but not quantitative. These are not internal inconsistencies in the method; they are gaps in evidence for an empirical claim. The appropriate status is conditional acceptance pending a controlled, fully reported head-to-head comparison. Since the reader already arrived at CONDITIONAL with the same weakest assumption, no verdict change is needed.","tokens_in":12659,"tokens_out":3469,"duration_ms":35032,"concrete_test":"Retrain uniGradICON from scratch under the exact protocol in Section IV-B (same composite dataset, N=1000 pairs per dataset per cycle, 800+200 epochs, learning rate 5e-5, lambda=1.5, same random seeds) with SAM disabled, and evaluate both that baseline and the SAM-trained model on the five datasets in Fig. 5. Report mean plus/minus standard deviation of Dice per structure and per dataset, along with a paired significance test across subjects. If the average Dice difference is below 0.01 or is not statistically significant after multiple-comparison correction, the claim of 'significantly higher' fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim, in Section IV-D, is that 'our method consistently achieves significantly higher Dice coefficients on all five datasets.' The evidence for this claim is Fig. 5, a boxplot that reports neither means, standard deviations, per-dataset sample sizes, nor any statistical significance test. The word 'significantly' is therefore unsupported by any quantitative statistic in the only direct head-to-head comparison with uniGradICON. In the one quantitative table (Table II), the primary baseline uniGradICON is absent: the comparison is against Elastix, VoxelMorph, DiffuseMorph, and TransMorph, and the reported margins are tiny (0.7360 vs 0.7322 over TransMorph on ACDC; 0.8722±0.0398 vs 0.8680±0.0362 over VoxelMorph on SLIVER, a gap of about 0.1 pooled standard deviation). Section IV-D does not state whether the uniGradICON baseline was retrained from scratch on the identical composite dataset, with the same 800/200 epochs, loss weights, random seeds, and sampling protocol, keeping SAM as the only difference. If the Fig. 5 uniGradICON results came from a released checkpoint or a different training schedule, the claimed generalization gain could be caused by training details rather than by SAM. The loss-landscape visualization in Fig. 6 is qualitative and does not quantify flatness, so it does not independently establish the mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes incorporating Sharpness-Aware Minimization (SAM) into the uniGradICON medical image registration foundation model. The method replaces the base optimizer's gradient with one computed at an adversarial perturbation, supposedly leading to flatter minima and improved cross-dataset generalization. The authors train on a composite dataset of four sources (COPDGene, OAI, HCP, L2R-Abdomen) and evaluate on five test sets (ACDC, SLIVER, and three Learn2Reg tasks). The central claim, stated in Section IV-D, is that adding SAM to uniGradICON 'consistently achieves significantly higher Dice coefficients on all five datasets,' supported by a boxplot (Fig. 5) and a qualitative loss-landscape visualization (Fig. 6). Quantitative comparisons against other methods on ACDC and SLIVER are given in Table II.","tokens_in":12976,"tokens_out":1790,"duration_ms":17879,"significance":"If substantiated, the result would be of practical interest: a drop-in modification (SAM) to an existing foundation model that improves cross-dataset generalization with no architectural change is a useful contribution. The paper also makes the method reproducible by releasing code. However, the core evidence for the central claim is currently descriptive rather than statistical. The reported margins over the best runner-up in Table II are small (0.0038 on ACDC, 0.0042 on SLIVER), and the primary baseline uniGradICON is absent from that table. The paper's strength is its clear framing of a plausible hypothesis (flat minima improve registration generalization) and the use of established components (GradICON/uniGradICON, SAM). The weakness is that the direct head-to-head comparison is not quantitatively reported, so the central claim is not yet established.","major_comments":[{"comment":"The claim that 'our method consistently achieves significantly higher Dice coefficients on all five datasets' is supported only by a boxplot. The figure reports no means, standard deviations, per-dataset sample sizes, or statistical tests. The word 'significantly' is therefore unsupported. Please provide the underlying numerical results (e.g., mean ± std for each dataset and each method) and, where appropriate, a paired significance test or confidence intervals, to substantiate the claim.","section":"Section IV-D, Fig. 5"},{"comment":"The primary baseline uniGradICON is missing from Table II, the only quantitative comparison table. Without numbers for uniGradICON on ACDC and SLIVER, the reader cannot assess whether SAM improves over its base model in the two datasets that do have quantitative results. Please add uniGradICON results to Table II, or clearly state why they are omitted.","section":"Section IV-C, Table II"},{"comment":"The paper does not state whether the uniGradICON baseline in Fig. 5 was retrained from scratch on the identical composite dataset with the same 800/200 epochs, loss weights (λ=1.5), learning rate (η=5e-5), random seeds, and sampling protocol, with SAM as the only difference. If the baseline numbers come from a published checkpoint or a different training procedure, the claimed generalization gain could be caused by training details rather than by SAM. Please specify the exact training setup for both models.","section":"Section IV-D, Fig. 5"},{"comment":"The loss-landscape visualization is qualitative and does not quantify flatness (e.g., via the trace of the Hessian, the largest eigenvalue, or a sharpness measure). As presented, Fig. 6 is suggestive but does not independently establish the mechanism of improved generalization. Please either add a quantitative flatness measure or temper the mechanistic claim.","section":"Section IV-D, Fig. 6"}],"minor_comments":[{"comment":"The abstract contains a malformed URL: 'https://github.com/Promise13/fm_sam}{https://github.com/Promise13/fm\\_sam'. Please correct the formatting.","section":"Abstract"},{"comment":"The section heading 'Taining loss' should be 'Training loss'.","section":"Section III-B"},{"comment":"The phrase 'this method has effectively improved model performance' is vague; please specify the tasks and the magnitude of improvements reported in prior work.","section":"Section II-D"},{"comment":"The 'SLIVER' dataset is listed as MRI, but the cited SLIVER dataset (Heimann et al., 2009) is a CT dataset. Please verify the modality label.","section":"Section IV-A, Table I"},{"comment":"The boxplot in Fig. 3 is described in the text, but no caption appears in the manuscript body. Please add a complete caption for Fig. 3.","section":"Section IV-C"}],"recommendation":"major_revision","confidential_remarks":"The central claim is plausible and the experimental design is reasonable, but the evidence as presented is insufficient to support the strong wording 'significantly higher' and the specific comparison with uniGradICON is under-reported. I would ask the authors to add quantitative numbers, statistical tests, and a clear statement of the baseline training protocol. If the baseline cannot be retrained identically, the claim should be scaled back. This is a fixable issue, hence major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper applies Sharpness-Aware Minimization to the uniGradICON registration foundation model and reports better cross-dataset Dice. That is a plausible and cheap idea, and the specific combination is new. But the evidence for the main claim is under-powered: the only direct head-to-head with uniGradICON is a boxplot with no numbers or significance test, and the quantitative table omits uniGradICON entirely.\n\nWhat is actually new: SAM is an existing optimizer, uniGradICON is an existing model, and the paper says so plainly. The contribution is the application of SAM to a medical registration foundation model, which the cited literature doesn't cover. The method section correctly reproduces SAM's equations and uses the GradICON loss with the inverse-consistency regularizer. The authors evaluate on five datasets spanning heart, liver, abdomen, hippocampus, and lung, which is reasonable for a generalization claim.\n\nThe soft spots are real but fixable. Section IV-D states the method 'consistently achieves significantly higher Dice coefficients on all five datasets,' but Fig. 5 is a boxplot reporting no means, no standard deviations, and no statistical test. The word 'significantly' is doing work that no statistic does. In Table II, the primary baseline — uniGradICON — is missing; the comparison is against Elastix, VoxelMorph, DiffuseMorph, and TransMorph, with margins of 0.0038 on ACDC and 0.0042 on SLIVER over the runner-up. On SLIVER that's about 0.1 pooled standard deviation, so the practical difference may be small. The paper also never states whether the uniGradICON baseline was retrained from scratch on the identical composite dataset, with the same epochs, seeds, and sampling protocol, keeping SAM as the only change. If the baseline numbers came from a released checkpoint, the gain could be training details rather than SAM. The loss landscape visualization in Fig. 6 is qualitative and doesn't quantify flatness.\n\nNone of this is fatal. The direction is consistent with SAM's known behavior, and the missing numbers are easy to add. I'd send this to a serious referee, but the referee should ask for a proper head-to-head: Table II with uniGradICON included, numbers and error bars for Fig. 5, significance tests, and a clear statement of the baseline training protocol. The paper also needs to fix the malformed code link in the abstract.\n\nFor whom: readers interested in cheap ways to improve registration foundation models, and reviewers who want a concrete test of whether SAM transfers to this domain. It's a modest paper, not a breakthrough, but it's not a waste of time.","headline":"Plausible SAM application to a registration foundation model, but the central 'significant improvement' claim rests on a boxplot with no statistics and a table that omits the primary baseline.","tokens_in":13514,"tokens_out":2538,"would_cite":false,"duration_ms":22975,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training the uniGradICON medical image registration foundation model with Sharpness-Aware Minimization improves cross-dataset registration accuracy.","keywords":["Sharpness-Aware Minimization (SAM)","medical image registration","foundation model","generalization","cross-dataset","deformable registration","uniGradICON","loss landscape"],"falsifier":"Retrain the baseline uniGradICON model from scratch on the same composite training dataset, using the same random seed, number of epochs, learning rate, and loss weights, with SAM as the only added component, and measure Dice overlap on the same five test datasets. If the Dice differences between the two models disappear or become statistically insignificant, the paper's central claim would be falsified.","tokens_in":12428,"feed_emoji":"","tokens_out":13365,"duration_ms":107119,"temperature":0.7,"pith_summary":"This paper addresses a known weakness of medical image registration foundation models: they lose accuracy when applied to datasets, anatomies, or modalities they were not trained on. The authors propose training the uniGradICON foundation model with Sharpness-Aware Minimization (SAM), an optimizer that steers the model toward flat regions of the loss landscape, as a way to improve cross-dataset generalization. Their experiments on five held-out datasets — cardiac, liver, abdominal MR-CT, hippocampus, and lung CT — show that the SAM-trained model achieves consistently higher Dice overlap scores than the same model without SAM. The paper argues that this flat-minima training makes the model more robust across different anatomical structures and imaging conditions, and it visualizes a flatter loss landscape for the SAM-trained model in support.","feed_headline":"SAM training lifts medical registration Dice on five datasets","feed_subtitle":"The base registration model trained with SAM beats its original on every one of five cross-dataset tests.","key_machinery":"Sharpness-Aware Minimization (SAM) is the load-bearing mechanism: an optimization scheme that minimizes, not the training loss at the current parameters, but the worst-case loss in a small ball around them, by perturbing weights along the gradient direction by a radius ρ before computing each update. The paper applies SAM on top of the uniGradICON foundation model, whose architecture is built from ICON's downsampling (Down) and two-step (TS) operators and trained with the GradICON loss — a localized normalized cross-correlation similarity term plus a gradient inverse-consistency regularizer. SAM's adversarial perturbation is computed with the efficient first-order approximation of the inner maximization, so the extra cost per update is roughly one additional forward-backward pass. The flatness of the resulting loss landscape is the explanatory link the paper draws to generalization.","core_discovery":"The central claim is that incorporating Sharpness-Aware Minimization into the training of the uniGradICON foundation model improves its generalization across medical image registration tasks. SAM replaces the ordinary gradient update with a two-step procedure: first compute the gradient at the current weights, take a small step of size ρ in that gradient direction to find a perturbation that maximizes the training loss (an adversarial point within a neighborhood), then compute the actual update gradient at that perturbed point. This minimax objective biases training toward minima whose neighborhoods are uniformly low in loss — flat minima — which the paper finds produces a flatter loss landscape than the base model. The paper evaluates the approach on the ACDC cardiac and SLIVER liver datasets against four comparison baselines and against uniGradICON on five datasets, reporting that the SAM-trained model achieves the highest average Dice scores and near-zero or zero percentages of folded voxels.","pith_inferences":["A controlled re-training study with identical seeds and hyperparameters, varying only SAM, would isolate the optimizer's true contribution from any training-schedule differences in the published comparison.","The method's success on both MRI and CT across four anatomical regions suggests the flat-minima recipe is modality-agnostic and might transfer to 2D imaging (X-ray, ultrasound) with minimal adjustment.","Quantifying the sharpness of the found minimum directly — for instance, by measuring the curvature of the loss surface — rather than relying on visual loss-landscape plots would let future work verify the generalization mechanism more rigorously."],"forward_implications":["The SAM-trained model reports higher average Dice scores than all compared baselines on the ACDC and SLIVER datasets, with folded-voxel percentages near or exactly zero.","Across the five ablation datasets, the method consistently achieves significantly higher Dice coefficients than uniGradICON, indicating stronger cross-dataset generalization.","The loss-landscape visualizations are flatter for the SAM-trained model, matching the proposed mechanism that flat minima generalize better.","Because SAM adds only one hyperparameter and is compatible with any first-order optimizer, the same recipe can be applied to other registration networks without architectural changes."],"supporting_citations":[{"why":"Supplies the Sharpness-Aware Minimization method, including the minimax objective and the first-order perturbation step that the paper adds to training.","marker":"[64]"},{"why":"Defines the uniGradICON foundation model and its composite training set; the paper's method is this model retrained with SAM.","marker":"[60]"},{"why":"Provides the GradICON registration network, its loss (LNCC similarity plus gradient inverse consistency), and the default hyperparameters the paper adopts.","marker":"[82]"},{"why":"Defines the Down and two-step (TS) operators that compose the multi-resolution, multi-step registration network from atomic UNets.","marker":"[83]"},{"why":"Supplies the efficient approximation used to compute the adversarial perturbation in the SAM update with negligible extra cost.","marker":"[85]"}],"fun_headline_variants":["SAM sharpness trick improves medical image registration","Sharpness-aware training enhances registration foundation model","Flat minima make registration model robust across datasets","SAM training lifts Dice on 5 registration datasets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that SAM alone causes the improvement rests on the assumption that the comparison uniGradICON model was trained under identical conditions, with the same data sampling, epochs, loss weights, and random seeds, so that adding SAM is the only change.","fun_headline_variants_meta":{"raw":{"variants":["SAM sharpness trick improves medical image registration","Sharpness-aware training enhances registration foundation model","Flat minima make registration model robust across datasets","SAM training lifts Dice on 5 registration datasets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00053,"raw_usage":{"total_tokens":2548,"prompt_tokens":932,"completion_tokens":1616,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":1560}},"tokens_in":548,"tokens_out":1616,"duration_ms":13246,"temperature":1.0,"reasoning_tokens":1560,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:39:05.766011+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the baseline uniGradICON model from scratch on the same composite training dataset, using the same random seed, number of epochs, learning rate, and loss weights, with SAM as the only added component, and measure Dice overlap on the same five test datasets. If the Dice differences between the two models disappear or become statistically insignificant, the paper's central claim would be falsified.","supporting_citations":[{"cited_title":"unigradicon: A foundation model for medical image registration,","cited_arxiv_id":null,"evidence_quote":"Defines the uniGradICON foundation model and its composite training set; the paper's method is this model retrained with SAM."},{"cited_title":"Gradicon: Approximate diffeomorphisms via gradient inverse consistency,","cited_arxiv_id":null,"evidence_quote":"Provides the GradICON registration network, its loss (LNCC similarity plus gradient inverse consistency), and the default hyperparameters the paper adopts."},{"cited_title":"Icon: Learning regular maps through inverse consistency,","cited_arxiv_id":null,"evidence_quote":"Defines the Down and two-step (TS) operators that compose the multi-resolution, multi-step registration network from atomic UNets."},{"cited_title":"High- performance large-scale image recognition without normaliza- tion,","cited_arxiv_id":null,"evidence_quote":"Supplies the efficient approximation used to compute the adversarial perturbation in the SAM update with negligible extra cost."}],"review_version":1}