{"id":"e1ae463e-22dd-404b-9ee3-0f1d8f6b18bc","arxiv_id":"2608.08319","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Graph-Blueprint Pruning lets a 425,000-scan brain MRI foundation model expand across clinical domains with less forgetting than sequential tuning or EWC.","lead":"Alcmaeon is a 3D brain MRI model trained on more than 425,000 scans that can learn new clinical domains one after another while retaining earlier abilities. It uses Graph-Blueprint Pruning to freeze important network modules, reporting lower forgetting than sequential fine-tuning and elastic weight consolidation, which points toward MRI foundation models that can grow with new data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"EWC baseline appears effectively broken (10.1, 8.3, 13.2 dB PSNR on D1–D3 after D4), implying undertuning or misimplementation; without a properly tuned EWC the headline advantage is not established.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: EWC is not tuned or specified fairly, and its reported catastrophic collapse (10.1, 8.3, 13.2 dB on D1–D3) is atypical for a reasonable EWC implementation. The main text provides no lambda, Fisher sample size, or loss details, and the background-inclusive results show EWC also fails to learn D4 well, which is more consistent with over-regularization or a bug than with a meaningful trade-off. This directly undermines the abstract's contrast with EWC. Other concerns, such as missing code/weights, limited downstream sample sizes, or lack of comparison to other continual-learning methods, are secondary because the headline claim is specifically about beating SEQ and EWC on reconstruction retention. The background-retained evaluation is good practice but does not fix the EWC baseline issue. Since the paper could be made valid by re-running EWC with proper tuning, a conditional verdict remains appropriate rather than outright rejection. The concrete test would settle the matter by checking whether any reasonable EWC configuration narrows the reported gap.","tokens_in":35605,"tokens_out":8741,"duration_ms":75031,"concrete_test":"Re-run the single-channel D1–D4 continual learning with EWC using a grid of regularization strengths (e.g., lambda = 0.1, 1.0, 10, 100) and a Fisher computed on a fixed subset (e.g., 1,000 samples per domain) with the reconstruction loss; report D1–D4 PSNR trade-off curves. If any lambda achieves D1–D3 PSNR above 20 dB while D4 PSNR stays above 22 dB, the reported EWC collapse is a tuning artifact and the GBP-vs-EWC claim is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that GBP 'showed less forgetting than sequential adaptation and elastic weight consolidation across voxel-level reconstruction measures.' The EWC numbers in the single-channel configuration are pathological: after D4 adaptation, EWC retains only 10.1, 8.3 and 13.2 dB PSNR on D1–D3, which is near-noise level and corresponds to peak forgetting F=10.20 dB. In the background-inclusive evaluation, EWC also learns D4 poorly (19.26 dB vs 24.45 for SEQ and 22.75 for GBP), indicating over-regularization or an incorrectly computed Fisher. Methods states only 'a diagonal Fisher approximation' with no regularization strength (lambda), Fisher sample count, or loss used for the Fisher. If the EWC baseline is undertuned or misimplemented, then the comparison to EWC is not informative and the claimed 'largest advantage after adaptation to tumour imaging' is inflated relative to a reasonable EWC configuration. This is load-bearing because the abstract's headline claim directly contrasts GBP with EWC; a fair EWC baseline could substantially narrow the gap and alter the practical conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Alcmaeon, a three-dimensional brain MRI foundation model pretrained on over 425,000 volumes, and proposes Graph-Blueprint Pruning (GBP) as a continual-learning mechanism that freezes network modules selected by activation–gradient salience while leaving other modules trainable. The model is sequentially expanded across four clinical domains (healthy ageing, neurodegeneration, developmental/psychiatric imaging, and brain tumours), and the authors report that GBP exhibits less forgetting than sequential fine-tuning (SEQ) and elastic weight consolidation (EWC) on voxel-level reconstruction metrics (PSNR, MSE, RMSE, MAE), with the largest advantage after adaptation to tumour imaging. The paper also evaluates representations from the continually adapted models on cross-modal synthesis, Alzheimer's disease classification, glioma survival modelling, and postoperative outcome prediction, finding task-dependent utility rather than a single universally optimal representation. The manuscript is accompanied by extensive extended-data figures and a supplementary information document.","tokens_in":35854,"tokens_out":5642,"duration_ms":53551,"significance":"If the central claim is upheld, GBP would offer a practical and inspectable structural-memory mechanism for expanding large volumetric foundation models across clinical domains, which is a relevant problem for medical imaging. The work is unusually broad in scope: it combines a large-scale pretraining pipeline, a new continual-learning method, multiple input-channel configurations, a background-retained robustness check, and several clinical downstream tasks. The authors are transparent about many limitations, including the single domain order, the small postoperative cohort, and the exploratory nature of the anatomical attribution maps. However, the headline comparison against EWC currently rests on an EWC implementation whose key hyperparameters are undisclosed and whose reported behaviour is consistent with severe over-regularization, so the quantitative claim of 'less forgetting than EWC' is not yet established on the evidence presented.","major_comments":[{"comment":"The EWC baseline is not specified sufficiently to be considered a fair comparator. Methods states only that EWC 'used a diagonal Fisher approximation' and gives no regularization strength (lambda), Fisher sample count, or the loss used for the Fisher. In the single-channel evaluation, EWC retains only 10.1, 8.3 and 13.2 dB PSNR on D1–D3 after D4 adaptation, and in the background-inclusive evaluation its D4 learning accuracy is 19.26 dB versus 24.45 dB for SEQ and 22.75 dB for GBP (Fig. 2 and Extended Data Fig. 3). These values are consistent with a heavily over-regularized or misimplemented EWC. Because the abstract's headline claim directly contrasts GBP with EWC, a properly tuned EWC, with lambda reported and a sensitivity analysis over lambda, is required before 'less forgetting than ... EWC' can be considered established.","section":"Methods: Continual-learning strategies; Results: Graph-Blueprint Pruning limits forgetting"},{"comment":"The GBP method depends on hyperparameters that are not reported: the upper-quantile threshold for selecting salient units, the PCA variance or number of components used to compress salience vectors, and the number of probe samples aggregated per domain. These choices control how much capacity is frozen at each stage and directly affect the forgetting–adaptation trade-off. The manuscript gives no sensitivity analysis around the quantile threshold, so it is unclear how robust the reported advantage over SEQ/EWC is to the main methodological knob. Please report these values and evaluate at least a small range of thresholds.","section":"Methods: Continual-learning strategies; Results: Graph-Blueprint mapping"},{"comment":"The continual-learning comparisons in Fig. 2 and Extended Data Figs. 3–4 report standard deviations per cell but no multiple-seed information or statistical comparison of the forgetting differences. The central claim is comparative ('GBP showed less forgetting'), and the reader cannot tell whether the reported F, BWT, and ACC values come from a single run or from repeated runs, or whether the differences are statistically significant. Please state the number of independent runs/seeds, clarify whether the SDs are across subjects, folds, or seeds, and provide paired tests or confidence intervals for the key forgetting/backward-transfer/accuracy differences.","section":"Evaluation metrics; Supplementary Information S6"}],"minor_comments":[{"comment":"'Overal survival' should be 'Overall survival'.","section":"Fig. 5 and Extended Data Fig. 5 captions"},{"comment":"The salience formula is typeset as [s_n(x) = E∗j |a∗n,j(x) ∂an,j L|]; the expectation index and the meaning of the starred quantities are unclear. Please define a_n,j and clarify the aggregation.","section":"Methods: Continual-learning strategies"},{"comment":"The sentence 'We acknowledge the use of the facilities of the Research Computing Services (RCS) of University of Cambridge, UK.' appears twice; delete the duplicate.","section":"Acknowledgements"},{"comment":"Use consistent capitalization for 'Swin'/'SWIN' (e.g., '3D-SWIN' vs '3D-Swin') and for 'single-channel' vs '1-channel' across text and figures.","section":"Throughout the manuscript"},{"comment":"The methods for cross-modal synthesis state that competing generative models were trained with 'recommended repository settings' but do not list the actual hyperparameters for VQVAE, AEKL, MAISI, LDM, MOTFM, and Rectified Flow; specify them for reproducibility.","section":"Methods: Clinical fine-tuning"},{"comment":"The caption contains the fragment 'mEWC one-channel modelOne-channel model'; this appears to be a copy-paste artifact and should be fixed.","section":"Extended Data Fig. 6 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is broad and ambitious, and the core continual-learning claim could be publishable if the EWC baseline is properly tuned and fully reported. I would encourage the editor to require the authors to make code and trained weights available to reviewers during the revision process, since reproducibility is central for a methodological claim. The scope is very large (foundation model pretraining, continual learning, synthesis, survival, classification, few-shot adaptation); the central novelty, GBP, deserves a focused evaluation with disclosed hyperparameters and a sensitivity analysis before the broader downstream claims are emphasized."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Should know: this is a substantial empirical paper, not a conceptual breakthrough. It pretrains a 3D brain MRI foundation model on 425k volumes and evaluates a new continual-learning mechanism (GBP) that freezes high-salience modules using activation–gradient salience and PCA-compressed blueprints. The results are honestly reported and carefully hedged. The main weakness is the EWC baseline, which looks broken in the reported numbers.\n\nWhat's new: GBP combines existing ideas (module-level salience, masking/freezing, PCA compression) in a concrete way and applies it to a volumetric MRI model at a scale that hasn't been done before. The downstream evaluation is broad: synthesis, Alzheimer's classification, glioma survival, and few-shot postoperative prediction. The paper includes a background-retained robustness check, explicitly acknowledges that SSIM is aligned with the training objective, and states its own limitations (single domain order, small 49-patient cohort, blueprints as inspectable memory rather than causal evidence). That level of hedging is rare and earns credit.\n\nSoft spots: the EWC comparison is the big one. EWC drops to 10.1, 8.3, and 13.2 dB PSNR on D1–D3 after D4, which is essentially noise. Methods gives no regularization strength, no Fisher sample count, and no loss for the Fisher. Either EWC is undertuned or misimplemented. That means the abstract's claim of 'less forgetting than ... elastic weight consolidation' is not yet established; a properly tuned EWC could narrow the gap substantially. However, the advantage over sequential fine-tuning is independently shown (peak forgetting 0.03 vs 5.93 dB in single-channel), so the core idea survives even if the EWC comparison is invalid. The SSIM alignment issue is real but acknowledged, and the voxel-wise metrics give a consistent picture. Lack of released code/weights and reliance on inaccessible Supplementary Information for key hyperparameters is a reproducibility problem.\n\nWho is this for: anyone working on continual learning for medical imaging or on scaling 3D foundation models. It deserves a serious referee, but the authors should be required to release code/weights, disclose and tune EWC, and provide the missing methodological details. Even then, this is a solid incremental contribution, not a paradigm shift.","headline":"Large-scale, honestly hedged empirical study of continual learning for 3D brain MRI; the GBP method looks promising, but the EWC baseline appears undertuned, weakening part of the headline claim.","tokens_in":36401,"tokens_out":3418,"would_cite":true,"duration_ms":28862,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07","68U10","92C55"],"pacs":[],"model":"deepseek-v4-flash","headline":"A 3D brain MRI foundation model pretrained on 425,000 volumes can be sequentially expanded from healthy ageing to tumour imaging while retaining earlier skills, provided newly important network modules are frozen and the rest kept…","keywords":["brain MRI foundation model","continual learning","catastrophic forgetting","Graph-Blueprint Pruning","latent diffusion","self-supervised learning","3D Swin transformer","neuroimaging domain expansion"],"falsifier":"Run the D1–D4 expansion again with EWC's regularisation strength swept over at least an order of magnitude, recomputing the Fisher information on each domain's training set, and compare retention on D1–D3 after D4; if a well-tuned EWC reaches the 19–26 dB retention levels that GBP reports, the claimed advantage is an artifact of the baseline's tuning.","tokens_in":35425,"feed_emoji":"🧠","tokens_out":6635,"duration_ms":59339,"temperature":0.7,"pith_summary":"This paper tries to establish that a large, self-supervised brain MRI foundation model can keep learning new clinical domains without losing what it already knows. The authors pretrain Alcmaeon on more than 425,000 three-dimensional volumes and derived maps, then adapt it sequentially to healthy ageing, neurodegeneration, developmental and psychiatric imaging, and adult brain tumours. Their central proposal, Graph-Blueprint Pruning (GBP), ranks the network's computational modules by how much they contribute to each domain and freezes the highest-ranking modules before the next domain is learned, while leaving the remaining modules trainable. On voxel-level reconstruction measures, GBP shows less catastrophic forgetting than sequential fine-tuning and elastic weight consolidation, with the largest advantage after the tumour domain, the most disruptive transition. If the claim holds, foundation models for neuroimaging need not be treated as fixed, one-shot systems: they can grow incrementally as new data arrive, and the frozen modules provide an inspectable record of how capacity was protected and reused.","feed_headline":"Brain MRI model expands across diseases with less forgetting","feed_subtitle":"Graph-Blueprint Pruning freezes only needed modules, beating EWC and sequential fine-tuning on retention.","key_machinery":"The load-bearing mechanism is Graph-Blueprint Pruning (GBP), a structural-memory operator that partitions the network into computational modules (attention, feed-forward, normalisation and adaptive-modulation components) and ranks their domain-associated contributions by an activation–gradient salience score such as $s_n(x) = \\mathbb{E}_j^*\\bigl|a^*_{n,j}(x)\\,\\partial_{a_{n,j}} L\\bigr|$. Salience vectors are aggregated over probe samples, compressed by principal component analysis into a task blueprint, and an upper-quantile rule selects the salient modules among the currently trainable complement; those modules are frozen exactly by masking their optimizer updates, and the frozen sets accumulate monotonically across domains. Separate blueprint streams are maintained for the 3D-SWIN encoder–decoder and the 3D-DiT latent diffusion generator, and after four domains roughly 58% of the predefined modules in each stream remain trainable. This is what distinguishes GBP from elastic weight consolidation: instead of softly penalising parameter changes, it imposes a hard, inspectable structural constraint while keeping most of the model's capacity available for future learning.","core_discovery":"The paper's central claim is that structural module freezing, rather than soft parameter anchoring, is what lets a volumetric MRI model expand across clinically distinct domains with limited forgetting. GBP computes activation–gradient salience for each predefined computational module, compresses the salience vectors with principal component analysis, and freezes the high-salience modules by masking their optimizer updates, so each new domain is learned in the residual trainable subspace. After adaptation to adult brain tumours, the most disruptive domain, GBP retained 25.6, 18.5 and 19.3 dB PSNR on the three earlier domains in the single-channel setting, while EWC collapsed to 10.1, 8.3 and 13.2 dB, and GBP still reached 21.4 dB on the new tumour domain. The same ordering held in two-channel and hybrid input configurations, with GBP peak forgetting between 0.03 and 0.87 dB compared with roughly 10–12 dB for EWC, although SSIM told a different story and showed GBP paying a structural-similarity cost under two-channel learning and channel collapse. The paper also shows that no single representation level is best for every clinical objective: encoder features supported cross-modal synthesis, intermediate features supported Alzheimer's classification, combined representations supported glioma risk ranking, and GBP's clearest downstream advantage came in few-shot postoperative outcome prediction.","pith_inferences":["The success of GBP suggests the salience ranking itself is doing causal work; a natural control experiment the paper does not run is freezing a randomly chosen set of modules of the same size, which should forget more if the blueprint selection matters.","Because blueprints accumulate monotonically, longer sequences than the four domains tested here will eventually exhaust trainable modules, so compressing, sharing, or safely releasing protected capacity is the next design problem rather than a proven solution.","The strong postoperative few-shot result with microstructural maps hints that protected modules retain low-level diffusion and NODDI features; this could be tested by ablating blueprint-protected modules during postoperative inference and measuring the drop in F1.","Domain order is likely to change the relative ranking of methods; the paper places tumour imaging last, and a milder final domain would presumably shrink GBP's advantage over sequential fine-tuning and EWC."],"forward_implications":["If GBP works as reported, brain MRI foundation models can be updated incrementally with new diseases, cohorts and protocols instead of being re-pretrained or replaced by one model per domain.","Hard structural freezing of selected modules is a viable alternative to soft regularisation when the new domain differs strongly in pathology and anatomy, with the largest measured gap on tumour imaging.","The cumulative blueprint gives an inspectable map of protected versus available capacity, which could support auditing and targeted release of modules in regulated clinical settings.","Because no single representation wins every task, practical deployments should select feature depth and component per endpoint: encoder features for synthesis, intermediate features for classification, combined representations for survival ranking.","Background-inclusive SSIM saturates and masks forgetting, so retention benchmarks should report signal-region voxel metrics alongside structural ones to avoid mistaking limited plasticity for successful retention."],"supporting_citations":[{"why":"Supplies the elastic weight consolidation baseline whose retention GBP must beat.","marker":"[8]"},{"why":"Provides the winning-subnetwork idea that motivates protecting selected modules.","marker":"[9]"},{"why":"Provides hard-attention task masking, the structural-protection lineage GBP builds on.","marker":"[10]"},{"why":"Supplies the pathway-protection continual learning setting that motivates the paper's approach.","marker":"[11]"},{"why":"Provides the shifted-window transformer blocks underlying the 3D-SWIN encoder–decoder.","marker":"[12]"},{"why":"Supplies the SwinUNETR-style 3D self-supervised architecture that 3D-SWIN adapts.","marker":"[13]"},{"why":"Provides the denoising diffusion objective used by the 3D-DiT latent generator.","marker":"[17]"},{"why":"Provides the variational autoencoder formulation compared against the deterministic latent bottleneck.","marker":"[19]"},{"why":"Supplies the UK Biobank pretraining corpus that gives Alcmaeon its population-scale starting representation.","marker":"[33]"},{"why":"Supplies the skull-stripping preprocessing used to harmonise the training volumes.","marker":"[44]"}],"fun_headline_variants":["Graph-Blueprint Pruning lets brain MRI model grow with less forgetting","Blueprint pruning beats EWC in expanding brain MRI model","Alcmaeon: brain MRI model expands without losing earlier skills","Module freezing, not soft anchoring, reduces forgetting in brain MRI","Structural pruning helps brain MRI model adapt across diseases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes elastic weight consolidation was implemented and tuned fairly; the reported single-channel EWC collapse to 10.1, 8.3 and 13.2 dB PSNR on D1–D3 after D4 is atypical for reasonable EWC settings, so the headline advantage over EWC could be inflated if EWC was undertuned.","fun_headline_variants_meta":{"raw":{"variants":["Graph-Blueprint Pruning lets brain MRI model grow with less forgetting","Blueprint pruning beats EWC in expanding brain MRI model","Alcmaeon: brain MRI model expands without losing earlier skills","Module freezing, not soft anchoring, reduces forgetting in brain MRI","Structural pruning helps brain MRI model adapt across diseases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001037,"raw_usage":{"total_tokens":4405,"prompt_tokens":1026,"completion_tokens":3379,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":3295}},"tokens_in":642,"tokens_out":3379,"duration_ms":24847,"temperature":1.0,"reasoning_tokens":3295,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:08:33.589217+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the D1–D4 expansion again with EWC's regularisation strength swept over at least an order of magnitude, recomputing the Fisher information on each domain's training set, and compare retention on D1–D3 after D4; if a well-tuned EWC reaches the 19–26 dB retention levels that GBP reports, the claimed advantage is an artifact of the baseline's tuning.","supporting_citations":[{"cited_title":"Proceedings of the National Academy of Sciences of the United States of America 114, 3521–3526 (2017)","cited_arxiv_id":null,"evidence_quote":"Supplies the elastic weight consolidation baseline whose retention GBP must beat."},{"cited_title":"Forget-free continual learning with winning subnetworks, Vol","cited_arxiv_id":null,"evidence_quote":"Provides the winning-subnetwork idea that motivates protecting selected modules."},{"cited_title":"& Karatzoglou, A.Overcoming catastrophic forget- ting with hard attention to the task, Vol","cited_arxiv_id":null,"evidence_quote":"Provides hard-attention task masking, the structural-protection lineage GBP builds on."},{"cited_title":"Learning without isolation: Pathway protection for continual learn- ing, Vol","cited_arxiv_id":null,"evidence_quote":"Supplies the pathway-protection continual learning setting that motivates the paper's approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SwinUNETR-style 3D self-supervised architecture that 3D-SWIN adapts."},{"cited_title":"& Abbeel, P.Denoising diffusion probabilistic models, Vol","cited_arxiv_id":null,"evidence_quote":"Provides the denoising diffusion objective used by the 3D-DiT latent generator."},{"cited_title":"https://www.ukbiobank.ac.uk/","cited_arxiv_id":null,"evidence_quote":"Supplies the UK Biobank pretraining corpus that gives Alcmaeon its population-scale starting representation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the skull-stripping preprocessing used to harmonise the training volumes."}],"review_version":1}