{"id":"6d84afab-6349-46d8-bab0-31bc2c521e3f","arxiv_id":"2501.12840","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"AMM-Diff imputes missing MRI modalities from any available subset using an image-frequency fusion network pretrained with masked-image modeling and a diffusion decoder, outperforming Pix2Pix and UMM-CSGM on BraTS 2021.","lead":"AMM-Diff is a diffusion-based system that fills in missing MRI brain scans, such as generating a missing FLAIR or T2 sequence from the scans a patient already has. It is meant to make tumor segmentation and clinical workflows more robust when not all standard MRI sequences can be acquired.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central 'any number of input modalities' claim is only quantitatively tested for single-missing configurations; the 10 multi-missing configurations are acknowledged but not evaluated.","rationale":"The reader's weakest assumption correctly identifies that the IFFN's unified representation is only shown to work for single-missing configurations, while the central claim explicitly promises arbitrary input configurations. My stress-test converges on the same load-bearing concern: the quantitative evaluation in Table 1 covers only the 4 configurations with three available modalities, and the paper itself acknowledges the remaining 10 configurations were not tested. This is not a matter of internal inconsistency; the architecture may well generalize. However, the authors' own qualitative observation about random interpolation when fewer inputs are available suggests the generalization is not guaranteed and may be a genuine weakness. The absence of error bars and statistical tests compounds the issue, since even the single-missing results cannot be assessed for stability. A concrete evaluation across all 14 configurations would settle whether the central claim holds. The correct verdict remains conditional acceptance: the method is promising and the single-missing results are consistent, but the paper must be re-evaluated with multi-missing quantitative experiments before the central claim is accepted. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":5502,"tokens_out":2802,"duration_ms":30885,"concrete_test":"Run AMM-Diff on all 14 nonempty proper subsets of {FLAIR, T1, T1CE, T2} on the BraTS 2021 test split, reporting per-configuration MSE, PSNR, SSIM, and LPIPS for each missing modality, along with mean ± std over at least three training seeds or patient-level bootstrapping. Compare the one-input and two-input configurations against the reported single-missing results and against per-configuration baselines (e.g., UMM-CSGM retrained for each target). If the average MSE for two-input configurations is more than 2× the average MSE for three-input configurations, or if one-input configurations show qualitatively degraded generation consistent with 'random interpolation of tumor regions,' then the central 'any number of input modalities' claim is not supported by the current evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that AMM-Diff is 'capable of handling any number of input modalities and generating the missing ones' (Abstract). The only quantitative evidence in Table 1 is for single-missing configurations, where exactly three modalities are available and one is imputed. The paper explicitly states in Section 3.2: 'Although further experiments could have explored 10 additional configurations, we limited our tests to single-target comparisons due to the 4-page limit.' Those 10 configurations are precisely the cases with one or two input modalities (4 one-input and 6 two-input configurations), which are the regimes where the adaptive architecture and the IFFN's unified representation are most stressed. Figure 3 shows qualitative examples of these configurations, but no MSE/PSNR/SSIM/LPIPS numbers are reported. Without quantitative evaluation of one-input and two-input configurations, the central adaptability claim is unsupported: a model that works well with three inputs may degrade severely, or fail entirely, when fewer inputs are available. The authors even observe qualitatively in Section 3.3 that 'when fewer input modalities are provided, there is a tendency for random interpolation of tumor regions.' This is a limitation explicitly acknowledged in the text and should be treated as a missing piece of evidence, not an artifact. The concern is not that the method is internally inconsistent, but that the evaluation does not match the strength of the claim. In addition, no error bars, confidence intervals, or statistical tests are provided, so it is unclear whether the reported improvements over baselines are significant or stable across subjects and random seeds.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes AMM-Diff, a diffusion-based generative model for imputing missing MRI modalities in brain tumor imaging. The method centers on an Image-Frequency Fusion Network (IFFN) that is pretrained with a masked-image-modeling cosine-similarity objective across all four BraTS modalities, then fine-tuned together with a diffusion decoder. The decoder always outputs all four sequences, so the same trained model can be used for any missing-modality configuration. Quantitative comparisons on BraTS 2021 against Pix2Pix and UMM-CSGM are reported for single-missing-modality settings, along with an ablation replacing IFFN with a U-Net and removing the pretraining step; qualitative examples illustrate one-input and two-input configurations.","tokens_in":11,"tokens_out":2927,"duration_ms":88504,"significance":"If fully supported, the central claim is valuable: a single model that imputes any missing subset of MRI sequences would be practically useful and would go beyond fixed translation pairs. The paper has several strengths: it uses a public dataset (BraTS 2021), compares against two established baselines, includes an ablation isolating the contribution of the pretrained IFFN, and reports all four standard metrics. The main weakness is that the 'any number of input modalities' claim is only quantitatively tested for the single-missing case; the one-input and two-input regimes, where adaptability matters most, are supported only by qualitative examples and are explicitly acknowledged to degrade. The absence of error bars and statistical tests further limits the strength of the quantitative conclusions.","major_comments":[{"comment":"The central claim that AMM-Diff is 'capable of handling any number of input modalities and generating the missing ones' is only quantitatively evaluated for single-missing configurations in Table 1, where exactly three modalities are present and one is imputed. The ten remaining configurations (four one-input and six two-input cases) are acknowledged in Section 3.2 but are not evaluated numerically. Because the architecture's adaptive reconstruction and the IFFN's unified representation are most stressed precisely in the one- and two-input regimes, the adaptability claim is not supported by the current quantitative evidence. Please provide quantitative results for at least a representative subset of these configurations, or restrict the claim to the evaluated setting.","section":"Abstract, Section 3.2, Table 1"},{"comment":"No variance estimates, multiple seeds, or statistical tests are reported. Several differences between AMM-Diff and UMM-CSGM are small in absolute terms (e.g., FLAIR PSNR 30.707 vs. 29.416, T1CE SSIM 0.9311 vs. 0.9112, T2 LPIPS 0.0235 vs. 0.0268), so it is not clear whether these differences are reproducible. Please report mean and standard deviation over at least three independent training runs and use a paired significance test (e.g., Wilcoxon signed-rank or bootstrapped confidence intervals) for the headline claims.","section":"Section 3.2, Table 1"},{"comment":"The paper explicitly states that 'when fewer input modalities are provided, there is a tendency for random interpolation of tumor regions, especially without key modality pairs like FLAIR/T2 or T1/T1CE.' This is an admitted degradation in exactly the low-input configurations that distinguish the method from fixed single-target translators. The claim of robustness for arbitrary input configurations is therefore directly qualified. Please quantify this failure mode (e.g., MSE/SSIM on the problematic configurations) and discuss whether it is clinically acceptable, or modify the claim to reflect the observed limitation.","section":"Section 3.3"},{"comment":"The assertion that pretraining 'enables the IFFN to learn the underlying correlations between different modalities, which will be distilled as prior knowledge in the case of missing modalities' is not tested directly. The learned representation is never probed or compared across input configurations, and the quantitative evaluation only covers the three-input case. The transfer of the pretrained IFFN to unseen one- and two-input configurations is therefore an untested assumption. A representation-level analysis (e.g., comparing IFFN embeddings for different modality subsets or ablating the pretraining objective on low-input configurations) would strengthen the causal link between the pretext task and the adaptability claim.","section":"Section 2.2"}],"minor_comments":[{"comment":"The bullet 'spatial and spectral feature respresention' contains a typo: 'respresention' should be 'representation.'","section":"Section 1, Contributions"},{"comment":"The spacing in 'V AEs' is inconsistent; it should be 'VAEs' throughout.","section":"Section 1"},{"comment":"The notation is under-specified: Heq is stated as histogram equalization but its precise implementation is not given, and fN is described only as 'a smooth high-pass filter' without its shape or cutoff behavior. Please provide a concrete definition or a reference.","section":"Equation (3)"},{"comment":"The dataset name is written as 'BRATS2021' in Section 3.1 but as 'BraTS 2021' in the rest of the paper; please use a consistent spelling.","section":"Section 3.1"},{"comment":"The sentence 'we limited our tests to single-target comparisons due to the 4-page limit' is unusual in a research paper. If there is a page limit, a supplementary document or appendix with the remaining configurations would be more appropriate and would address the main evaluation gap.","section":"Section 3.2"},{"comment":"Please make the modality labels explicit in the figure captions. Figure 2 says 'Columns represent FLAIR, T1CE, T1, and T2 sequences (left to right),' but Figure 3 does not state which modality each column shows, which makes the qualitative claims about T1CE, T2, FLAIR/T2, and T1/T1CE difficult to verify.","section":"Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of a medical-imaging or computer-vision venue, and the proposed approach is reasonable. However, the gap between the advertised 'any number of input modalities' capability and the quantitative evidence (single-missing only, no error bars) is substantial. The authors should either add the missing quantitative experiments, ideally in a supplementary document, or clearly narrow the claim to the evaluated setting. The qualitative observation about tumor-region interpolation in low-input configurations should also be addressed head-on."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take on arXiv:2501.12840. The real contribution is the combination: a Siamese IFFN that fuses spatial and high-frequency Fourier features, pretrained with masked-image modeling, then fine-tuned with the diffusion decoder, plus a fixed-output adaptive reconstruction head that lets one model impute any missing MR sequence. The ablation shows the pretraining matters—pretrained IFFN beats the non-pretrained variant on every metric in Table 1, and the gap is not tiny. That is a concrete, reproducible finding, and the architecture is genuinely more flexible than the per-sequence baselines they compare against.\n\nThe soft spot is exactly where the stress-test lands. The central claim is 'any number of input modalities', but Table 1 only quantifies single-missing configurations (three inputs, one output). The paper itself says the 10 additional configurations were skipped due to the 4-page limit. Those are precisely the one-input and two-input cases where the adaptive design is most stressed. Figure 3 shows a few qualitative examples and the authors admit 'when fewer input modalities are provided, there is a tendency for random interpolation of tumor regions.' That is an honest observation, but it means the headline claim is not backed by numbers. No error bars or statistical tests either, so the reported improvement over UMM-CSGM might be within noise for some sequences, though the consistency across all four sequences and all metrics makes me think the effect is real.\n\nOne more thing: the paper describes UMM-CSGM as 'specialized for a single sequence', but the title of that arXiv paper is a unified conditional score-based framework for multi-modal completion. The comparison might be fair in practice if that prior work is trained separately per target, but the characterization is at least imprecise and should be checked.\n\nNet: this is a worthwhile paper for a workshop or conference with a serious revision. The method is plausible, the ablation is informative, and the limitation is stated rather than hidden. What it needs before the central claim can be trusted: quantitative results for the multi-missing configurations, error bars or repeated runs, and ideally a reproducibility release. I'd send it to peer review, asking for those additions.","headline":"Solid engineering with a clean adaptive imputation architecture, but the headline capability—any number of input modalities—is only quantified for single-missing cases, and the authors say so themselves.","tokens_in":6373,"tokens_out":1606,"would_cite":true,"duration_ms":16211,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that one diffusion model, guided by a self-supervised fusion representation, can impute any missing MRI sequence from whatever combinations of modalities are available.","keywords":["missing modality imputation","diffusion models","multi-modal MRI","Siamese network","masked image modeling","image-frequency fusion","BraTS 2021"],"falsifier":"Train the same AMM-Diff on BraTS 2021 using only single-missing cases as reported, then evaluate imputation on held-out cases where two or three modalities are missing (for example, only FLAIR and T2 available) and compare pixel- and perception-wise error against a model trained specifically for that configuration. If the adaptive model's error is not close to the specialized model's, the unified-representation claim fails.","tokens_in":5271,"feed_emoji":"🧠","tokens_out":8740,"duration_ms":67389,"temperature":0.7,"pith_summary":"This paper tries to show that missing MRI sequences can be imputed by a single diffusion model no matter which sequences are present, instead of training one translator per target. It builds an Image-Frequency Fusion Network (IFFN) that is pretrained with a masked-image self-supervision task so that any subset of input modalities maps to one shared feature representation, then fine-tunes the IFFN together with a diffusion decoder that reconstructs all four MR sequences. On BraTS 2021, the resulting AMM-Diff outperforms Pix2Pix and UMM-CSGM on every metric for each single-missing-sequence case while using one model instance. If true, this removes the need to know the missing-modality configuration in advance and makes imputation practical in clinical settings where acquisition protocols vary.","feed_headline":"One model fills in any missing MRI sequence from the rest","feed_subtitle":"A single trained model imputes whichever MRI sequences are absent on BraTS 2021, no per-target translators.","key_machinery":"The load-bearing object is the Siamese Image-Frequency Fusion Network (IFFN), a two-branch encoder with downscale factor two that maps any combination of input modalities to a single unified feature map. It combines spatial fusion with a spectral branch that extracts high-frequency Fourier components through a smooth high-pass filter followed by histogram equalization, which helps preserve anatomical structure and tumor boundaries. A masked-image-modeling pretext task trains the Siamese branches to minimize cosine similarity loss between patch-masked and complete multimodal inputs, giving the representation its invariance to missing inputs. The IFFN is then fine-tuned end-to-end with a diffusion decoder whose output is fixed to the full set of MR sequences, so the same model instance both reconstructs available modalities and synthesizes missing ones.","core_discovery":"The central discovery claimed is that a unified feature representation learned from complete multimodal data, combined with an adaptive reconstruction strategy that always outputs the full set of sequences, lets one conditional diffusion model handle arbitrary missing-modality configurations. The IFFN's self-supervised pretext task—minimizing a cosine similarity loss between patch-masked multimodal inputs and their complete counterparts—distills the correlations among FLAIR, T1, T1CE, and T2 into prior knowledge. During fine-tuning, the diffusion model's noise-prediction loss back-propagates through the IFFN, and the decoder is trained to reconstruct both missing and available sequences so that convolutional outputs stay fixed at the total number of modalities. Quantitatively, AMM-Diff with the pretrained IFFN reports the best MSE, PSNR, SSIM, and LPIPS for each of the four sequences against Pix2Pix and UMM-CSGM on BraTS 2021; qualitatively, it preserves brain structure and tumor boundaries across varied input configurations.","pith_inferences":["Inference: if the unified representation truly generalizes across arbitrary input subsets, the same pretrained IFFN should impute missing modalities in unseen datasets or scanner protocols without retraining, but the paper only demonstrates this on BraTS 2021.","Inference: a direct test of the adaptability claim would be to train on all single-missing cases and evaluate on held-out multi-missing configurations; the paper shows only qualitative examples of those configurations with no quantitative metric.","Inference: reconstructing available modalities alongside missing ones may act as an implicit regularizer that keeps the decoder faithful to the input anatomy; isolating this effect by ablating reconstruction of available inputs would quantify its contribution.","Inference: the Fourier high-frequency branch likely matters most for tumor-boundary sharpness, and an ablation that removes only the spectral path while keeping the Siamese pretraining could reveal how much of the gain is spectral versus self-supervised."],"forward_implications":["A single trained AMM-Diff instance can replace the per-target models needed by Pix2Pix and UMM-CSGM, cutting training and deployment cost for multi-sequence MRI imputation.","Imputed sequences can be fed to downstream brain-tumor segmentation pipelines, potentially recovering accuracy lost when modalities are absent at test time.","Because the architecture treats 3D volumes as stacks of 2D slices, the same method can be applied to full 3D image translation without redesign.","The method's stated generality suggests it can be extended beyond MR sequences to other cross-modal translations such as CT or PET."],"supporting_citations":[{"why":"The UMM-CSGM conditional score-based baseline, the main multi-modal completion comparison.","marker":"[3]"},{"why":"Denoising diffusion probabilistic models, the diffusion formulation the translation decoder is built on.","marker":"[7]"},{"why":"Masked language modeling pretext task, which motivates the masked image modeling used to pretrain the IFFN.","marker":"[10]"},{"why":"Diffusion generative feedback procedure, the end-to-end fine-tuning mechanism that updates the IFFN with the diffusion loss.","marker":"[11]"},{"why":"BraTS 2021 dataset with four MR sequences used for training, validation, and testing.","marker":"[12]"},{"why":"Pix2Pix image-to-image translation baseline compared against in the experiments.","marker":"[13]"}],"fun_headline_variants":["One diffusion model imputes any missing MRI sequence","Adaptive diffusion fills absent MRI modalities from what remains","AMM-Diff: one model for any missing MRI input","Single adaptive model completes any missing MRI combo"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument relies on the pretrained IFFN representation transferring to input configurations it has not been quantitatively tested on: all quantitative results use the single-missing case, while the central adaptability claim covers any number of missing modalities.","fun_headline_variants_meta":{"raw":{"variants":["One diffusion model imputes any missing MRI sequence","Adaptive diffusion fills absent MRI modalities from what remains","AMM-Diff: one model for any missing MRI input","Single adaptive model completes any missing MRI combo"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000468,"raw_usage":{"total_tokens":2344,"prompt_tokens":971,"completion_tokens":1373,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":1310}},"tokens_in":587,"tokens_out":1373,"duration_ms":10471,"temperature":1.0,"reasoning_tokens":1310,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:43:43.919535+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same AMM-Diff on BraTS 2021 using only single-missing cases as reported, then evaluate imputation on held-out cases where two or three modalities are missing (for example, only FLAIR and T2 available) and compare pixel- and perception-wise error against a model trained specifically for that configuration. If the adaptive model's error is not close to the specialized model's, the unified-representation claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The UMM-CSGM conditional score-based baseline, the main multi-modal completion comparison."},{"cited_title":"A Novel Unified Conditional Score-based Generative Framework for Multi-modal Medical Image Completion","cited_arxiv_id":"2207.03430","evidence_quote":"Denoising diffusion probabilistic models, the diffusion formulation the translation decoder is built on."},{"cited_title":"Denoising diffusion probabilistic models,","cited_arxiv_id":null,"evidence_quote":"Diffusion generative feedback procedure, the end-to-end fine-tuning mechanism that updates the IFFN with the diffusion loss."},{"cited_title":"Diffusion models in medical imaging: A com- prehensive survey,","cited_arxiv_id":null,"evidence_quote":"BraTS 2021 dataset with four MR sequences used for training, validation, and testing."},{"cited_title":"Diffusion models beat gans on image synthesis,","cited_arxiv_id":null,"evidence_quote":"Pix2Pix image-to-image translation baseline compared against in the experiments."}],"review_version":1}