{"id":"6154fd7b-1814-4ba8-b8fe-ebe9693ce315","arxiv_id":"2507.00185","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A self-supervised vision model with a memory module reports AUROC from 0.86 to 0.99 across seven medical imaging modalities, often matching or beating specialty foundation models.","lead":"MerMED-FM is a single AI model trained without labels on 3.3 million medical images spanning CT, chest X-ray, ultrasound, pathology, eye scans, and skin photos. The authors report it matches or beats several specialty-specific models across 25 public and 7 local datasets, with AUROC values between 0.86 and 0.99 depending on the imaging type.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The pretraining corpus is never enumerated (deferred to missing Supplementary Table 6), so test-set contamination cannot be excluded; the 0.002 AUROC margin over Dino could be erased by a small overlap.","rationale":"After reading the full text, the strongest part of the paper is the breadth of external evaluation across 25 datasets and 7 modalities, and the architecture section gives enough detail to be plausibly implemented. The load-bearing soft spot is pretraining corpus provenance. The reader's weakest assumption independently identifies the same gap, and I agree. Other concerns (missing code/weights, typos such as OCTDL 0.911 vs 0.991, malformed Table 3 rows) are real but secondary: they affect reproducibility and presentation, not the logical validity of the comparison. A leakage audit is the single check that can settle the central claim. I do not see a reason to move the reader's conditional verdict; if the overlap check fails, the verdict should be reconsidered, but on current evidence the correct disposition remains the same as the reader's.","tokens_in":22837,"tokens_out":12389,"duration_ms":135639,"concrete_test":"Release Supplementary Table 6 with every pretraining source (name, URL/version, download date) and run an exact plus near-duplicate image-level overlap analysis (e.g., perceptual hash or pretrained-feature retrieval) between the pretraining corpora and all 25 public evaluation sets; then re-fine-tune MerMED-FM on the non-overlapping subset of each evaluation training split and recompute the Table 2 AUROCs. If the headline mean AUROC and the 0.002 Dino margin survive unchanged, the leakage concern is resolved; if any evaluation image appears in pretraining, the reported comparison is uninterpretable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"MerMED-FM's central claim is the headline comparison: mean AUROC 0.935 vs Dino 0.933 and BiomedCLIP 0.919, plus reported advantages over specialty FMs. This claim depends on pretraining images being disjoint from the 25 public evaluation datasets. The Methods section states only that the 3.3M-image corpus comes from 'publicly available unlabelled datasets' and defers the list to Supplementary Table 6, which is absent from the provided text. No exclusion, deduplication, or overlap analysis is described anywhere. Many evaluation datasets (RSNA, SIIM, APTOS2019, IDRiD, MESSIDOR2, HAM10000, Dermnet, BUSI) are public and widely reused; several are plausible components of the stated 401,059 dermoscopic and 333,700 fundus pretraining images. Because SSL pretraining on even a small fraction of test images can improve fine-tuned AUROC for those exact images, and the reported margin over Dino is only 0.002, a few percent overlap could account for the headline difference. The missing provenance is therefore not cosmetic: without Table 6 and an overlap audit, the central claim is uncheckable from the manuscript.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents MerMED-FM, a vision-only self-supervised foundation model trained on 3.3 million unlabelled medical images spanning CT, CXR, ultrasound, histopathology, color fundus photography, OCT, and dermatology, using a teacher–student ViT architecture augmented with a FIFO memory module. The authors report downstream fine-tuning results across 25 public and seven local datasets, comparing against DINO, BiomedCLIP, and six single-modality foundation models, with a headline mean AUROC of 0.935 versus 0.933 for DINO and 0.919 for BiomedCLIP. They also report data-efficiency experiments at 10%, 30%, and 50% of the fine-tuning data and claim general equivalence or superiority to specialty models on several tasks.","tokens_in":23176,"tokens_out":7782,"duration_ms":86924,"significance":"If the reported benchmark results and pretraining hygiene are confirmed, this is a substantial contribution: a single vision-only SSL encoder that matches or exceeds several specialty foundation models across seven imaging modalities would be practically useful, and the paper is admirably candid about the tasks where the model does not win, such as pathology and dermatology. The authors also provide a relatively large pretraining corpus, a memory-augmented SSL recipe, and a downstream evaluation protocol with five independent runs, paired t-tests, confidence intervals, and effect sizes. However, the manuscript as submitted does not enumerate the pretraining corpus or provide an overlap audit, and several reported numbers are mutually inconsistent; these issues currently block verification of the central claim. The missing Supplementary Table 6 is not a cosmetic omission because the reported margin over DINO is only 0.002 AUROC.","major_comments":[{"comment":"The pretraining corpus is not enumerated. The Methods state that the 3.3 million images come from 'publicly available unlabelled datasets' and defer the list to Supplementary Table 6, but that table is absent from the submitted manuscript. Because the 25 evaluation datasets include widely used public collections such as RSNA, SIIM, APTOS2019, IDRiD, HAM10000, and BUSI, and because several of these are plausible components of the stated 401,059 dermoscopic and 333,700 fundus pretraining images, the manuscript cannot rule out that a fraction of evaluation images appeared in pretraining. Since the reported overall margin over DINO is 0.002 AUROC, even a small overlap could account for the headline difference. This is a load-bearing transparency issue: please provide the full corpus enumeration, a deduplication procedure, and an overlap analysis against every evaluation dataset.","section":"Methods — Pre-Training Dataset / Supplementary Table 6"},{"comment":"Several headline numbers disagree across the manuscript. The Abstract reports modality AUROCs of 0.943 (CT) and 0.858 (CXR), whereas Table 1 gives a CT public AUROC of 0.990 and a CXR public AUROC of 0.908; Table 1 also reports histopathology 1.000 while the main text reports a mean AUROC of 0.999. In addition, the Results state that MerMED-FM achieved a mean AUROC of 0.935 overall, but no table presents this overall mean or its confidence interval. Please reconcile these values and state exactly which experiment each number comes from.","section":"Abstract / Results / Table 1"},{"comment":"The OCTDL paragraph is internally inconsistent. The text says MerMED-FM 'achieved an AUROC of 0.911 (CI: 0.990-0.992)' and 'mean AUROC of 0.988' for OCT, but Table 2 lists the same OCTDL entry as 0.991 with CI 0.990-0.992. The printed 0.911 is incompatible with its own confidence interval and with Table 2. The sentence also appears grammatically incomplete ('where achieved an AUROC'). Please correct the value and recheck the surrounding OCTDL CIs and p-values.","section":"Results — Ocular disease diagnosis (OCTDL)"},{"comment":"The claimed statistical significance for the IQ-OTHNCCD comparison with BiomedCLIP is not supported by Table 3. The text states that MerMED-FM outperformed BiomedCLIP 'by 0.93% (p<0.01, T = 0.94, Cohen's d = 0.951)', but Table 3 reports for the same comparison a T statistic of 0.94 and p = 7.46e-01, which is not p<0.01. The same Table 3 row is garbled ('0.939 0.93 0.94 7.46E-01 0.951'), and several other rows contain AUROC-like numbers where mean differences are expected. Please repair Table 3 and ensure that every p-value quoted in the text is the one reported in the table.","section":"Results — CT lung carcinoma / Table 3"},{"comment":"The memory module is presented as a key contribution, but the only support for its benefit is the sentence 'The memory size was fixed at K = 65536, as determined through ablation studies.' No ablation study, table, or comparison with and without the memory module is reported anywhere in the manuscript. Since the memory size K, block size Nb, and temperature schedules are design choices, the reader cannot verify that these choices, rather than pretraining data composition or fine-tuning protocol, drive the reported results. Please include the ablation evidence or temper the claim accordingly.","section":"Methods — Model Architecture (memory module)"}],"minor_comments":[{"comment":"Please standardize the baseline name ('Dino' vs 'DINO') and specify the exact architecture and checkpoint used for each baseline, including patch size and pretraining dataset, since the comparison is load-bearing and the current text refers only to 'Dino' via a general reference.","section":"Throughout"},{"comment":"There are numerous typographical slips, including 'its to leading', 'with a AUROC', 'carder diagnosis', 'Supplemantary', 'carcioma', and 'both both'. Please run a careful proofreading pass.","section":"Throughout"},{"comment":"The normalization formula is malformed: it reads '3𝑥−𝑚𝑖𝑥(𝑥)8' and '0.2+ ... 0.8max...' with unbalanced parentheses. Please rewrite it in standard mathematical notation.","section":"Figure 1"},{"comment":"The statement 'Additional data may reasonably be requested from the corresponding author' is insufficient for reproducibility; please specify whether the pretraining corpus list, fine-tuning data splits, random seeds, and model weights will be released.","section":"Data Availability"},{"comment":"The Conclusion states that MerMED-FM 'outperforms existing single-modality and multi-specialty models', which is stronger than the Discussion's own limitations, where the authors concede that the model did not outperform UNI in pathology and PanDERM in dermatology. Please qualify the claim to match the reported results.","section":"Conclusion / Discussion"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my read. The paper reports a competently built SSL vision foundation model across seven medical modalities, evaluated on a large set of public and local datasets. The actual new content is the specific model and the benchmark suite, not the method: teacher-student SSL with a FIFO memory queue is a known combination of DINO/BYOL/MoCo ideas. That is fine—medical imaging needs these empirical tests—but the novelty is incremental.\n\nThe strongest part is the evaluation breadth: 25 public datasets plus a Singapore hospital cluster, with data-efficiency runs at 10-50% and comparisons against specialty models. If the results hold, this is a genuinely useful data point for hospital deployment.\n\nThe problem is that the results cannot be audited as written. The pretraining corpus is described as 3.3M images from 'publicly available unlabelled datasets' but the list is deferred to Supplementary Table 6, which is not present here. Since several evaluation sets (RSNA, SIIM, APTOS, HAM10000, BUSI) are public and plausible pretraining sources, you cannot rule out leakage. The headline margin over Dino is 0.002 mean AUROC; a small overlap could erase it. That is a load-bearing transparency issue, not a cosmetic one.\n\nThere are also internal inconsistencies: the abstract CT AUROC is 0.943 vs Table 1's 0.990; CXR is 0.858 vs 0.908; the OCTDL entry prints AUROC 0.911 with CI 0.990-0.992; and a T=0.94 is labeled p<0.01, which is unlikely. These are probably typos, but they undermine confidence in the reported numbers. No code or weights are released, and the memory module's claimed ablation (K=65536) is mentioned but not shown.\n\nWho should read this? Practitioners working on medical SSL or foundation models will find the benchmark useful, and the external validation is a plus. It deserves a serious referee because the empirical claim is important and the method is coherent. But I would not accept the current numbers without the pretraining list, an overlap audit, and a corrected table.\n\nMy recommendation: send it to review, but with the explicit requirement that the authors disclose the pretraining sources and rule out overlap. If they do, this could be a solid contribution.","headline":"Broad medical SSL benchmark that deserves a referee, but missing pretraining provenance and internal number inconsistencies make the headline AUROC claims uncheckable as written.","tokens_in":23755,"tokens_out":2795,"would_cite":false,"duration_ms":32042,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single self-supervised vision model, pretrained on 3.3 million unlabelled images across seven modalities, reports a mean AUROC of 0.935 and matches or beats specialty-specific medical imaging models on most tasks tested.","keywords":["medical imaging foundation model","self-supervised learning","multimodal imaging","multi-disease classification","memory module","vision transformer","transfer learning","AUROC benchmarking"],"falsifier":"Run an exact-duplicate and perceptual near-duplicate search between the 3.3 million pretraining images and the test images of the public evaluation sets (RSNA, SIIM, TBX11K, IQ-OTH/NCCD, SARS-COV-2, BUSI, BreakHis, APTOS2019, IDRiD, HAM10000, Dermnet, OCTID, and OCTDL). If even a small number of evaluation images, or near-duplicates of them, appear in the pretraining pool, the central comparison is no longer evidence for the model's generalization.","tokens_in":22672,"feed_emoji":"🩻","tokens_out":9908,"duration_ms":102956,"temperature":0.7,"pith_summary":"MerMED-FM is an attempt to answer a practical question: can one vision model, trained without labels, serve radiology, pathology, ophthalmology, ultrasound, and dermatology at once? The paper reports yes: a single ViT-B encoder pretrained on 3.3 million unlabelled images across seven modalities and over ten specialties reaches a mean AUROC of 0.935 across 25 public and local evaluation datasets, edging out the multimodal BiomedCLIP (0.919) and general-purpose DINO (0.933) and matching or beating specialty models in CT, CXR, OCT, CFP, and ultrasound tasks. The authors frame the contribution as a step toward a deployable all-in-one imaging model that needs no text prompts and far less labelled data for new tasks. If the result holds, it would mean hospitals could run one imaging model across departments rather than a stack of specialty-specific systems.","feed_headline":"One vision model matches specialty AIs across seven scans","feed_subtitle":"MerMED-FM hits 0.935 mean AUROC after self-supervised pretraining on 3.3M unlabelled medical images.","key_machinery":"The central mechanism is a memory-augmented joint-embedding teacher-student self-supervised learning loop. A student ViT and a teacher ViT encode multiple augmented views of each image; only the student is updated by gradients, while the teacher is an exponential moving average of the student. A dynamic memory module, a non-differentiable FIFO store of 65,536 compact representation vectors, holds recent embeddings from all modalities. The student's views are compared with stored representations, and the model enforces that similarity distributions over memory blocks agree across views. Balanced modality- and specialty-aware batch sampling prevents chest X-rays or other large modalities from dominating training. This combination is what the paper credits with stabilising joint training, preventing mode collapse and catastrophic forgetting, and letting a single encoder learn transferable features across seven modalities without labels.","core_discovery":"The paper's central claim is that a self-supervised, vision-only foundation model can be genuinely multimodal and multi-disease without the usual label burden. MerMED-FM is a ViT-B encoder pretrained on 3.3 million unlabelled images spanning CT, chest X-ray, ultrasound, histopathology patches, colour fundus photography, OCT, and dermoscopy. After fine-tuning with a small task-specific head, it reports the best mean AUROC among the compared models, 0.935, above BiomedCLIP's 0.919 and DINO's 0.933, with per-modality AUROCs of 0.988 (OCT), 0.982 (pathology), 0.951 (ultrasound), 0.943 (CT), 0.931 (skin), 0.894 (CFP), and 0.858 (CXR). The paper's authors interpret this as showing that a shared visual encoder trained with a memory-augmented self-supervised objective can transfer across specialties and match or beat models built for one modality, while explicitly noting that it did not surpass the pathology or dermatology specialists and has not yet been tested on true same-patient multimodal reasoning.","pith_inferences":["The paper's 'multimodal' claim is about a shared encoder, not integrated multimodal reasoning; it explicitly has not tested pairing, say, a patient's CT and biopsy in one decision. A natural next experiment is to feed same-patient images from two modalities and check whether diagnostic accuracy improves over either modality alone.","A testable consequence of the memory module is continual learning: because stored representations are FIFO and sampling is balanced, the same training loop could in principle absorb a new imaging modality later without full retraining; the paper does not demonstrate this, but the mechanism invites it.","If a leakage screen comes back clean, the result would support a stronger inference than the paper draws: text supervision and paired image-text data may be unnecessary for strong medical imaging encoders, with broad unlabelled modality coverage plus self-supervision doing the work."],"forward_implications":["If the central claim holds, a hospital could deploy one vision encoder for CT, CXR, ultrasound, pathology, fundus photography, OCT, and dermoscopy instead of a separate specialized model per department, simplifying maintenance and oversight.","The self-supervised pretraining route would lower the annotation barrier: with 10–50% of the fine-tuning data, MerMED-FM retains most of its full-data AUROC, so new diseases and modalities could be added more quickly in low-resource settings.","Because MerMED-FM needs no text prompts, it can be applied to raw images at acquisition time, before a report exists, which suits triage and screening workflows.","The claim is not that one model beats every specialist: the paper reports parity with but not superiority over the pathology and dermatology specialists, so the all-in-one benefit is bought with some specialty-specific trade-offs."],"supporting_citations":[{"why":"Supplies the teacher-student self-supervised protocol (augmented views, EMA teacher, projection heads) that the pretraining loop builds on.","marker":"[53]"},{"why":"Supplies the DINO general-domain baseline used in every benchmark and the self-distillation design behind the projection heads.","marker":"[23]"},{"why":"Supplies the BiomedCLIP multispecialty image-text baseline; the paper reports mean AUROC comparisons against it across all seven modalities.","marker":"[22]"},{"why":"Supplies the RETFound ophthalmology specialist baseline for OCT and CFP comparisons.","marker":"[8]"},{"why":"Supplies the Merlin CT specialist baseline whose lung-cancer and COVID-19 AUROCs MerMED-FM claims to surpass.","marker":"[7]"},{"why":"Supplies the RadDINO CXR specialist baseline for pneumonia, pneumothorax, and tuberculosis.","marker":"[10]"},{"why":"Supplies the USFM ultrasound specialist baseline for breast cancer on three public datasets.","marker":"[17]"},{"why":"Supplies the PanDERM dermatology specialist baseline on four skin-lesion datasets.","marker":"[18]"},{"why":"Supplies the UNI pathology specialist baseline on breast and tissue-type classification.","marker":"[19]"}],"fun_headline_variants":["One vision model, seven scans, no labels: MerMED-FM","MerMED-FM hits 0.935 AUROC across 7 imaging modalities","Self-supervised AI learns one model for CT, X-ray, ultrasound, and more","MerMED-FM: 3.3M unlabelled images teach a single multi-specialty AI","Multimodal medical AI matches specialists without hand-labelled data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the 3.3 million pretraining images, taken from publicly available unlabelled datasets named only in Supplementary Table 6, are disjoint from the 25 public evaluation datasets; if any evaluation images were also in pretraining, the reported AUROC improvements would be inflated by leakage.","fun_headline_variants_meta":{"raw":{"variants":["One vision model, seven scans, no labels: MerMED-FM","MerMED-FM hits 0.935 AUROC across 7 imaging modalities","Self-supervised AI learns one model for CT, X-ray, ultrasound, and more","MerMED-FM: 3.3M unlabelled images teach a single multi-specialty AI","Multimodal medical AI matches specialists without hand-labelled data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000458,"raw_usage":{"total_tokens":2339,"prompt_tokens":1027,"completion_tokens":1312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":1205}},"tokens_in":643,"tokens_out":1312,"duration_ms":13846,"temperature":1.0,"reasoning_tokens":1205,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:22:04.642821+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an exact-duplicate and perceptual near-duplicate search between the 3.3 million pretraining images and the test images of the public evaluation sets (RSNA, SIIM, TBX11K, IQ-OTH/NCCD, SARS-COV-2, BUSI, BreakHis, APTOS2019, IDRiD, HAM10000, Dermnet, OCTID, and OCTDL). If even a small number of evaluation images, or near-duplicates of them, appear in the pretraining pool, the central comparison is no longer evidence for the model's generalization.","supporting_citations":[{"cited_title":"*These authors contributed equally to this work as joint first authors †Corresponding author Corresponding Author: Assoc Prof","cited_arxiv_id":null,"evidence_quote":"Supplies the DINO general-domain baseline used in every benchmark and the self-distillation design behind the projection heads."},{"cited_title":"2; Supplementary Table 5)","cited_arxiv_id":null,"evidence_quote":"Supplies the RETFound ophthalmology specialist baseline for OCT and CFP comparisons."}],"review_version":1}