{"id":"4be36a4c-cc2a-4804-9029-54c379d6fabb","arxiv_id":"2506.13306","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of brain imaging foundation models covering 86 models and 161 datasets, with a performance tournament, dataset atlas, and duplicated-data warnings.","lead":"This paper systematically reviews 86 AI foundation models and 161 datasets built for brain scans. It maps which models win on which tasks, flags repeated datasets that can leak into training, and highlights under-served areas like PET imaging and mental health.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"","rationale":"The reader's weakest assumption and my concern coincide: the Section 6 rankings assume that self-reported evaluation numbers from different papers, datasets, and metrics are comparable. This is genuinely load-bearing because the abstract and introduction advertise identification of the leading models per task, and the paper itself concedes in Section 8 that the meta-analysis relies entirely on reported figures and claims. The concern does not invalidate the review's other contributions: the curated dataset inventory, PRISMA-style screening, architecture and training taxonomy, and the honest discussion of biases and evaluation gaps are useful regardless of the rankings. However, the 'leading models' claim should be read, and ideally reworded, as 'reported leaders under heterogeneous evaluation protocols', not as an established performance ranking. Since the reader already gave a CONDITIONAL verdict with moderate confidence and identified the same comparability issue, my stress-test does not move the verdict; it reinforces the condition. The concrete test would settle whether the specific named leaders survive a like-for-like comparison, and until that is run, the conditional framing is appropriate.","tokens_in":36549,"tokens_out":2107,"duration_ms":24207,"concrete_test":"","verdict_should_be":"UNCHANGED","load_bearing_attack":"","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a systematic literature review of foundation models (FMs) applied to brain imaging. Following a PRISMA-based protocol (Section 3), the authors screened 541 records from Semantic Scholar, Crossref, and PubMed, augmented the list via surveys and snowballing, and included 86 FM architectures and 161 brain imaging datasets in their analysis. The review maps general trends (publication timeline, venues, code availability), architectural choices (backbones, encoders), training strategies (contrastive/masked learning, adapters/LoRA, MoE), task and pathology coverage, and a dataset inventory with modality, license, and deduplication statistics. It also proposes a 'tournament' approach (Section 6.1) and benchmark tables (Tables 2-3) to identify the best-performing models per task, and discusses pitfalls including bias, pathology imbalance, and evaluation weaknesses. The central claims are that this is the first comprehensive curated review of brain-imaging FMs, that it systematically analyzes 161 datasets and 86 model architectures, and that it highlights the leading models for each task.","tokens_in":36458,"tokens_out":9101,"duration_ms":87063,"significance":"If the inventory and trend analysis are accurate, the paper provides a valuable consolidated map of an active, fast-moving field. The dataset analysis—covering deduplication, licensing, anonymization, modality and pathology distribution—is particularly useful, as is the model analysis of backbone reliance, parameter-efficient fine-tuning, and task coverage. The explicit search protocol and the public interactive atlas on Notion are strengths that support reproducibility and updateability. However, the leaderboard component is not yet methodologically reliable: the tournament and benchmark tables rank models across different BraTS versions, different metric types (ACC vs F1 vs BERT similarity), and different evaluation conditions, using self-reported numbers. The paper itself concedes this in Section 8 ('our synthesis relies entirely on the figures and claims reported in the reviewed literature'), yet the 'best model' statements remain a prominent contribution. The citation-based 'not outperformed' criterion in Section 6.1 is defensible only as a literature-status snapshot, not as a performance ranking.","major_comments":[{"comment":"The statement 'The results show that the best performing models are MoME and BrainSegFounder' is inconsistent with the numbers in Table 2, where MAE-Seg Africa24 (92.90) and OBJ-SAM (91.90) report higher DICE scores than either MoME (92.10) or BrainSegFounder (91.15). The caveats listed in the text (missing contrasts, small training subsets, population shift) do not explain the omission of MAE-Seg and OBJ-SAM from the 'best' designation, nor is any explicit exclusion criterion given. Please specify the exact inclusion and exclusion rules for the leaderboard (e.g., benchmark version, full-contrast evaluation, metric type) and apply them consistently to all rows of the table.","section":"Section 6.2, Table 2"},{"comment":"The rankings mix incomparable quantities. In Table 3, most rows report Accuracy, but RadFM and Med-PaLMM report F1, Med-Flamingo reports a BERT similarity score, and LLaVA-Med's open-ended column is Recall; in Table 2, BraTS versions from 2013 through 2024 are combined, and the text notes that some DICE values were converted from AUC. Ranking these together and then asserting which model is 'best' (e.g., 'Med-VLP achieving the best accuracy', 'MUMC achieving the best performance', 'RadFM achieving the best score of 78.09') is not methodologically sound. Because the identification of leading models is a central contribution (Abstract, Section 1, Section 6), please either (a) restrict all leaderboard claims to same-version, same-metric comparisons, (b) present the tables without any cross-metric 'best' statements, or (c) add a sensitivity analysis that omits non-comparable entries and shows whether the rankings are stable.","section":"Section 6.2, Table 3 and Section 6.1"},{"comment":"The text says 'We restrict our research to the last 5 years' for the database search on 18 February 2025, which would exclude papers from 2019. Yet Section 4.1 includes Genesis and Med3D, both published in 2019, as the first models in scope. This is presumably because the augmentation step (surveys and snowballing) was not time-restricted, but the manuscript does not say so. Please state explicitly that the five-year window applies only to the database query and not to the augmentation/snowballing step, and adjust the wording in Section 4.1 accordingly.","section":"Section 3.1"},{"comment":"The single-reviewer screening is disclosed in the Limitations section, but Section 3 claims adherence to PRISMA 2020, which recommends at least two independent reviewers for screening. This deviation is load-bearing for the reliability of the included-study list. Please move or repeat the disclosure in the Methodology section, describe any mitigation (e.g., verification of a random subset by a second author), and discuss the potential impact of screening bias on the final set of 86 included models.","section":"Section 8 and Section 3"}],"minor_comments":[{"comment":"The sentence 'Among the 32 vision-language brain FMs included in our study, 34% are based on CLIP [167, 42, 86, 40, 17, 122, 118, 70]' is arithmetically inconsistent with the cited list, which contains 8 items: 8/32 is 25%, not 34%. Please correct this percentage and audit nearby percentages in Section 4 for consistency.","section":"Section 4.2"},{"comment":"The captions for Fig. 2 ('Cumulative number of FMs publications over the years') and Fig. 3 ('Most cited FM over the time period') appear to be mismatched with their content: the text around Fig. 2 discusses cumulative publication counts, while Fig. 3 shows a task-distribution chart that resembles Fig. 7. Please verify all figure-caption pairings.","section":"Figures 2 and 3"},{"comment":"The word 'interative' should be 'interactive'. In addition, the Notion links are hosted on a third-party platform; given the review's reproducibility claims, consider archiving the interactive atlas (e.g., in Zenodo or as a supplementary PDF) so the data behind Tables 1-3 remains accessible.","section":"Section 3.3"},{"comment":"The text states that DICE scores were 'converted from AUC if the original only uses AUC metrics', but no conversion formula or reference is provided. AUC-to-DICE conversion is not a standard, well-defined operation; please either justify it with a citation or remove the converted entries from the leaderboard.","section":"Section 6.2"},{"comment":"Several typos and grammatical slips should be corrected in a final pass: 'metholodgy' (Section 3), 'keywods' and 'thee' (Section 3.1), 'anomymization' (Sections 5.2 and 8), 'Traditionnaly' (Section 7.1), 'dependant' (Section 4.2), 'speciliaized' (Section 7.2), and 'demyelating' (Section 5.2).","section":"Global"}],"recommendation":"major_revision","confidential_remarks":"The paper's main value is the dataset/model inventory and the trend analysis, not the leaderboard, which is built on heterogeneous and self-reported metrics. The internal inconsistency in Table 2 (the 'best' statement contradicts the table numbers) is a concrete issue that must be fixed. I also suggest the authors verify the claim of being 'the first comprehensive curated review' against any very recent competing arXiv surveys before resubmission, and consider replacing the Notion links with a permanent archive. The revision should be manageable within the current scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about: this is the first systematic map I've seen dedicated to brain imaging foundation models, and it's genuinely useful as a reference. The authors catalog 86 FMs and 161 datasets, split the datasets by 2D/3D, modality, pathology, and access, and flag that about 6% of 3D imaging studies come from duplicate or derivative datasets — a concrete warning that should be heeded by anyone training on mixed brain MRI pools. The interactive Notion tables are a plus, and the PRISMA flow is reported, though not perfectly executed.\n\nWhat's new is the dataset-level synthesis and the per-task/per-benchmark leaderboard attempts. The pathology-access imbalance (cancer-heavy FM work vs. abundant Alzheimer/Parkinson data) and the PET/mental-health gaps are well argued and actionable. The limitations section is honest: single-reviewer screening, reliance on self-reported numbers, and the comparability problem are all admitted.\n\nSoft spots, in order. First, the tournament rankings in Section 6 are the weakest part. Table 2 mixes BraTS 2018, 2019, 2021, and 2023 with different amounts of training data, and some scores are converted from AUC to Dice; the authors acknowledge the issue but still put 'best models' in the headline. I'd treat those rankings as illustrative, not authoritative. Second, the screening was single-reviewer, and the PRISMA numbers have a small arithmetic inconsistency (excluded counts sum to 451, text says 449). Third, there are more typos and jargon slips than I'd like ('healthcare care', 'metholodgy', 'anomymization'), which suggests a rushed final pass, but nothing that changes the substance. The LoRA/MoE equations are standard quotations, so no circularity concern.\n\nOverall, the central claim — that brain FM evaluation lacks standardized benchmarks and that the literature is skewed — holds up. The paper is a reference resource, not a source of new measurements. I'd send it to peer review with a request to soften the 'best model' language, add a comparability matrix for benchmark versions, and fix the inconsistencies. For anyone starting in this area, it's worth having on the desk.","headline":"A useful first map of brain imaging foundation models with an honest limitations section; the model rankings are illustrative rather than authoritative.","tokens_in":37020,"tokens_out":2248,"would_cite":true,"duration_ms":22516,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review systematically maps 86 foundation models and 161 datasets for brain imaging and identifies the leading models per task, the first consolidated survey of the field.","keywords":["brain imaging","deep learning","foundation models","brain cancer","neurodegenerative diseases","neurovascular diseases","systematic review","benchmarking"],"falsifier":"Re-run the models the review crowns as leaders (e.g., MoME, BrainSegFounder, Med-VLP, MUMC, RadFM) on one shared benchmark — BraTS 2021 for segmentation and VQA-RAD for question answering — under identical preprocessing and metrics; if the tournament order changes materially, the paper's 'best model' conclusions fail.","tokens_in":36297,"feed_emoji":"🧠","tokens_out":8307,"duration_ms":73246,"temperature":0.7,"pith_summary":"This paper claims to be the first systematic review devoted entirely to foundation models for brain imaging, spanning 161 datasets and 86 model architectures. It maps the field's design choices, training strategies, and tasks, and then tries to identify which models lead on each major benchmark. The payoff for a reader is a consolidated atlas of a scattered literature: it shows where brain foundation models have concentrated effort (MRI-based cancer segmentation), where they have neglected available data (PET imaging, mental-health conditions), and where evaluation is too inconsistent to support confident comparisons. If the review is right, it gives both clinicians and machine learning researchers a shared starting point and a list of blind spots to fix.","feed_headline":"86 brain-imaging foundation models mapped across 161 datasets","feed_subtitle":"It names per-task leaders, yet finds inconsistent benchmarks and neglected PET and mental-health data.","key_machinery":"Two devices carry the analysis. First, a tournament graph of reported performance: a directed edge from model A to model B means B beat A in at least one publication, and 'green' nodes mark peer-reviewed models with over 50 citations that no surveyed approach has yet outperformed; this is what produces the lists of leading models. Second, benchmark tables that line up self-reported scores on the few common testbeds — BraTS for segmentation and VQA-RAD for question answering — where the review can compare variants side by side. The inclusion criterion that defines a foundation model (training on at least two imaging modalities, two organs, or two brain pathologies) is what lets the review sweep in the 86 architectures in the first place.","core_discovery":"The central discovery the authors claim is that brain imaging foundation models now form a recognizable but uneven research field: 86 architectures, most built on a handful of backbones (SAM, CLIP, U-Net, BERT), with a burst of growth after 2023. On benchmarks, the review crowns specific leaders — MoME and BrainSegFounder on BraTS segmentation, Med-VLP and MUMC on VQA-RAD question answering, RadFM among generative multimodal systems — while cautioning that the comparisons inherit whatever datasets and metrics the original papers chose. It also documents structural problems: about six percent of 3D imaging studies are duplicates or derivatives, creating leakage risk; only six models address demographic bias; only seven include human evaluation; and pathology coverage in models diverges from both disease prevalence and dataset availability. The authors therefore frame the review as a snapshot and a call for standardized benchmarks, broader pathology coverage, and clinically grounded evaluation.","pith_inferences":["A fair side-by-side re-evaluation on fixed splits could reshuffle the tournament leaders, since the ranking is built from heterogeneous self-reported numbers; the paper's method naturally extends into a living leaderboard.","The mismatch between pathology prevalence and model coverage suggests that future dataset contributions (e.g., PET and mental-health cohorts) may advance the field more than new architectures.","The bias-aware practices of the six models that address demographic balance provide a testable template: requiring stratified performance reporting for every brain FM would let the field verify fairness claims.","The inclusion criteria treat models trained on two modalities, organs, or pathologies as foundation models, so the review's 'foundation model' category spans very different scales; separating genuinely large pretrained systems from smaller multi-task models might change the conclusions."],"forward_implications":["Researchers can use the 161-dataset, 86-model atlas as an entry point when choosing backbones and benchmarks for a new brain imaging project.","The tournament singles out a short list of leaders (such as MoME and BrainSegFounder for segmentation, Med-VLP and MUMC for question answering) that new methods should be compared against.","The finding that about six percent of 3D imaging studies come from duplicate datasets warns that multi-dataset training should include explicit deduplication to avoid leakage.","The scarcity of standardized benchmarks, with no dataset used by more than five models outside the two dominant ones, implies that the field needs shared evaluation protocols before progress becomes measurable.","Because only seven of 86 models included any human expert evaluation, clinical translation claims rest on algorithmic metrics that may not reflect diagnostic utility."],"supporting_citations":[{"why":"Supplies the systematic literature review procedure that structures the search and screening.","marker":"[69]"},{"why":"Supplies the reporting and screening checklist used to document the review's inclusion decisions.","marker":"[104]"},{"why":"Defines the foundation model criteria the review adapts for inclusion.","marker":"[14]"},{"why":"A large healthcare foundation model review with only 32 of 200 papers on brain imaging, establishing the gap the paper fills.","marker":"[55]"},{"why":"A trustworthiness-focused survey with only ten of 76 papers on brain imaging, supporting the claim that brain imaging is underrepresented.","marker":"[114]"},{"why":"A CLIP-focused survey covering only seven brain foundation models, evidence that even model-family reviews miss the area.","marker":"[163]"},{"why":"A SAM-focused survey identifying only nine brain foundation models, completing the gap argument across review types.","marker":"[161]"}],"fun_headline_variants":["Brain imaging AI: 86 models, but benchmarks don’t add up","86 foundation models for brain scans, yet gaps remain","Brain foundation models: leaders named, standards lacking","161 datasets, 86 models, one big benchmarking mess","Brain imaging AI: growth spurt after 2023, but weak evaluation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings assume that the numbers reported by different papers, measured on different dataset versions and with different metrics, can be compared side by side; the paper itself concedes in Section 8 that the synthesis relies entirely on published claims.","fun_headline_variants_meta":{"raw":{"variants":["Brain imaging AI: 86 models, but benchmarks don’t add up","86 foundation models for brain scans, yet gaps remain","Brain foundation models: leaders named, standards lacking","161 datasets, 86 models, one big benchmarking mess","Brain imaging AI: growth spurt after 2023, but weak evaluation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000756,"raw_usage":{"total_tokens":3354,"prompt_tokens":935,"completion_tokens":2419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":2333}},"tokens_in":551,"tokens_out":2419,"duration_ms":17338,"temperature":1.0,"reasoning_tokens":2333,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:04:25.423214+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the models the review crowns as leaders (e.g., MoME, BrainSegFounder, Med-VLP, MUMC, RadFM) on one shared benchmark — BraTS 2021 for segmentation and VQA-RAD for question answering — under identical preprocessing and metrics; if the tournament order changes materially, the paper's 'best model' conclusions fail.","supporting_citations":[],"review_version":1}