{"id":"a2556580-ab03-4182-aeb6-1bf040b2a677","arxiv_id":"2411.18730","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of foundation models in radiology, covering their definition, training data and methods, capabilities, evaluation, and risks, and outlining a roadmap for responsible deployment.","lead":"Foundation Models in Radiology is a review that explains what AI foundation models are in medical imaging and how they are trained, adapted, and evaluated. It also lays out the benefits and risks, aiming to align technical advances with clinical needs.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposed 'standardized terminology' is not operational: Section 2's four defining properties are applied inconsistently to image-only and non-emergent radiology FMs in Sections 3 and 5.","rationale":"The reader's weakest assumption concerned transfer of scaling laws from general NLP/vision to radiology, with the caveat that terminology and risk catalog would remain useful even if scaling failed. That is a reasonable secondary concern, but the paper's primary deliverable is the standardized terminology itself. The load-bearing issue is that this terminology is not internally consistent: Section 2's four properties are not applied uniformly to models the paper itself calls FMs, and the notion of emergent abilities is presented as settled despite a well-known controversy. A review whose stated purpose is to provide shared terminology should first be internally coherent and explicit about which properties are definitional. This is a correctable but substantive flaw, so the appropriate verdict is CONDITIONAL acceptance pending a revision that clarifies the definition and its boundary conditions. The proposed test is straightforward: apply the Section 2 criteria to a representative set of models and see whether the classification matches field usage and annotator agreement. If it does, the concern is resolved; if not, the central claim is weakened.","tokens_in":14152,"tokens_out":3430,"duration_ms":31851,"concrete_test":"Construct a decision procedure from Section 2's four properties and apply it to a list of 20-30 systems that the paper or prominent surveys call radiology FMs (e.g., Medical SAM, CheXagent, Med-PaLM-M, MAE-based medical models, RadImageNet transfer models). Have two independent annotators classify each as FM/not-FM under (a) all four required, (b) any two required. Measure inter-annotator agreement and the fraction of commonly accepted FMs excluded. If the all-four rule excludes more than 20% of models the field calls FMs, or if annotators cannot agree, the definition is not standardized. Section 2 would then need to state necessary vs. characteristic properties and address the emergence controversy (e.g., cite Schaeffer et al. and define operational criteria).","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is to establish a standardized terminology for radiology FMs. For that claim to hold, the proposed definition must be precise and applied consistently. Section 2 lists four properties: 'incorporating large-scale model architectures and training data, extracting knowledge from multiple data modalities, employing self-supervised training strategies..., and exhibiting emergent capabilities beyond their training objectives.' This reads as a conjunctive definition. However, Section 3 presents image-only pre-training methods (MAE, SimCLR, BYOL) as core FM training, and Section 5 calls Medical SAM a foundation model, a single-modality segmentation model with no emergent language ability. If the four properties are only illustrative rather than definitional, the paper does not tell readers which are necessary, sufficient, or weighted. If they are necessary, the terminology excludes many models commonly called radiology FMs; if they are optional, it is hard to distinguish FMs from conventional self-supervised models. Additionally, 'emergent abilities' is cited to Wei et al. without acknowledging that emergence may be a metric artifact (Schaeffer et al., 2023), so a key term in the proposed terminology is contested. This undermines the 'standardized' and 'unify' goals independent of whether scaling laws transfer.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a narrative review of foundation models (FMs) in radiology. It proposes a framework of four defining properties (large scale, multimodality, self-supervision, emergence), surveys pre-training and adaptation methods (generative and contrastive pre-training, zero-shot inference, fine-tuning, instruction tuning, RLHF), catalogs public datasets and their limitations, lists clinical capabilities and evaluation strategies, and discusses risks (hallucination, bias, automation bias, environmental cost) and future directions. The stated aim is to standardize terminology and align technical development with clinical deployment.","tokens_in":14397,"tokens_out":4989,"duration_ms":43932,"significance":"The review fills a gap for a consolidated, clinically oriented overview of FMs in radiology. It is well referenced, organizes a large literature into a coherent structure, and includes useful figures summarizing training and adaptation pathways. Its strengths include the emphasis on dataset requirements, evaluation benchmarks, and responsible deployment, as well as the broad coverage of both technical and ethical considerations. If the terminology were clarified, it could serve as an accessible entry point for radiologists and machine learning researchers. However, as currently written, the definitional inconsistency prevents the paper from fully achieving its central claim of a standardized terminology.","major_comments":[{"comment":"The paper's central claim of establishing a standardized terminology is not supported by the definitional structure. Section 2 introduces four properties as characteristics of FMs, including multimodality and emergent abilities. Yet Section 3.1 presents masked autoencoders (e.g., Medical MAE) and Section 3.2 presents SimCLR and BYOL as FM pre-training methods, and Section 5 explicitly refers to Medical SAM as a foundation model—all of which are single-modality models with no demonstrated emergent capabilities beyond their training objectives. If the four properties are intended as necessary, these models are excluded; if they are optional, the definition does not provide a basis for distinguishing FMs from conventional self-supervised models. The manuscript should clarify which properties are necessary, sufficient, or merely typical, or revise the terminology to be consistent with the models actually discussed.","section":"Section 2 vs. Sections 3 and 5"},{"comment":"The definition of emergent abilities cites Wei et al. (ref. 16) as an established phenomenon, but Schaeffer et al. (2023) have shown that apparent emergence in large language models can be an artifact of the evaluation metric used. Since emergence is one of the four defining properties in the paper's framework, presenting it without this caveat is a contested choice and weakens the standardization claim. The authors should either acknowledge the debate and justify their definitional stance or reframe emergence as a contested/optional property.","section":"Section 2, emergent abilities"}],"minor_comments":[{"comment":"The heading 'Human-Computer Reaction with FMs' should be 'Human-Computer Interaction with FMs'; 'Reaction' appears to be a typo.","section":"Section 6"},{"comment":"The statement that 'Recently released datasets (63) consist of around 20,000 CXRs' appears inconsistent with the cited reference (CANDID-PTX), which is a small curated collection. Please verify the dataset and the number, or correct the reference to an appropriate large dataset.","section":"Section 4.1, datasets"},{"comment":"The sentence 'Model performance follows power laws, consistently improving as model and data size increase' states a general-domain empirical observation as a universal fact; the authors should qualify this as an observation from NLP/general vision that may not transfer to radiology, where data sizes are smaller and distribution shifts are more severe.","section":"Section 2, scaling laws"},{"comment":"The description of BYOL as bringing 'the representations of the two views closer in the representation space' is correct but could be more precise by noting that BYOL uses a predictor network and stop-gradient to avoid collapse; a brief clarification would help readers unfamiliar with the method.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The review is comprehensive and likely to be useful after revisions. The main concern is the inconsistency between the proposed definition of foundation models and the models included in Sections 3 and 5; this directly affects the paper's stated goal of standardizing terminology. If the authors clarify the status of the four properties and add a caveat about emergent abilities, the paper would be suitable for publication. I would also recommend verifying the CANDID-PTX dataset description before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Your quick read on 2411.18730 is fair: this is a well-structured review that covers how radiology foundation models are built, adapted, evaluated, and where they fail. What is actually new is modest - the synthesis is the contribution - but for a reader new to the area, it is a genuinely useful map.\n\nThe paper does several things well. The training section is accurate: generative and contrastive pre-training, adapters, instruction tuning, RLHF are described correctly. The dataset section is concrete and helpful. The cautions are balanced, covering hallucination, automation bias, data bias, and cost. The reference list is broad, and I do not see citation padding or self-citation abuse.\n\nThe stress-test note lands. Section 2 gives four defining properties of a foundation model: large scale, multi-modality, self-supervision, and emergent abilities. But later sections treat single-modality, non-emergent models like MAE, SimCLR, BYOL, and Medical SAM as core examples. If those properties are meant as necessary conditions, the terminology excludes common radiology FMs; if they are merely typical features, the paper never tells the reader which are essential. That inconsistency undermines the claim of a \"standardized terminology.\" Also, the paper cites Wei et al. for emergent abilities without acknowledging Schaeffer et al.'s argument that emergence can be a metric artifact. For a term central to the proposed glossary, that is a real omission. The scaling-law claim is also stated flatly, though the review does not depend on it heavily, so that concern is minor.\n\nNone of this sinks the paper. The core content is accurate and the structure is sound. The definitional looseness is fixable with a sentence stating that the four properties are typical characteristics, not necessary and sufficient criteria, and a citation to the emergence debate. With that revision, the paper becomes a solid reference. I would send it to peer review. It is worth a serious referee's time.","headline":"A competent and useful review of radiology foundation models, but the 'standardized terminology' claim is undercut by an inconsistent definition of what counts as a foundation model.","tokens_in":666,"tokens_out":2169,"would_cite":false,"duration_ms":42394,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that radiology needs a shared foundation-model vocabulary, and supplies one together with training and evaluation pathways.","keywords":["foundation models","radiology","self-supervised learning","multimodal AI","evaluation benchmarks","responsible deployment","vision-language pre-training","clinical AI"],"falsifier":"A careful empirical study that trains radiology foundation models at increasing scale on the same multi-site data and finds that diagnostic accuracy plateaus or drops beyond a modest model size, while smaller task-specific models match them, would falsify the scaling-law premise the review relies on, although the terminology it proposes could still stand.","tokens_in":14003,"feed_emoji":"🩻","tokens_out":5089,"duration_ms":60512,"temperature":0.7,"pith_summary":"Foundation models trained on large unlabeled datasets have moved into radiology, but the field lacks shared terms for what counts as a foundation model, how one is trained, and how it should be tested. This review's central aim is to supply that standardized terminology: it defines a foundation model by four properties — large scale, multimodality, self-supervised training, and emergent capabilities — and organizes the literature around what, how, when, why, and why not to build radiology foundation models. A sympathetic reader would care because without a common vocabulary, models cannot be compared, regulatory discussions stall, and clinical deployment proceeds without agreed evaluation standards. The review also compiles the risks — hallucination, automation bias, data bias, cost — that any responsible deployment plan must address.","feed_headline":"Radiology foundation models get a shared playbook","feed_subtitle":"One review defines what they are, how to build and test them, and where they can fail.","key_machinery":"The central object is the taxonomy itself: the four-property definition of a foundation model — large-scale architecture and data, multimodality, self-supervision, and emergent abilities — inside a what, how, when, why, and why-not structure. This framework does the work of turning a scattered literature into a shared checklist: what counts as a foundation model, how training data and pre-training objectives determine capabilities, when datasets are sufficient, and why caution is needed. The same taxonomy is used to derive concrete data requirements (3D/4D and longitudinal multi-site datasets with demographics) and evaluation categories (coarse- and fine-grained tasks, visual question answering, generative similarity, human and model-based assessment, and subgroup bias analysis).","core_discovery":"The review's central claim is that existing and future radiology foundation models can be understood through a single descriptive framework, rather than as a list of isolated systems. It identifies four defining characteristics — large-scale architectures and data, multimodal data integration, self-supervised training, and emergent abilities — and describes the building blocks that produce them: modality-specific encoders, fusion modules, and multimodal decoders, trained generatively or contrastively. It then argues that pre-trained models are adapted to clinical tasks by zero-shot inference, linear probing, fine-tuning, instruction tuning, or reinforcement learning from human feedback, and that evaluation must cover discriminative accuracy, generative similarity, human ratings, and bias. The paper's conclusion is that radiology-specific foundation models are feasible and potentially valuable, provided training data are large, multi-site, and representative, and deployment is governed by explicit safeguards.","pith_inferences":["If the scaling-law assumption carries over to radiology, a natural next step would be compute-optimal training studies on radiology-specific data, since the review stops short of quantifying how much data or compute a target capability requires.","The terminology could serve as a shared language for regulatory bodies and payers, even though the review itself does not engage specific approval or reimbursement pathways.","A testable extension would be a benchmark suite that scores a radiology foundation model on all five capability categories and all three risk categories at once, making the review's separate checklists operational.","The claim that emergent abilities are clinically valuable depends on their reliability; a targeted evaluation of emergent zero-shot diagnostic capabilities across rare diseases would test that dependency."],"forward_implications":["Papers and regulators gain a common set of terms, so different radiology foundation models can be compared on the same axes: scale, modality coverage, training paradigm, adaptability, and evaluation.","Training radiology foundation models should prioritize large multi-site datasets spanning 3D/4D imaging, ultrasound, demographics, and longitudinal follow-ups, since current resources mostly cover 2D chest imaging.","Evaluation should combine automatic discriminative benchmarks, generative content-similarity measures, human radiologist review or interactive demos, and subgroup and bias analysis before clinical deployment.","Responsible deployment requires safeguards against hallucinated findings, automation bias and overreliance, anthropomorphism of model outputs, and cost-driven centralization of AI development.","Future research the review points to includes efficient architectures, privacy-preserving training such as federated learning and differential privacy, continual learning, and monitoring for data drift after deployment."],"supporting_citations":[{"why":"Supplies the evidence that biomedical foundation models show emergent zero-shot capabilities, such as tuberculosis prediction on chest X-rays without explicit training.","marker":"(8)"},{"why":"Supplies the scale of existing radiology image datasets, with over one million images from RadImageNet available for pre-training.","marker":"(10)"},{"why":"Supplies the main vision-language radiology dataset, MIMIC-CXR, with over 300,000 chest X-rays paired with free-text reports.","marker":"(11)"},{"why":"Supplies the contrastive vision-language pre-training method, CLIP, that radiology models adapt for image-text alignment.","marker":"(42)"},{"why":"Supplies an instruction-tuned radiology foundation model and the evaluation approach referenced for capabilities and limitations.","marker":"(22)"},{"why":"Supplies the definition and source of emergent abilities in large models, which the review uses as a defining property of foundation models.","marker":"(16)"},{"why":"Supports the caution that foundation models can encode data disparities and societal bias in chest radiography.","marker":"(77)"}],"fun_headline_variants":["Radiology foundation models get a unified framework","Four pillars of radiology foundation models","Building radiology AI: a foundation model guide","A roadmap for safe radiology foundation models","What radiology foundation models can and cannot do"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's load-bearing premise is that the scaling laws and self-supervised methods that work on web-scale general data also transfer to radiology, where datasets are smaller, labels are expensive, and imaging varies across sites and machines.","fun_headline_variants_meta":{"raw":{"variants":["Radiology foundation models get a unified framework","Four pillars of radiology foundation models","Building radiology AI: a foundation model guide","A roadmap for safe radiology foundation models","What radiology foundation models can and cannot do"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00028,"raw_usage":{"total_tokens":1628,"prompt_tokens":879,"completion_tokens":749,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":681}},"tokens_in":495,"tokens_out":749,"duration_ms":6537,"temperature":1.0,"reasoning_tokens":681,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:55:56.162710+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A careful empirical study that trains radiology foundation models at increasing scale on the same multi-site data and finds that diagnostic accuracy plateaus or drops beyond a modest model size, while smaller task-specific models match them, would falsify the scaling-law premise the review relies on, although the terminology it proposes could still stand.","supporting_citations":[],"review_version":1}