{"id":"bc92bd65-c671-4d12-94d7-f9983f70475d","arxiv_id":"2502.09242","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.","lead":"This paper is a structured review of 144 recent studies on generative AI in medicine, mapping how the field is moving from text-only language models to multimodal systems that combine images, text, and structured data. It is useful as a map of current methods, datasets, and evaluation metrics for anyone entering or funding this space.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed shift from unimodal to multimodal AI is not supported by the described methods: 58% of included papers entered via an unspecified manual search plus deliberate topical balancing, so the trend is an artifact of selection.","rationale":"The reader flagged representativeness of the curated literature as the weakest assumption; I agree and sharpen it. The concern is not merely that the manual search is non-reproducible, but that the inclusion criteria explicitly select on the outcome variable: 'proportional inclusion from prevalent fields' (Section 2.4) and the two-part database search (Section 2.3) both guarantee that multimodal content is well represented. A scoping review that aims to describe a field cannot then use the same selection to demonstrate that the field is shifting toward multimodal. The review does list limitations, noting overrepresentation of radiology, but it does not acknowledge that the balancing step directly undermines the central trend claim. The concrete test isolates the database-only subset; if the trend disappears, the claim must be revised. This does not invalidate the rest of the review; the method/dataset/evaluation tables are valuable. Therefore I recommend keeping the reader's CONDITIONAL verdict: the paper should either remove or substantially soften the 'shift' claim, or provide a reproducible protocol that does not select on the outcome. I see no reason to move to REJECT because the descriptive content is solid.","tokens_in":27047,"tokens_out":3894,"duration_ms":34680,"concrete_test":"Re-run the trend analysis using only the 60 database-derived papers, with explicit inclusion/exclusion criteria and no manual balancing, and report the unimodal vs multimodal distribution over time before any curation. If the unimodal-to-multimodal shift is absent or reversed in this subset, the central claim is an artifact of the manual search and balancing step. Also reconcile the PRISMA counts in Section 3 and Fig. 2 to confirm whether 83 or 84 manual papers were actually included.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim, stated in the abstract as 'a shift from unimodal to multimodal approaches,' rests entirely on the composition of the 144 included papers. The search protocol (Section 2.3) was constructed as two targeted subsearches: one for text-only LLMs and one for multimodal models. The retrieved set is therefore a union of two categories, not a sample from which prevalence can be inferred. More decisively, 83 of 144 papers (58%) came from a manual search described only as capturing 'recent and high-impact publications' (Section 2.3), and Section 2.4 states 'we aimed for proportional inclusion from prevalent fields, such as X-ray report generation.' Deliberately balancing inclusion on topic prevalence makes the observed shift a property of the curation rule, not of the literature. The PRISMA flow (Fig. 2) also disagrees with the text: it shows 84 records from 'other methods' (1 website + 83 citations) and 144 total, while the text says 83 manual papers, which would give 143 total. This off-by-one inconsistency further weakens confidence in the quantitative-looking trend statements. The paper's qualitative mapping of methods, datasets, and evaluation metrics is useful and independent of the trend claim, but the headline finding is not supported by the reported methodology.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This scoping review (arXiv:2502.09242) applies PRISMA-ScR guidance to survey generative AI in medicine, covering text-only LLMs, multimodal CLIP and MLLM architectures, associated datasets, and evaluation metrics. The authors report screening 4,384 database records and 84 additional records from manual searches, ultimately including 144 papers. The review synthesizes these papers in six tables organized by application, dataset, and evaluation approach, and the abstract and discussion make a central claim of a field-level 'shift from unimodal to multimodal approaches' in medical AI. The paper also describes persistent challenges such as data heterogeneity, interpretability, and evaluation gaps.","tokens_in":27259,"tokens_out":3179,"duration_ms":31429,"significance":"If the reported trend claim were reliable, the review would offer a valuable, up-to-date map of a rapidly moving field, with useful taxonomies of models, datasets, and clinical evaluation metrics. Its strengths include the reproducible database queries in Supplementary Table S.2, the explicit PRISMA-ScR checklist, and a structured categorization of 144 primary papers into coherent categories (Tables 1-6). The qualitative discussion of evaluation metrics (Section 6) is informative and reasonably current, and the identification of radiology-heavy dataset bias is a credible observation. However, the headline 'shift' claim depends on a literature-selection process that includes an unspecified manual search and a deliberate balancing step, which the reported methodology does not adequately support as a population-level trend. The review's descriptive synthesis remains useful regardless, but the central claim needs substantial clarification and additional transparency about how the included set was curated.","major_comments":[{"comment":"The reported screening numbers are internally inconsistent. The text says 2,656 articles were excluded during title/abstract screening, while Figure 2 reports 2,657 excluded records. The text also states that 83 papers came from manual searches and that the total is 144, but 60 database papers plus 83 manual papers gives 143; Figure 2 shows 84 records from 'other methods' (1 website + 83 citations), which would give 144. These discrepancies affect the credibility of the PRISMA-ScR flow and must be reconciled before the review can be considered reproducible.","section":"Section 3 and Figure 2"},{"comment":"The central claim of a 'shift from unimodal to multimodal approaches' is not supported by the described selection design. The database search was constructed as two targeted subsearches, one for text-only LLMs and one for multimodal models, so the retrieved set is a union of two intentionally defined categories rather than a sample from which prevalence can be inferred. Moreover, 83 of the 144 included papers (58%) came from a manual search described only as capturing 'recent and high-impact publications' (Section 2.3), and Section 2.4 states that inclusion was deliberately balanced 'to ensure proportional inclusion from prevalent fields.' Under this design, the observed preponderance of multimodal papers could be an artifact of the inclusion rules. To keep the claim, the authors should either restrict all trend statements to the curated set with an explicit caveat that the set is not a random sample, or provide a sensitivity analysis based only on database-derived papers and a prespecified manual search protocol.","section":"Sections 2.3-2.4 and Abstract"},{"comment":"The balancing criterion 'proportional inclusion from prevalent fields, such as X-ray report generation' is too vague to be a reproducible selection rule. The authors should specify how 'prevalent' was operationalized, who made the balancing decisions, what data were used to determine prevalence, and how many candidate papers were included or excluded because of this rule. Without this detail, a reader cannot determine whether the resulting 144-paper corpus is representative of the literature or of the authors' interests.","section":"Section 2.4"},{"comment":"The 'shift' claim is presented as a temporal trend, but no time-based analysis is shown. The review spans 2020-2024, yet no table or figure reports the distribution of unimodal versus multimodal papers by publication year, nor is there an analysis of whether the share of multimodal papers changed over time within the included set. If the authors intend to claim a shift, they need to provide direct evidence of changing proportions over time, or explicitly reframe the claim as a comparison between two categories present in the included set.","section":"Sections 5 and 7"}],"minor_comments":[{"comment":"The caption contains a typo: 'ebbedding space' should be 'embedding space.'","section":"Figure 3 caption"},{"comment":"The eligibility criteria state that 'manually selected preprints with high relevance and potential impact' were included, but the criteria do not define 'relevance' or 'impact.' Consider providing an operational definition.","section":"Section 2.1"},{"comment":"The column header 'Application' is so broad that entries like 'Report evaluation' and 'Conversation evaluation' are not fully descriptive; consider subdividing by evaluation target (e.g., text generation, image generation, calibration).","section":"Section 6, Table 6"},{"comment":"Some dataset sizes are given as raw counts without a clear unit (e.g., '224K triplets' vs. '25M pairs'); standardizing the unit label (e.g., 'image-text pairs,' 'patients') would improve clarity.","section":"Section 5.4, Table 5"},{"comment":"Several references are to arXiv preprints without a note of peer-reviewed status, which conflicts with the stated inclusion criterion of 'peer-reviewed conference and journal publications, alongside manually selected preprints.' Clarify which included papers are preprints and how they were evaluated for 'relevance and potential impact.'","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a scoping review, not a derivation, so circularity concerns are not applicable. The main risk is that the paper's high-level claim (a 'shift to multimodal AI') is likely to be cited as a field-level trend even though the reported selection process does not license that inference. I would advise the editor that the counting inconsistencies in Figure 2 and Section 3 should be fixed before publication, and that the authors should be asked to either temper the trend language or provide a more rigorous sensitivity analysis. The self-citations (refs [2], [10], [34], [133]) appear relevant and not disproportionate, but it would be appropriate to ensure the manual search process does not disproportionately favor the authors' own work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a competent, genuinely useful scoping review. The tables alone are worth the price of admission: they give a current map of text-only LLMs, CLIP-style models, MLLMs, datasets, and evaluation metrics in medical generative AI. The evaluation-metrics section, with clinical metrics like RaTEScore and GREEN contrasted against lexical metrics like BLEU and ROUGE, is a practical organizing contribution. The PRISMA-ScR checklist and full search strings in the supplement make the database portion reproducible. Where it gets soft is the trend claim. The abstract says the findings underscore a shift from unimodal to multimodal approaches. But 83 of 144 papers (58%) entered via a manual search described only as capturing recent and high-impact publications, and Section 2.4 says the authors deliberately aimed for proportional inclusion from prevalent fields like X-ray report generation. When you construct the corpus as a union of a text-only-LLM subsearch and a multimodal subsearch, and then hand-balance the topics, any statement about the prevalence or direction of the field is a property of the selection rule, not the literature. The shift may well be real, but this paper's methods cannot support it. There is also a small arithmetic inconsistency: the text says 60 database papers plus 83 manual papers equals 144, but that sum is 143. Figure 2 shows 84 records from other methods (1 website plus 83 citations), which would make 60 plus 84 equals 144. So the text and figure disagree by one. Minor, but the kind of thing a referee should flag. The paper does acknowledge limitations, including radiology overrepresentation and the manual search, so the authors are not hiding the ball. But the shift language is stronger than the methods justify. The qualitative synthesis of methods, datasets, and metrics is solid and independent of that claim. Who should read this: anyone wanting a quick, current map of medical generative AI, especially for teaching, grant writing, or entering the area. It is not a deep technical contribution, but as a scoping review it does its job. For peer review: yes, send it out. The material is current and useful, and the flaws are fixable. The authors should either soften the trend claim to the reviewed literature emphasizes multimodal approaches or reanalyze with a sensitivity check that excludes the hand-added papers. The off-by-one should be corrected. With those changes, I would be comfortable seeing it published.","headline":"Useful scoping review of medical generative AI, but the 'shift to multimodal' claim rests on a hand-curated subset and an off-by-one PRISMA count; worth refereeing after a revision.","tokens_in":719,"tokens_out":1051,"would_cite":true,"duration_ms":36077,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This scoping review of 144 papers claims generative AI in medicine is shifting from text-only large language models to multimodal systems that combine imaging, text, and structured data in a single model.","keywords":["generative AI in medicine","multimodal large language models","scoping review","clinical evaluation metrics","radiology report generation","medical datasets","generalist medical AI","contrastive learning"],"falsifier":"Count the share of included papers that integrate at least two modalities by publication year; if the share does not rise across 2022 to 2024, the claimed shift is an artifact of curation rather than a property of the literature.","tokens_in":26823,"feed_emoji":"🩺","tokens_out":4437,"duration_ms":39617,"temperature":0.7,"pith_summary":"This scoping review claims that generative AI in medicine is moving from text-only large language models to multimodal systems that combine images, text, and structured data in one model. The authors argue this shift is already visible in diagnostic support, medical report generation, drug discovery, and conversational AI, and that evaluation practice has not kept up. To ground this claim they systematically screened 4,384 records and included 144 papers published up to the end of 2024. The review matters because it identifies where clinical deployment is likely to succeed and where it will stall: data heterogeneity, interpretability, ethics, and real-world validation remain open problems.","feed_headline":"144 studies show medical AI moving beyond text-only models","feed_subtitle":"Scans, notes, and lab results are being merged into single models; evaluation metrics haven't caught up.","key_machinery":"The central objects are the two multimodal architectures the review uses to organize the field: contrastive language-image pretraining (CLIP) models, which align different modalities in a shared embedding space, and multimodal large language models (MLLMs), which encode images or other non-text data into the language model's embedding space. These carry the argument by defining what multimodal means and by structuring the taxonomy of methods, datasets, and applications. A second mechanism is the pairing of evaluation metrics: lexical metrics such as BLEU and ROUGE versus clinically grounded metrics such as entity-based scoring, grounded factual checks, and error-notation frameworks, which the review uses to show that evaluation lags behind model development.","core_discovery":"The paper's central finding is that the center of gravity of medical generative AI has shifted from unimodal large language models to multimodal models, organized around two architectural families: contrastive models that align images with text in a shared embedding space, and multimodal large language models that project non-text features into the language model's embedding space. This shift is documented across 144 included papers and appears in a growing family of generalist models that handle classification, segmentation, report generation, and visual question answering within one architecture. The authors also find that standard lexical evaluation metrics such as BLEU and ROUGE miss clinical correctness, and that newer clinically grounded metrics, while promising, are not yet standardized across sites and specialties.","pith_inferences":["Editorial inference: If the field's center of gravity is indeed shifting to multimodal models, text-only clinical language models may become components of larger systems rather than standalone products.","Editorial inference: The review's selection method makes the shift claim sensitive to curation, so a formal bibliometric analysis of all published medical AI abstracts would test whether the multimodal trend is a property of the literature or an artifact of inclusion choices.","Editorial inference: Clinical adoption may hinge less on model capability than on the availability of evaluation benchmarks that predict real-world diagnostic utility, a gap the review identifies but does not quantify.","Editorial inference: The heavy representation of radiology suggests the next bottleneck is data: without analogous multimodal datasets in pathology, genomics, and primary care, the shift to multimodal AI will proceed unevenly across medicine."],"forward_implications":["Clinical AI development will increasingly center on models that ingest images, text, and structured data together rather than text-only assistants.","Evaluation of medical generative models should combine lexical metrics with clinically grounded metrics that check factual correctness and relevance.","Dataset construction should move beyond radiology-centric resources toward other specialties and modalities to avoid generalization failures.","Generalist models that unify multiple tasks and modalities are becoming the default architectural direction for medical AI.","Synthetic image generation will be used to augment scarce data and simulate rare conditions, provided its clinical utility is validated on downstream tasks."],"supporting_citations":[{"why":"Supplies the scoping-review reporting standard that structures the review's methods.","marker":"[21]"},{"why":"Provides the transformer architecture that underlies the large language models the review traces.","marker":"[25]"},{"why":"Defines the contrastive language-image pretraining architecture used to align image and text modalities.","marker":"[76]"},{"why":"Provides the paired chest X-ray and report dataset that anchors many reviewed multimodal systems.","marker":"[8]"},{"why":"Exemplifies the generalist multimodal biomedical model that supports the claimed shift.","marker":"[13]"},{"why":"Supplies both a multimodal CT model and a paired CT-text dataset used across reviewed applications.","marker":"[14]"},{"why":"Demonstrates grounded radiology report generation, a key multimodal application the review highlights.","marker":"[72]"},{"why":"Provides a clinically grounded entity-based metric for radiology report evaluation.","marker":"[155]"},{"why":"Offers a clinically grounded evaluation metric with error notation, supporting the review's evaluation-gap claim.","marker":"[154]"},{"why":"Provides a public benchmark for neutral evaluation of radiology report generation, supporting the call for standardized benchmarks.","marker":"[152]"}],"fun_headline_variants":["Medical AI goes multimodal: 144 studies chart the shift","From text to images: How medical AI is merging data types","One model for scans and notes: Medical AI's next leap","Review finds medical AI now integrates images, text, and more","Multimodal medical AI: 144 papers reveal a field in flux"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's trend claims depend on the representativeness of the curated set of 144 papers, since 83 entered through a manual search described as capturing recent and high-impact work and the authors deliberately balanced inclusion across prevalent fields.","fun_headline_variants_meta":{"raw":{"variants":["Medical AI goes multimodal: 144 studies chart the shift","From text to images: How medical AI is merging data types","One model for scans and notes: Medical AI's next leap","Review finds medical AI now integrates images, text, and more","Multimodal medical AI: 144 papers reveal a field in flux"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000724,"raw_usage":{"total_tokens":3244,"prompt_tokens":942,"completion_tokens":2302,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":2215}},"tokens_in":558,"tokens_out":2302,"duration_ms":13810,"temperature":1.0,"reasoning_tokens":2215,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T22:10:28.364976+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Count the share of included papers that integrate at least two modalities by publication year; if the share does not rise across 2022 to 2024, the claimed shift is an artifact of curation rather than a property of the literature.","supporting_citations":[{"cited_title":"medRxiv, 2024–06 (2024)","cited_arxiv_id":null,"evidence_quote":"Provides a clinically grounded entity-based metric for radiology report evaluation."}],"review_version":1}