REVIEW 19 cited by
Advancing Multimodal Medical Capabilities of Gemini
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Many clinical tasks require an understanding of specialized data, such as medical images and genomics, which is not typically found in general-purpose large multimodal models. Building upon Gemini's multimodal models, we develop several models within the new Med-Gemini family that inherit core capabilities of Gemini and are optimized for medical use via fine-tuning with 2D and 3D radiology, histopathology, ophthalmology, dermatology and genomic data. Med-Gemini-2D sets a new standard for AI-based chest X-ray (CXR) report generation based on expert evaluation, exceeding previous best results across two separate datasets by an absolute margin of 1% and 12%, where 57% and 96% of AI reports on normal cases, and 43% and 65% on abnormal cases, are evaluated as "equivalent or better" than the original radiologists' reports. We demonstrate the first ever large multimodal model-based report generation for 3D computed tomography (CT) volumes using Med-Gemini-3D, with 53% of AI reports considered clinically acceptable, although additional research is needed to meet expert radiologist reporting quality. Beyond report generation, Med-Gemini-2D surpasses the previous best performance in CXR visual question answering (VQA) and performs well in CXR classification and radiology VQA, exceeding SoTA or baselines on 17 of 20 tasks. In histopathology, ophthalmology, and dermatology image classification, Med-Gemini-2D surpasses baselines across 18 out of 20 tasks and approaches task-specific model performance. Beyond imaging, Med-Gemini-Polygenic outperforms the standard linear polygenic risk score-based approach for disease risk prediction and generalizes to genetically correlated diseases for which it has never been trained. Although further development and evaluation are necessary in the safety-critical medical domain, our results highlight the potential of Med-Gemini across a wide range of medical tasks.
Forward citations
Cited by 19 Pith papers
-
Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering
DeepTumorVQA, a 9,262-volume 3D medical VQA benchmark, shows that current vision-language models handle measurement but remain far from clinical-grade lesion recognition and reasoning.
-
BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
BioProBench provides a large-scale benchmark and training resource for evaluating and improving language models' reasoning about biological experimental protocols.
-
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...
-
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
Current medical multimodal models, including GPT-4o and Claude 3.5 Sonnet, fail simple perceptual tasks on medical images that human experts solve almost perfectly.
-
Insights into a radiology-specialised multimodal large language model with sparse autoencoders
Applying Matryoshka sparse autoencoders to a radiology-specialised multimodal LLM reveals a minority of interpretable clinical features, while steering them produces unreliable and often off-target report changes.
-
VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge
A four-stage training recipe plus on-demand expert-model feedback lets a small open-source VLM beat much larger medical VLMs on several VQA, classification, and report generation benchmarks.
-
MAIRA-Seg: Enhancing Radiology Report Generation with Segmentation-Aware Multimodal Large Language Models
MAIRA-Seg shows that feeding segmentation mask tokens alongside chest X-rays improves mask-relevant clinical metrics for MLLM radiology report generation, but not lexical overlap.
-
The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models
Continued pretraining of open LLMs and VLMs on biomedical data yields little or no consistent improvement over their base models on closed-ended medical QA in zero-/few-shot and supervised fine-tuning regimes.
-
Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer
CT Foundation embeddings predict 2-year distant metastasis in head and neck cancer with AUC 0.791, matching radiomics-plus-deep-learning (0.794) while requiring no tumor contours.
-
Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
Harrison.Rad 1.5 is a radiology-specific multimodal LLM that passes simulated FRCR 2B Short Case examinations and outperforms general-purpose frontier models on plain-film radiography reporting tasks.
-
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
ChatEXAONEPath fine-tunes LLaVA with WSI-level features from TCGA slides, reaching 62.9% acceptance by an AI judge on generated pathology reports, though the judge's reliability is questioned in the paper.
-
Towards Fair Medical AI: Adversarial Debiasing of 3D CT Foundation Embeddings
A variational autoencoder with adversarial branches reduces sex and age signal in 3D CT foundation embeddings while preserving lung cancer risk prediction accuracy.
-
ACE-$M^3$: Automatic Capability Evaluator for Multimodal Medical Models
A branch-merge and reward-token DPO-based multimodal evaluator that surpasses GPT-4-Turbo on image-text medical QA, but relies on LLM-generated labels.
-
Anatomy-Guided Radiology Report Generation with Pathology-Aware Regional Prompts
A report generation pipeline that uses detected pathologies mapped to anatomical regions as prompt tokens improves several NLG and clinical efficacy metrics on MIMIC-CXR.
-
Architecting Clinical Collaboration: Multi-Agent Reasoning Systems for Multimodal Medical VQA
A systematic study on dermatology VQA finds that multi-agent reasoning and retrieval architectures outperform fine-tuned open-source vision-language models, maintaining 70% accuracy under distribution shift.
-
Chest X-ray Foundation Model with Global and Local Representations Integration
CheXFound, a ViT-Large model pretrained on 987K CXRs with DINOv2 plus the GLoRI head, outperforms prior CXR foundation models on long-tailed disease classification and transfer tasks.
-
Demographic Predictability in 3D CT Foundation Embeddings
Patient age and sex can be predicted from 3D CT foundation embeddings with high accuracy, while race prediction is weaker.
-
Health AI Developer Foundations
Health AI Developer Foundations packages six domain-specific medical embedding models into one platform, claiming large data and compute savings for downstream health ML tasks.
-
From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.
Discussion (0). Continue with ORCID to comment.