Pith. sign in

REVIEW 13 cited by

Med-Flamingo: a Multimodal Medical Few-shot Learner

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.15189 v1 pith:ZBX3NHJ3 submitted 2023-07-27 cs.CV cs.AI

classification cs.CVcs.AI
keywords medicalmed-flamingofew-shotgenerativemodelsmultimodalapplicationsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medicine, by its nature, is a multifaceted domain that requires the synthesis of information across various modalities. Medical generative vision-language models (VLMs) make a first step in this direction and promise many exciting clinical applications. However, existing models typically have to be fine-tuned on sizeable down-stream datasets, which poses a significant limitation as in many medical applications data is scarce, necessitating models that are capable of learning from few examples in real-time. Here we propose Med-Flamingo, a multimodal few-shot learner adapted to the medical domain. Based on OpenFlamingo-9B, we continue pre-training on paired and interleaved medical image-text data from publications and textbooks. Med-Flamingo unlocks few-shot generative medical visual question answering (VQA) abilities, which we evaluate on several datasets including a novel challenging open-ended VQA dataset of visual USMLE-style problems. Furthermore, we conduct the first human evaluation for generative medical VQA where physicians review the problems and blinded generations in an interactive app. Med-Flamingo improves performance in generative medical VQA by up to 20\% in clinician's rating and firstly enables multimodal medical few-shot adaptations, such as rationale generation. We release our model, code, and evaluation app under https://github.com/snap-stanford/med-flamingo.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

    cs.CV 2026-05 conditional novelty 7.0 of 10

    Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.

  2. NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives

    cs.NE 2026-08 conditional novelty 6.0 of 10

    NeuroMosaic links MRI regions to diagnostic language via an anatomical graph router and concept memory, reporting external macro-F1 up to 0.784, IDH AUROC 0.918, and 0.703 pointing accuracy, with a 0.036 macro-F1 gain...

  3. Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A query-conditioned latent evidence aggregator after frozen frame selection improves long-video QA by up to +5.2 average / +10.1 LVBench with 0.11–0.40% token overhead.

  4. DentiAsk: A VQA Benchmark for Multimodal Reasoning in Panoramic Dental Radiographs

    q-bio.QM 2026-06 conditional novelty 6.0 of 10

    A 1,000-image, 10,000-QA dental VQA benchmark shows current VLMs handle descriptive recognition far better than spatial localization or numerical counting on panoramic radiographs.

  5. Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    DIVE improves in-context vector distillation for medical report generation via decisive-token supervision on pathology terms and EOS plus state-conditioned dynamic steering, achieving top BLEU-4, ROUGE-L and RadGraph ...

  6. Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Introduces Wasserstein equilibrium decoding that improves accuracy and convergence speed for small VLMs on medical VQA benchmarks by using semantic consensus instead of lexical order.

  7. Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A multi-view vision-language model trained on 20,000 fetal ultrasound reports generates clinical text and diagnoses, reportedly outperforming general and medical baselines.

  8. Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A train-free method that compresses 3D medical volumes into small embeddings via a frozen 2D foundation model and random projections, outperforming several medical-volume pretrained models on benchmark tasks.

  9. Ask4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Ask4VG learns a risk estimator from counterfactual visual probes to rerank question rewrites, reducing held-out hallucination risk from 0.658 to 0.623 and raising accuracy from 0.337 to 0.356 on VQA-RAD.

  10. BiomedAP: A Vision-Informed Dual-Anchor Framework with Gated Cross-Modal Fusion for Robust Medical Vision-Language Adaptation

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    BiomedAP improves robustness of biomedical VLMs to prompt variations using gated cross-modal fusion and dual-anchor constraints, outperforming baselines on 11 benchmarks.

  11. FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    FLAME is an MoE architecture using modality-specific routers and low-rank compression of expert knowledge to support efficient continual multimodal multi-task learning while reducing catastrophic forgetting.

  12. Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

    cs.CV 2026-07 conditional novelty 4.0 of 10

    On MedFrameQA, order-vote (57.89%) beats fixed prompting (52.73%) and order-rerank (55.79%), and a single 100-generation run drops final-test accuracy to 56.02%.

  13. Data-Centric Foundation Models in Computational Healthcare: A Survey

    cs.LG 2024-01 unverdicted novelty 3.0 of 10

    The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.

Pith tools