Pith. sign in

REVIEW 14 cited by

Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.09909 v3 pith:SRLXYMNN submitted 2023-10-15 cs.CV cs.CL

classification cs.CVcs.CL
keywords gpt-4vmedicaldiagnosisdiseasemultimodalusedanatomyapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Driven by the large foundation models, the development of artificial intelligence has witnessed tremendous progress lately, leading to a surge of general interest from the public. In this study, we aim to assess the performance of OpenAI's newest model, GPT-4V(ision), specifically in the realm of multimodal medical diagnosis. Our evaluation encompasses 17 human body systems, including Central Nervous System, Head and Neck, Cardiac, Chest, Hematology, Hepatobiliary, Gastrointestinal, Urogenital, Gynecology, Obstetrics, Breast, Musculoskeletal, Spine, Vascular, Oncology, Trauma, Pediatrics, with images taken from 8 modalities used in daily clinic routine, e.g., X-ray, Computed Tomography (CT), Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Digital Subtraction Angiography (DSA), Mammography, Ultrasound, and Pathology. We probe the GPT-4V's ability on multiple clinical tasks with or without patent history provided, including imaging modality and anatomy recognition, disease diagnosis, report generation, disease localisation. Our observation shows that, while GPT-4V demonstrates proficiency in distinguishing between medical image modalities and anatomy, it faces significant challenges in disease diagnosis and generating comprehensive reports. These findings underscore that while large multimodal models have made significant advancements in computer vision and natural language processing, it remains far from being used to effectively support real-world medical applications and clinical decision-making. All images used in this report can be found in https://github.com/chaoyi-wu/GPT-4V_Medical_Evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink

    cs.LG 2025-01 conditional novelty 7.0 of 10

    Adversarial images optimized to induce attention sink behavior increase hallucination rates in multiple MLLMs, including commercial APIs, without visibly degrading response quality.

  2. Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On a new 65,464-item binary benchmark that swaps one medical term per caption, four medical multimodal models scored 49–62%, close to chance.

  3. Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 3D encoder pretrained with GPT-4V slice captions and partial optimal transport alignment beats vision-only SSL baselines on several medical tasks, but a key evaluation dataset may overlap with pretraining.

  4. Active Learning for Neurosymbolic Program Synthesis

    cs.PL 2025-08 unverdicted novelty 6.0 of 10

    The abstract claims a new active learning technique, constrained conformal evaluation (tool SmartLabel), that finds the ground-truth program in 98% of benchmarks, but the delivered full text is a different paper, leav...

  5. MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Current medical multimodal models, including GPT-4o and Claude 3.5 Sonnet, fail simple perceptual tasks on medical images that human experts solve almost perfectly.

  6. Are MLMs Trapped in the Visual Room?

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new Reddit-derived sarcasm benchmark and two-tier evaluation show that multimodal models can perceive scenes accurately while still failing to grasp sarcastic intent.

  7. How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Large language models complete only 29-67% of simulated radiology agent tasks, and prompting tricks and a simulated tool builder do not close the gap on complex workflows.

  8. MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MMedPO weights preference-optimization training samples by clinical relevance scores, combining hallucinated text answers and locally noised lesion images, and reports improved medical VQA and report generation metrics.

  9. MpoxVLM: A Vision-Language Model for Diagnosing Skin Lesions from Mpox Virus Infection

    eess.IV 2024-11 reject novelty 6.0 of 10

    MpoxVLM reports top accuracy for mpox detection from skin images and clinical data, but the design feeds the answer into the model through a mpox-specific lesion stage feature, making the reported result unreliable.

  10. Detecting Children with Autism Spectrum Disorder based on Script-Centric Behavior Understanding with Emotional Enhancement

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A zero-shot LLM pipeline that converts audio-visual behavior into textual scripts and emotion descriptions detects ASD in two-year-olds with 95.24% F1, but the estimate is weakened by threshold tuning on the evaluation data.

  11. Cross-Modal Consistency in Multimodal Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    GPT-4V answers identical questions much less accurately when they are presented as images than as text, even when it can extract the image content nearly perfectly.

  12. Discrete Prompt Tuning via Recursive Utilization of Black-box Multimodal Large Language Model for Personalized Visual Emotion Recognition

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Recursive generation, evaluation, and refinement of discrete prompts tunes a black-box multimodal LLM to each user, raising personalized visual emotion recognition accuracy on Affection from 40.6% to 44.9%.

  13. Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

    cs.CL 2025-01 conditional novelty 5.0 of 10

    UMed-LVLM uses GPT-4V-generated abnormality data and abnormal-aware rewards to improve medical image diagnosis and localization.

  14. SpatialFly: Implicit 3D Prior-Guided Visual Reparameterization for Continuous UAV Vision-and-Language Navigation

    cs.CV 2026-03 unverdicted novelty 4.0 of 10

    SpatialFly reparameterizes 2D visual tokens with implicit geometric priors and reports lower navigation error and higher success than prior UAV VLN systems on unseen splits.

Pith tools