Pith. sign in

REVIEW 5 cited by

Multimodal ChatGPT for Medical Applications: an Experimental Study of GPT-4V

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19061 v1 pith:V7EWKRYC submitted 2023-10-29 cs.CV

classification cs.CV
keywords gpt-4vmedicalaccuracyansweringdatasetsexperimentsmultimodalquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA) task. Our experiments thoroughly assess GPT-4V's proficiency in answering questions paired with images using both pathology and radiology datasets from 11 modalities (e.g. Microscopy, Dermoscopy, X-ray, CT, etc.) and fifteen objects of interests (brain, liver, lung, etc.). Our datasets encompass a comprehensive range of medical inquiries, including sixteen distinct question types. Throughout our evaluations, we devised textual prompts for GPT-4V, directing it to synergize visual and textual information. The experiments with accuracy score conclude that the current version of GPT-4V is not recommended for real-world diagnostics due to its unreliable and suboptimal accuracy in responding to diagnostic medical questions. In addition, we delineate seven unique facets of GPT-4V's behavior in medical VQA, highlighting its constraints within this complex arena. The complete details of our evaluation cases are accessible at https://github.com/ZhilingYan/GPT4V-Medical-Report.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 33 citations worldwide. Full citation record

  1. Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On a new 65,464-item binary benchmark that swaps one medical term per caption, four medical multimodal models scored 49–62%, close to chance.

  2. Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training

    eess.IV 2025-08 conditional novelty 6.0 of 10

    On abdominal CT, a vision-language pre-training method using organ-level normal/abnormal contrastive learning and a VQ-VAE normality model achieves 84.9% average zero-shot AUC, beating prior methods by 3.6%.

  3. MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Grouping instruction-tuning datasets by redundancy, uniqueness, or synergy of text-image interaction improves vision-language model accuracy over single-task and unselective multi-task tuning.

  4. Deep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical Translation

    cs.CV 2026-07 accept novelty 4.0 of 10

    A comprehensive review synthesizing AI-driven sperm analysis across computer vision tasks, multimodal fusion, and a staged clinical translation roadmap.

  5. Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models

    cs.CV 2025-07 reject novelty 4.0 of 10

    Expert-CFG combines entropy-based uncertainty selection with classifier-free guidance over expert-highlighted text to refine MedVLM outputs, reporting gains on VQA-RAD, SLAKE, and PathVQA.

Pith tools