Pith. sign in

REVIEW 6 cited by

A ChatGPT Aided Explainable Framework for Zero-Shot Medical Image Diagnosis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.01981 v1 pith:G7W3LTVS submitted 2023-07-05 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords medicalimagezero-shotdiagnosisexplainablechatgptframeworkapplications
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Zero-shot medical image classification is a critical process in real-world scenarios where we have limited access to all possible diseases or large-scale annotated data. It involves computing similarity scores between a query medical image and possible disease categories to determine the diagnostic result. Recent advances in pretrained vision-language models (VLMs) such as CLIP have shown great performance for zero-shot natural image recognition and exhibit benefits in medical applications. However, an explainable zero-shot medical image recognition framework with promising performance is yet under development. In this paper, we propose a novel CLIP-based zero-shot medical image classification framework supplemented with ChatGPT for explainable diagnosis, mimicking the diagnostic process performed by human experts. The key idea is to query large language models (LLMs) with category names to automatically generate additional cues and knowledge, such as disease symptoms or descriptions other than a single category name, to help provide more accurate and explainable diagnosis in CLIP. We further design specific prompts to enhance the quality of generated texts by ChatGPT that describe visual medical features. Extensive results on one private dataset and four public datasets along with detailed analysis demonstrate the effectiveness and explainability of our training-free zero-shot diagnosis pipeline, corroborating the great potential of VLMs and LLMs for medical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Brain Imaging Foundation Models, Are We There Yet? A Systematic Review of Foundation Models for Brain Imaging and Biomedical Research

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A systematic review of brain imaging foundation models covering 86 models and 161 datasets, with a performance tournament, dataset atlas, and duplicated-data warnings.

  2. MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A concept-guided distillation method lets a fundus-image model learn from unpaired OCT scans during training, improving retinal disease classification when only fundus photos are available at test time.

  3. BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    BiomedCoOp improves few-shot biomedical image classification by aligning learnable prompts with selectively pruned LLM-generated prompt ensembles and distilling their knowledge into BiomedCLIP.

  4. Fair-MoE: Fairness-Oriented Mixture of Experts in Vision-Language Models

    cs.CV 2025-02 reject novelty 5.0 of 10

    Fair-MoE reports improved accuracy and fairness on Harvard-FairVLMed for some protected attributes by adding sparse mixture-of-experts layers and a variance-based fairness loss to CLIP, but the all-attribute improveme...

  5. Explainability for Vision Foundation Models: A Survey

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A structured review of 122 papers on explainability for vision foundation models, with a taxonomy and the finding that quantitative evaluation is rare (36%).

  6. A Survey of Medical Vision-and-Language Applications and Their Techniques

    cs.CV 2024-11 conditional novelty 4.0 of 10

    This survey reviews medical vision-and-language models across five tasks and organizes existing methods, datasets, and evaluation metrics without introducing new techniques.

Pith tools