REVIEW 7 cited by
Robust and Interpretable Medical Image Classifiers via Concept Bottleneck Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Medical image classification is a critical problem for healthcare, with the potential to alleviate the workload of doctors and facilitate diagnoses of patients. However, two challenges arise when deploying deep learning models to real-world healthcare applications. First, neural models tend to learn spurious correlations instead of desired features, which could fall short when generalizing to new domains (e.g., patients with different ages). Second, these black-box models lack interpretability. When making diagnostic predictions, it is important to understand why a model makes a decision for trustworthy and safety considerations. In this paper, to address these two limitations, we propose a new paradigm to build robust and interpretable medical image classifiers with natural language concepts. Specifically, we first query clinical concepts from GPT-4, then transform latent image features into explicit concepts with a vision-language model. We systematically evaluate our method on eight medical image classification datasets to verify its effectiveness. On challenging datasets with strong confounding factors, our method can mitigate spurious correlations thus substantially outperform standard visual encoders and other baselines. Finally, we show how classification with a small number of concepts brings a level of interpretability for understanding model decisions through case studies in real medical data.
Forward citations
Cited by 7 Pith papers
-
A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
A 'tool bottleneck' framework—VLM tool selection plus learned spatial fusion—matches or beats black-box classifiers, especially on scarce data.
-
MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images
A concept-guided distillation method lets a fundus-image model learn from unpaired OCT scans during training, improving retinal disease classification when only fundus photos are available at test time.
-
Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models
ConceptGuard protects concept bottleneck models from concept-level backdoor attacks by clustering concepts, training an ensemble of sub-classifiers, and taking a majority vote.
-
Stable Vision Concept Transformers for Medical Diagnosis
A vision transformer with a concept bottleneck and denoised diffusion smoothing is claimed to give stable concept explanations under input perturbations while keeping diagnostic accuracy.
-
Explainability for Vision Foundation Models: A Survey
A structured review of 122 papers on explainability for vision foundation models, with a taxonomy and the finding that quantitative evaluation is rare (36%).
-
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
HiCA, a hierarchical contrastive fine-tuning method for large vision-language models, is claimed to achieve state-of-the-art few-shot medical image classification, but the paper lacks the experimental detail needed to...
-
Towards Interpretable Radiology Report Generation via Concept Bottlenecks using a Multi-Agentic RAG
A concept-bottleneck classifier coupled with a multi-agent RAG pipeline generates chest X-ray reports, with LLM-based evaluation scoring the multi-agent approach higher than single-agent or GPT-4 baselines.
Discussion (0). Continue with ORCID to comment.