Pith. sign in

REVIEW 7 cited by

Robust and Interpretable Medical Image Classifiers via Concept Bottleneck Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03182 v1 pith:RB4XBSXU submitted 2023-10-04 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords imagemedicalconceptsmodelsclassificationmodelwhenclassifiers
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Medical image classification is a critical problem for healthcare, with the potential to alleviate the workload of doctors and facilitate diagnoses of patients. However, two challenges arise when deploying deep learning models to real-world healthcare applications. First, neural models tend to learn spurious correlations instead of desired features, which could fall short when generalizing to new domains (e.g., patients with different ages). Second, these black-box models lack interpretability. When making diagnostic predictions, it is important to understand why a model makes a decision for trustworthy and safety considerations. In this paper, to address these two limitations, we propose a new paradigm to build robust and interpretable medical image classifiers with natural language concepts. Specifically, we first query clinical concepts from GPT-4, then transform latent image features into explicit concepts with a vision-language model. We systematically evaluate our method on eight medical image classification datasets to verify its effectiveness. On challenging datasets with strong confounding factors, our method can mitigate spurious correlations thus substantially outperform standard visual encoders and other baselines. Finally, we show how classification with a small number of concepts brings a level of interpretability for understanding model decisions through case studies in real medical data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

    cs.CV 2025-12 reject novelty 6.0 of 10

    A 'tool bottleneck' framework—VLM tool selection plus learned spatial fusion—matches or beats black-box classifiers, especially on scarce data.

  2. MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A concept-guided distillation method lets a fundus-image model learn from unpaired OCT scans during training, improving retinal disease classification when only fundus photos are available at test time.

  3. Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models

    cs.CR 2024-11 reject novelty 5.0 of 10

    ConceptGuard protects concept bottleneck models from concept-level backdoor attacks by clustering concepts, training an ensemble of sub-classifiers, and taking a majority vote.

  4. Stable Vision Concept Transformers for Medical Diagnosis

    cs.CV 2025-06 reject novelty 4.0 of 10

    A vision transformer with a concept bottleneck and denoised diffusion smoothing is claimed to give stable concept explanations under input perturbations while keeping diagnostic accuracy.

  5. Explainability for Vision Foundation Models: A Survey

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A structured review of 122 papers on explainability for vision foundation models, with a taxonomy and the finding that quantitative evaluation is rare (36%).

  6. Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning

    cs.CV 2025-01 reject novelty 4.0 of 10

    HiCA, a hierarchical contrastive fine-tuning method for large vision-language models, is claimed to achieve state-of-the-art few-shot medical image classification, but the paper lacks the experimental detail needed to...

  7. Towards Interpretable Radiology Report Generation via Concept Bottlenecks using a Multi-Agentic RAG

    cs.IR 2024-12 conditional novelty 4.0 of 10

    A concept-bottleneck classifier coupled with a multi-agent RAG pipeline generates chest X-ray reports, with LLM-based evaluation scoring the multi-agent approach higher than single-agent or GPT-4 baselines.

Pith tools