Pith. sign in

REVIEW 2 cited by

Pretrained ViTs Yield Versatile Representations For Medical Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07034 v3 pith:TF2WOOUE submitted 2023-03-13 cs.CV

classification cs.CV
keywords cnnsmedicalimagetasksalternativeclassificationperformpretrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Convolutional Neural Networks (CNNs) have reigned for a decade as the de facto approach to automated medical image diagnosis, pushing the state-of-the-art in classification, detection and segmentation tasks. Over the last years, vision transformers (ViTs) have appeared as a competitive alternative to CNNs, yielding impressive levels of performance in the natural image domain, while possessing several interesting properties that could prove beneficial for medical imaging tasks. In this work, we explore the benefits and drawbacks of transformer-based models for medical image classification. We conduct a series of experiments on several standard 2D medical image benchmark datasets and tasks. Our findings show that, while CNNs perform better if trained from scratch, off-the-shelf vision transformers can perform on par with CNNs when pretrained on ImageNet, both in a supervised and self-supervised setting, rendering them as a viable alternative to CNNs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification

    cs.CV 2025-06 reject novelty 4.0 of 10

    A ViT feature ensemble plus ML classifier voting pipeline is evaluated on two binary brain MRI datasets, reporting up to 99.8% accuracy without a same-dataset comparison against prior methods.

  2. Advancements in Medical Image Classification through Fine-Tuning Natural Domain Foundation Models

    eess.IV 2025-05 conditional novelty 4.0 of 10

    Fine-tuning recent natural-domain foundation models, especially AIMv2, improves medical image classification accuracy across mammography, skin lesion, retinopathy, and chest X-ray benchmarks.

Pith tools