A 4.5M-parameter ViT combining convolutional tokenization, diagonal masking, temperature scaling, and sequence pooling matches or beats larger models on MedMNIST.
Scientific Reports 14(1), 12567 (May 2024)
1 Pith paper cite this work, alongside 57 external citations. Polarity classification is still indexing.
1
Pith paper citing it
57
external citations · OpenAlex
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
A 4.5M-parameter ViT combining convolutional tokenization, diagonal masking, temperature scaling, and sequence pooling matches or beats larger models on MedMNIST.