Smaller image patches improve transformer accuracy on biomedical tasks, attention-window size matters less, and Hyena/MambaVision operators match attention with up to 80% faster training.
We required a minimum batch size of two to fit on the GPU to enable batch normalization layers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Study on Context Length and Efficient Transformers for Biomedical Image Analysis
Smaller image patches improve transformer accuracy on biomedical tasks, attention-window size matters less, and Hyena/MambaVision operators match attention with up to 80% faster training.