MVFormer combines a weighted blend of three normalizations with a three-branch multiscale convolutional token mixer, achieving modest top-1 accuracy gains on ImageNet-1K and downstream vision tasks.
Coatnet: Marrying convolution and attention for all data sizes
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
MVFormer combines a weighted blend of three normalizations with a three-branch multiscale convolutional token mixer, achieving modest top-1 accuracy gains on ImageNet-1K and downstream vision tasks.