A spiking vision transformer with multi-scale attention fusion (MSVIT) reaches 85.06% top-1 ImageNet accuracy with a linear-complexity attention mechanism and reports state-of-the-art results among compared spiking transformers.
Multiscale vision transformers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
A spiking vision transformer with multi-scale attention fusion (MSVIT) reaches 85.06% top-1 ImageNet accuracy with a linear-complexity attention mechanism and reports state-of-the-art results among compared spiking transformers.