REVIEW 1 cited by
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has garnered significant attention due to their potential for energy-efficient and high-performance computing paradigms. However, a substantial performance gap still exists between SNN-based and ANN-based transformer architectures. While existing methods propose spiking self-attention mechanisms that are successfully combined with SNNs, the overall architectures proposed by these methods suffer from a bottleneck in effectively extracting features from different image scales. In this paper, we address this issue and propose MSVIT. This novel spike-driven Transformer architecture firstly uses multi-scale spiking attention (MSSA) to enhance the capabilities of spiking attention blocks. We validate our approach across various main datasets. The experimental results show that MSVIT outperforms existing SNN-based models, positioning itself as a state-of-the-art solution among SNN-transformer architectures. The codes are available at https://github.com/Nanhu-AI-Lab/MSViT.
Forward citations
Cited by 1 Pith paper
-
SAFformer:Improving Spiking Transformer via Active Predictive Filtering
SAFformer uses brain-inspired active predictive filtering in a spiking transformer to reach new state-of-the-art accuracy on CIFAR-10/100 and CIFAR10-DVS plus 80.5% top-1 on ImageNet-1K at 26.58M parameters and 5.88 m...
Discussion (0). Sign in to comment.