Pith. sign in

REVIEW 1 cited by

MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.14719 v3 pith:CXSZUQHC submitted 2025-05-19 cs.CV cs.AI

classification cs.CVcs.AI
keywords spikingarchitecturesattentionmsvittransformerexistingmethodsmulti-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has garnered significant attention due to their potential for energy-efficient and high-performance computing paradigms. However, a substantial performance gap still exists between SNN-based and ANN-based transformer architectures. While existing methods propose spiking self-attention mechanisms that are successfully combined with SNNs, the overall architectures proposed by these methods suffer from a bottleneck in effectively extracting features from different image scales. In this paper, we address this issue and propose MSVIT. This novel spike-driven Transformer architecture firstly uses multi-scale spiking attention (MSSA) to enhance the capabilities of spiking attention blocks. We validate our approach across various main datasets. The experimental results show that MSVIT outperforms existing SNN-based models, positioning itself as a state-of-the-art solution among SNN-transformer architectures. The codes are available at https://github.com/Nanhu-AI-Lab/MSViT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SAFformer:Improving Spiking Transformer via Active Predictive Filtering

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    SAFformer uses brain-inspired active predictive filtering in a spiking transformer to reach new state-of-the-art accuracy on CIFAR-10/100 and CIFAR10-DVS plus 80.5% top-1 on ImageNet-1K at 26.58M parameters and 5.88 m...

Pith tools