A spiking Transformer variant, SAFormer, drops the value matrix and uses downsampled spike query/key pairs plus depthwise convolution to reach 95.8% on CIFAR-10 and 81.3% on CIFAR10-DVS with lower estimated energy cost.
Attention is all you need
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.NE 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Combining Aggregated Attention and Transformer Architecture for Accurate and Efficient Performance of Spiking Neural Networks
A spiking Transformer variant, SAFormer, drops the value matrix and uses downsampled spike query/key pairs plus depthwise convolution to reach 95.8% on CIFAR-10 and 81.3% on CIFAR10-DVS with lower estimated energy cost.