SpikeVideoFormer replaces dot-product attention with normalized Hamming similarity in a spiking transformer, achieving linear temporal complexity and state-of-the-art SNN results on three video tasks.
Xception: Deep learning with depthwise separable convolutions
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
SpikeVideoFormer replaces dot-product attention with normalized Hamming similarity in a spiking transformer, achieving linear temporal complexity and state-of-the-art SNN results on three video tasks.