REVIEW 11 cited by
Spikformer: When Spiking Neural Network Meets Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We consider two biologically plausible structures, the Spiking Neural Network (SNN) and the self-attention mechanism. The former offers an energy-efficient and event-driven paradigm for deep learning, while the latter has the ability to capture feature dependencies, enabling Transformer to achieve good performance. It is intuitively promising to explore the marriage between them. In this paper, we consider leveraging both self-attention capability and biological properties of SNNs, and propose a novel Spiking Self Attention (SSA) as well as a powerful framework, named Spiking Transformer (Spikformer). The SSA mechanism in Spikformer models the sparse visual feature by using spike-form Query, Key, and Value without softmax. Since its computation is sparse and avoids multiplication, SSA is efficient and has low computational energy consumption. It is shown that Spikformer with SSA can outperform the state-of-the-art SNNs-like frameworks in image classification on both neuromorphic and static datasets. Spikformer (66.3M parameters) with comparable size to SEW-ResNet-152 (60.2M,69.26%) can achieve 74.81% top1 accuracy on ImageNet using 4 time steps, which is the state-of-the-art in directly trained SNNs models.
Forward citations
Cited by 11 Pith papers
-
SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event Reasoning
A spiking neural network with subtractive and additive attention performs all-in-one image restoration in one time step, matching older ANN baselines with much lower estimated energy.
-
Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition
Spike-HTR is a spiking transformer that recognizes handwritten text with T=2 timesteps, achieving 5.4/2.5/3.9 test CER on IAM/LAM/READ2016 and cutting mixer sequence length by 32% with negligible accuracy loss.
-
SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks
SMM Transformer uses spiking neurons, spike-driven token mixing, and a spiking mixture of experts to reach ANN-comparable accuracy on vision and vision-language tasks with lower estimated compute energy.
-
TEFormer: Structured Bidirectional Temporal Enhancement Modeling in Spiking Transformers
A spiking transformer with forward temporal EMA in attention and backward gated recurrence in the MLP improves accuracy across static, neuromorphic, and temporally complex datasets.
-
Neuromorphic Diffusion Language Models: Addressing Compute and Memory Bottlenecks via Sparsity and Block Denoising
N-MDLMs, spiking versions of block-diffusion language models, are estimated to improve throughput and energy efficiency, but the estimates rely on assumed event-driven hardware behavior and an internally inconsistent ...
-
Bio-SFT: Asymmetric Cortical Guidance and Retinal Adaptation for Robust HDR Reconstruction
Bio-SFT, a transformer with retinal adaptation, asymmetric high-to-low frequency guidance, and SNN hard gating, improves HDR-VDP-3 and ΔE_ITP on HDRTV1K with as few as 0.14M parameters.
-
Transformer-based EEG Decoding: A Survey
A survey that classifies Transformer-based EEG decoding models into backbone, hybrid, and customized categories and reviews their applications and limitations.
-
Spike-TBR: a Noise Resilient Neuromorphic Event Representation
Spike-TBR adds a spiking-neuron filter to the TBR event representation, making it robust to event-stream noise while preserving accuracy on clean data.
-
Energy-Efficient Deep Reinforcement Learning with Spiking Transformers
A spiking Transformer trained on A* maze demonstrations reaches 99.64% action accuracy on a custom 21x21 maze benchmark, but the energy-efficiency and baseline-comparison claims are not backed by proper experiments.
-
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
MDST++ reaches state-of-the-art audio-visual zero-shot learning on three benchmarks by converting RGB frames to synthetic events and processing them with a spiking transformer that shrinks timesteps.
-
Linguistics and Human Brain: A Perspective of Computational Neuroscience
A narrative review arguing that computational neuroscience, powered by LLM-based model–brain alignment, serves as the bridge between linguistic theory and neural data.
Discussion (0). Sign in to comment.