Pith. sign in

REVIEW 3 cited by

Spiking Vision Transformer with Saccadic Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12677 v1 pith:37GFKN2D submitted 2025-02-18 cs.CV cs.AI

Spiking Vision Transformer with Saccadic Attention

classification cs.CV cs.AI
keywords visionperformancesaccadicsnn-basedsssavitssnn-vitspike
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The combination of Spiking Neural Networks (SNNs) and Vision Transformers (ViTs) holds potential for achieving both energy efficiency and high performance, particularly suitable for edge vision applications. However, a significant performance gap still exists between SNN-based ViTs and their ANN counterparts. Here, we first analyze why SNN-based ViTs suffer from limited performance and identify a mismatch between the vanilla self-attention mechanism and spatio-temporal spike trains. This mismatch results in degraded spatial relevance and limited temporal interactions. To address these issues, we draw inspiration from biological saccadic attention mechanisms and introduce an innovative Saccadic Spike Self-Attention (SSSA) method. Specifically, in the spatial domain, SSSA employs a novel spike distribution-based method to effectively assess the relevance between Query and Key pairs in SNN-based ViTs. Temporally, SSSA employs a saccadic interaction module that dynamically focuses on selected visual areas at each timestep and significantly enhances whole scene understanding through temporal interactions. Building on the SSSA mechanism, we develop a SNN-based Vision Transformer (SNN-ViT). Extensive experiments across various visual tasks demonstrate that SNN-ViT achieves state-of-the-art performance with linear computational complexity. The effectiveness and efficiency of the SNN-ViT highlight its potential for power-critical edge vision applications.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Otters++: A Time-to-first-spike Based Energy Efficient Optical Spiking Transformer

    cs.AI 2026-06 unverdicted novelty 7.0

    Otters++ realizes TTFS via measured device decay in optical synapses, uses hybrid QNN-equivalent training with noise awareness, and reports 84.17% average GLUE score with energy gains over prior spiking transformers.

  2. Bias Redistribution in Visual Machine Unlearning: Does Forgetting One Group Harm Another?

    cs.LG 2026-04 unverdicted novelty 6.0

    Unlearning a demographic group in CLIP models redistributes bias primarily along gender boundaries rather than eliminating it.

  3. TEFormer: Structured Bidirectional Temporal Enhancement Modeling in Spiking Transformers

    cs.NE 2026-01 conditional novelty 6.0

    A spiking transformer with forward temporal EMA in attention and backward gated recurrence in the MLP improves accuracy across static, neuromorphic, and temporally complex datasets.