Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Training spiking networks with integers instead of binary spikes lets them match ANN accuracy at 86.2% top-1 on ImageNet.

desk verdict Big empirical gains for scaled SNNs, but the integer-to-spike equivalence is proven only for a single neuron and is not shown to compose through the attention layers. read the letter →

arxiv 2411.16061 v1 pith:JENYDP7Z submitted 2024-11-25 cs.CV

classification cs.CV
keywords SpikingneuralnetworkSpike-drivenTransformerNeuromorphiccomputingEfficientarchitectureandtrainingSpikeFiringApproximationMaskedimagemodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to close the accuracy gap between spiking neural networks (SNNs) and ordinary artificial neural networks (ANNs) without giving up the event-driven, addition-only inference that makes SNNs low-power. It identifies binary spike firing as the root flaw, and replaces it during training with integer-valued activations: a neuron outputs an integer in $[0,D]$ rather than a single bit, and at inference that integer is re-expanded into a train of $D$ spikes. On ImageNet-1k the resulting E-SpikeFormer family reaches 78.5%, 79.8%, 84.0%, and 86.2% top-1 accuracy at 10M, 19M, 83M, and 173M parameters, which the paper reports as the best results for directly trained spiking networks at these scales and comparable to classic ANN vision Transformer backbones. Training for static tasks is reduced to one timestep, giving about 4.5x faster ImageNet training and 3.9x lower inference energy at the 10M scale, and a spike-masked autoencoder with spike sparse convolution is added to keep accuracy from degrading as models scale. If the claim holds, SNNs become a practical low-power visual backbone rather than a research curiosity.

What carries the argument

The load-bearing object is the integer fire function $\operatorname{Fire}_D(U)=\lfloor\operatorname{clip}(U,0,D)\rceil$ together with the identity in Proposition 1: feeding the membrane potential $U$ to an integrate-and-fire neuron with soft reset and threshold 1, with nonzero input only at the first of $D$ timesteps, produces a spike train whose sum equals $\operatorname{Fire}_D(U)$. This is what lets the network be trained as a one-timestep integer network and then expanded into $D$ timesteps of spike-driven inference, so the energy advantage comes from sparse additions triggered only when spikes arrive. The second mechanism is the changed firing pattern it induces: SFA firing is asynchronous and concentrated at early timesteps, whereas ANN-to-SNN conversion and vanilla direct training fire randomly and need all timesteps to compute a rate. Supporting machinery includes the E-SpikeFormer block, built from SpikeSepConv and efficient spike-driven self-attention with a widened value branch, and the Spike Sparse Convolution used in masked autoencoding, which restricts convolution to unmasked positions to prevent information leakage.

What would settle it

Take a trained E-SpikeFormer and run the same ImageNet image through both the one-timestep integer forward pass used in training and the $D$-timestep spike deployment; if the equivalence composes, the two forward passes must give identical activations and the same prediction, and any systematic mismatch would show that the reported spike-deployed accuracy is not what the training objective optimized.

Watch

Extended reading notes

Core claim

The central claim is that binary firing is not just a quantization nuisance but a mechanistic defect in spiking neurons, harming both spatial representation (a spike cannot encode how strong the input was) and temporal dynamics (reset can only forget a fixed amount). The proposed cure, Spike Firing Approximation (SFA), trains with the integer fire function $\operatorname{Fire}_D(U)=\lfloor\operatorname{clip}(U,0,D)\rceil$ and deploys at inference by replacing each integer with $D$ binary spikes from an integrate-and-fire neuron with soft reset and threshold 1. Proposition 1 gives the per-neuron identity $S^l_D=\sum_{d=1}^{D}\hat{S}^l[d]$, and the paper argues that this makes a one-timestep integer forward pass equivalent to a $D$-timestep spike-driven forward pass for the network. On top of this the paper builds E-SpikeFormer, an efficient spike-driven Transformer that removes energy-hungry reparameterized convolutions, and a masked-image-modeling pretraining scheme with Spike Sparse Convolution that prevents the feature collapse, measured by effective rank, that scaling induces in binary-spike networks. The result, as the paper states, is that directly trained SNNs reach ANN-level accuracy while preserving the low-power, sparse-addition inference path.

Load-bearing premise

The method assumes that what holds for one neuron whose input arrives in a single moment also holds for a whole network whose layers receive spikes spread over many moments, but the proof only covers the single-neuron, single-moment case.

Editorial extensions

If this is right

  • Directly trained spiking networks can now reach 86.2% top-1 on ImageNet at 173M parameters, a higher accuracy than prior directly trained SNNs at this scale and comparable to several classical ANN backbones.
  • Static-image training collapses to one timestep, cutting ImageNet training time by about 4.5x at the 10M scale while inference still runs as a single image presentation followed by $D$ spike steps.
  • Inference power drops: the 10M E-SpikeFormer uses 3.0 mJ versus 11.9 mJ for the Meta-SpikeFormer baseline, and the spike firing rate decreases as timesteps advance, making later computation sparser.
  • SFA-trained models fit asynchronous neuromorphic chips, because the $D$ spikes can be emitted in a short window without a global clock, unlike vanilla multi-timestep direct training.
  • Combining SFA with masked-image-modeling pretraining and Spike Sparse Convolution avoids the performance degradation that still occurs when binary-spike networks are scaled with the same MIM strategy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extension: Because $\operatorname{Fire}_D$ is uniform $D$-level quantization, SFA suggests a broader recipe: any ANN trained with $D$-level integer activations and threshold-normalized weights could be deployed as a spike-driven network, putting spike-driven efficiency within reach of standard quantization-aware training.
  • Extension: The paper's firing-pattern analysis implies $D$ can be treated as a tunable accuracy-latency dial, so an adaptive version that spends fewer timesteps on easy inputs or early layers could extend the reported $1\times4$ versus $1\times8$ trade-off beyond the fixed settings tested.
  • Extension: The effective-rank view of why binary networks fail to scale points to a practical diagnostic: monitor the encoder's effective rank during masked pretraining as an early indicator of downstream fine-tuning quality, potentially guiding architecture or mask-ratio choices without full fine-tuning.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes Spike Firing Approximation (SFA), a training method that replaces binary spike activations with integer-valued activations (Fire_D) during training and converts them to spike trains during inference via IF-SR neurons with threshold 1. The authors also introduce an efficient spike-driven Transformer architecture (E-SpikeFormer) and a masked image modeling pre-training strategy with spike sparse convolution. They report state-of-the-art top-1 accuracies on ImageNet-1k at scales from 10M to 173M parameters (78.5% to 86.2%), with large improvements in training time and inference energy over prior SNNs, and validate the method on object detection, semantic segmentation, and neuromorphic action recognition. The central theoretical claim is that one-timestep integer training is losslessly equivalent to multi-timestep spike-driven inference via Proposition 1.

Significance. If the claimed training-inference equivalence and the reported results hold, this work would be a significant step toward scaling SNNs to practical vision backbones while retaining the low-power advantages of event-driven computation. The paper's strengths include extensive experiments across multiple tasks and model scales, concrete efficiency gains (e.g., 4.5x training acceleration, 3.9x inference energy improvement at the 10M scale), and a stated public code release. The analysis of spike firing patterns (asynchronous versus synchronous) and the discussion of neuromorphic chip implementation are also valuable. However, the central equivalence proof is limited to a single neuron under an unrealistic assumption, and the extension to whole deep networks, especially the nonlinear attention module, is not established. This gap undermines the claim that the reported inference accuracies follow from the training objective, so the theoretical foundation needs substantial strengthening before the results can be fully credited.

major comments (3)
  1. [Section 4.1, Proposition 1 and Fig. 6] The proof of the integer-to-spike equivalence assumes that for each neuron 'the input is zero at timesteps d = 2, ..., D' (after Eq. (16)). In a deep network this premise is violated for every layer after the first: each layer receives spikes from the previous layer across all D timesteps. The paper asserts, rather than proves, that one-timestep integer training equals D-timestep spike-driven inference for the whole network. For purely linear layers with nonnegative inputs the total spike count of an IF-SR neuron may be timing-invariant, but negative weights (which can occur) break this, and the E-SDSA module (Section 3.3) is quadratic: integer training computes a function of (sum_d Q[d], sum_d K[d], sum_d V[d]), while spike-driven inference accumulates sum_d (Q[d] K[d]^T V[d]); the cross-timestep products with d != e are not present in the inference computation, and the paper provides no argument that they vanish or are negligible. Please provide an end-to-end equivalence proof for linear layers with arbitrary weights and for the attention module, or, failing that, an experimental comparison of the one-step integer-trained model's accuracy versus the implemented D-step spike-driven inference accuracy for each model scale in Table 1.
  2. [Section 4.2, Eqs. (26)-(28)] The gradient-error derivation is dimensionally inconsistent. Eq. (27) defines Err^l as the integral over U of (Rect[0,D](U) - (1/D) round(clip(U,0,D))); the result of this integral is a constant (or divergent), but Eq. (28) reports a piecewise function of U, which is the integrand, not the integral. Consequently, the formal conclusion that larger D reduces the gradient error is not supported by the equations as written. Please correct the derivation or clarify what quantity is actually being computed; the empirical trend in the experiments may still hold, but the theoretical analysis needs repair.
  3. [Section 4.1, Eqs. (18)-(20)] The step from Eq. (19) to Eq. (20), where the 0.5 bias introduced by the round function is said to be 'incorporated into the weight,' is not formalized for a deep network. If every neuron's activation function is changed from round(clip(U,0,D)) to floor(clip(U,0,D)) by adding a constant 0.5 to the neuron's input, the network function changes unless the biases of the preceding layers are adjusted consistently. For the equivalence to be claimed, either the reparameterization must be specified explicitly for all layers, or the Fire_D function should be defined with floor from the outset.
minor comments (6)
  1. [Section 4.1, Eq. (16)] The notation in Eq. (16) appears to use {S^l[d]}_D for the spike train generated by IF-SR, but the output of IF-SR is the spike train {hat S^l[d]}_D; please correct the notation for consistency.
  2. [Fig. 3] The label 'al D = 0.3' in Fig. 3 seems to be a typo; it should likely read 'a^l_D = 0.3'.
  3. [Table 1] The row 'Spike-dirven Transformer' contains a typo; it should be 'Spike-driven Transformer'.
  4. [Section 3.4] The phrase 'K d=32' appears to be a formatting error; please clarify whether this is intended to denote the kernel size K^d = 32.
  5. [Section 4.2] Definitions 1 and 2 both use the symbol Err^l for different quantities (forward approximation error and backward gradient error); please use distinct symbols such as Err_fwd and Err_bwd to avoid ambiguity.
  6. [References, [24]] Reference [24] is a closely related prior work on integer-valued training and spike-driven inference by the same group; the paper cites it but does not discuss its relationship to the proposed SFA method. Please clarify the differences and the specific novelties introduced here.

Circularity Check

1 steps flagged · score 2.0 of 10

No significant circularity; the integer-to-spike equivalence is a designed identity, and the headline results are externally benchmarked.

  1. self definitional [Section 4.1, Proposition 1 (Eqs. 15-21)]
    "Feeding the membrane potential input of Eq. (5) into an IF-SR neuron with Vth = 1, the spike train generated is: {S^l[d]}_D = IF-SR(U^l), ... The underlying assumption is that there is non-zero input only at timestep d = 1 ... Then the sum of the spike train is: S^l_D = floor(clip{U^l, 0, D}) ... Combining Eq. (17) and Eq. (20), we have: S^l_D = floor(clip{U^l, 0, D}) = sum_{d=1}^D S^l[d]."

    Eq. (5) defines Fire_D as round(clip(U,0,D)); under the proof's assumption that nonzero input arrives only at d=1, an IF-SR neuron with Vth=1 fires floor(clip(U,0,D)) spikes, and the 0.5 difference between round and floor is absorbed into the weights. Hence S^l_D = sum_d S^l[d] holds by construction, not by an independent derivation. This designed quantizer-match is legitimate for a single neuron, but the paper extends it to whole networks by asserting Equivalence in Fig. 6 without proving that layer-wise spike-count identities compose through E-SDSA's quadratic attention, residual connections, and BN folding.

full rationale

The central SFA construction is not circular in the harmful sense: Fire_D is deliberately defined as the number of spikes an IF-SR neuron would emit, so Proposition 1 is a designed quantizer-to-spike identity rather than a fitted input disguised as a prediction. The ImageNet, COCO, ADE20K, and HAR-DVS results are measured on spike-driven inference and benchmarked against prior SNNs and ANNs, so the headline accuracies are external evidence, not self-confirming. Two concerns lower, but do not destroy, the paper's independent content. First, the layer-wise equivalence does not automatically compose through the nonlinear E-SDSA attention and residual paths; Fig. 6 asserts whole-network equivalence without a proof, leaving a correctness gap between the integer training objective and the deployed spike-driven computation. That is a verifiability issue, not a circular fit. Second, reference [24] by the same group already proposed integer-valued training and spike-driven inference for object detection, and the body introduces SFA without explicitly flagging that prior introduction; this is an attribution and novelty concern rather than load-bearing circularity. Overall, the paper's contributions are externally validated and not reduced to their inputs by definition, so the circularity score is low.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical entities. Its free parameters are mainly D, the integer width that directly controls accuracy and power, plus the MIM masking ratio and the E-SDSA channel expansion factor. The assumptions are mostly domain-level: rate coding, the deep-network equivalence of integer training, the binary-firing flaw narrative, the energy model, and the effective-rank measure. The deep-network equivalence is the least secure premise because the proof covers only a single layer with a zero-input assumption. The energy model is also a notable caveat because all power claims depend on it.

free parameters (3)
  • D (wave width / max integer activation) = 4 and 8 in main experiments
    D controls the number of inference timesteps and the maximum integer activation. The 83M model improves from 83.2% at D=4 to 84.0% at D=8, at increased power, so D is a manually chosen hyperparameter that directly affects the central results.
  • MIM masking ratio mu = not reported in main text
    The masking ratio in Eq. (12) determines the difficulty of the masked image modeling pretraining. Its value is deferred to supplementary Section S5, yet it affects pretraining quality and final accuracy.
  • E-SDSA channel expansion factor gamma = not reported in main text
    The iRMB-style expansion of the V channel dimension in the E-SDSA module is a design choice made to compensate for removing RepConv. The specific value is not stated in the main text, but it changes the parameter count and representational capacity.
assumptions (5)
  • domain assumption Spiking neurons transmit spike firing rates across spatial dimensions (Eq. 4).
    The entire SFA method begins by assuming that the signal between neurons is the firing rate al_D = (1/D) sum S^l[d]. This rate-coding assumption is stated in Section 3.2 but not derived from any biological or theoretical principle.
  • ad hoc to paper One-timestep integer training is equivalent to multi-timestep spike inference for the full deep network.
    Proposition 1 proves the equivalence only for a single layer with nonzero input at the first timestep. The extension to deep networks is asserted through Fig. 6 and the text around Eq. (7), but not proved. This is the load-bearing premise for the reported inference accuracy.
  • domain assumption Binary spike firing is a fundamental mechanistic flaw that impairs spatial representation and temporal dynamics.
    Section 3.1.2 frames binary firing as an intrinsic defect. The statement is qualitative and motivates SFA, but it is not a formal theorem and does not enter the equivalence proof.
  • domain assumption Power consumption can be estimated by counting spike-driven additions and multiplying by fixed per-operation energies.
    The power numbers in Tables 1-3 and the energy-efficiency claims depend on a per-operation energy model rather than measurements from an actual neuromorphic chip. The model details are placed in supplementary Section S1, which was not part of the reviewed text.
  • domain assumption Effective rank of the encoder output measures feature quality and predicts scaling behavior.
    Section 5.4 uses the effective rank in Eq. (29) to explain why binary spikes hurt scaling. The connection between effective rank and final accuracy is argued, not proven.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training." pith.science (2026). https://pith.science/paper/JENYDP7Z

@misc{pith2026241116061,
  author       = {Pith},
  title        = {Pith review of: Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JENYDP7Z}},
  note         = {Machine review of arXiv:2411.16061}
}
abstract

The ambition of brain-inspired Spiking Neural Networks (SNNs) is to become a low-power alternative to traditional Artificial Neural Networks (ANNs). This work addresses two major challenges in realizing this vision: the performance gap between SNNs and ANNs, and the high training costs of SNNs. We identify intrinsic flaws in spiking neurons caused by binary firing mechanisms and propose a Spike Firing Approximation (SFA) method using integer training and spike-driven inference. This optimizes the spike firing pattern of spiking neurons, enhancing efficient training, reducing power consumption, improving performance, enabling easier scaling, and better utilizing neuromorphic chips. We also develop an efficient spike-driven Transformer architecture and a spike-masked autoencoder to prevent performance degradation during SNN scaling. On ImageNet-1k, we achieve state-of-the-art top-1 accuracy of 78.5\%, 79.8\%, 84.0\%, and 86.2\% with models containing 10M, 19M, 83M, and 173M parameters, respectively. For instance, the 10M model outperforms the best existing SNN by 7.2\% on ImageNet, with training time acceleration and inference energy efficiency improved by 4.5$\times$ and 3.9$\times$, respectively. We validate the effectiveness and efficiency of the proposed method across various tasks, including object detection, semantic segmentation, and neuromorphic vision tasks. This work enables SNNs to match ANN performance while maintaining the low-power advantage, marking a significant step towards SNNs as a general visual backbone. Code is available at https://github.com/BICLab/Spike-Driven-Transformer-V3.

Figures

Figures reproduced from arXiv: 2411.16061 by the authors.

Figure 1
Figure 1. E-SpikeFormer versus other spiking Transformers on [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The impact of the reset mechanism on the spatio-temporal dynamics of spiking neurons. (a) The [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Spike firing patterns. Assuming an approximation [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: The overview of E-SpikeFormer. In general, we follow the design of Meta-SpikeFormer [ [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: The overview of MIM pre-train in E-SpikeFormer. It consists of a SNN encoder and a ANN decoder. The encoder [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Comparison of vanilla and SFA training in SNNs. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Vanilla and SFA training time on the ImageNet with [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Visualization of spike distribution in vanilla and SFA training on ImageNet. From top to bottom: vanilla and SFA [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Network Spike Firing Rate (NSFR) of vanilla ( [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Binary spike firing interferes with the scaling of [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 12
Figure 12. Figure 12: Comparison of Spike Sparse Conv (SSC) and Vanilla [PITH_FULL_IMAGE:figures/full_fig_p014_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantized Spike-driven Transformer

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 4-bit quantized spike-driven transformer with multi-bit training and binary inference achieves 80.3% ImageNet accuracy with 6.8M parameters.

  2. Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A spiking Mask2Former with normalized integer neurons achieves state-of-the-art SNN performance on ADE20K, CityScapes, and Pascal VOC, with claimed large energy savings.

Reference graph

Works this paper leans on

105 extracted references · 56 canonical work pages · cited by 2 Pith papers

  1. [1]

    Towards spike-based machine intelligence with neuromorphic computing,

    K. Roy, A. Jaiswal, and P . Panda, “Towards spike-based machine intelligence with neuromorphic computing,” Nature, vol. 575, no. 7784, pp. 607–617, 2019

  2. [2]

    Opportunities for neuromorphic computing algorithms and applications,

    C. D. Schuman, S. R. Kulkarni, M. Parsa, J. P . Mitchell, B. Kay et al. , “Opportunities for neuromorphic computing algorithms and applications,” Nature Computational Science, vol. 2, no. 1, pp. 10–19, 2022

  3. [3]

    Networks of spiking neurons: The third generation of neural network models,

    W. Maass, “Networks of spiking neurons: The third generation of neural network models,” Neural Networks , vol. 10, no. 9, pp. 1659–1671, 1997

  4. [4]

    A million spiking-neuron integrated circuit with a scalable communication network and interface,

    P . A. Merolla, J. V . Arthur, R. Alvarez-Icaza, A. S. Cassidy, J. Sawada, F. Akopyan, B. L. Jackson, N. Imam, C. Guo, Y. Naka- mura et al., “A million spiking-neuron integrated circuit with a scalable communication network and interface,” Science, vol. 345, no. 6197, pp. 668–673, 2014

  5. [5]

    Loihi: A neuromorphic manycore processor with on-chip learning,

    M. Davies, N. Srinivasa, T.-H. Lin, G. Chinya, Y. Cao, S. H. Choday, G. Dimou, P . Joshi, N. Imam, S. Jain et al. , “Loihi: A neuromorphic manycore processor with on-chip learning,” IEEE Micro, vol. 38, no. 1, pp. 82–99, 2018

  6. [6]

    Towards artificial general intelligence with hybrid tianjic chip architecture,

    J. Pei, L. Deng et al., “Towards artificial general intelligence with hybrid tianjic chip architecture,” Nature, vol. 572, no. 7767, pp. 106–111, 2019

  7. [7]

    Accurate and efficient time- domain classification with adaptive spiking recurrent neural networks,

    B. Yin, F. Corradi, and S. M. Boht ´e, “Accurate and efficient time- domain classification with adaptive spiking recurrent neural networks,” Nature Machine Intelligence, vol. 3, no. 10, pp. 905–913, 2021

  8. [8]

    A long short-term memory for ai applications in spike-based neuromorphic hard- ware,

    A. Rao, P . Plank, A. Wild, and W. Maass, “A long short-term memory for ai applications in spike-based neuromorphic hard- ware,” Nature Machine Intelligence, vol. 4, no. 5, pp. 467–479, 2022

Show all 105 references
  1. [9]

    Neurozoom: Denoising and super resolving neuromor- phic events and spikes,

    P . Duan, Y. Ma, X. Zhou, X. Shi, Z. W. Wang, T. Huang, and B. Shi, “Neurozoom: Denoising and super resolving neuromor- phic events and spikes,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 12, pp. 15 219–15 232, 2023

  2. [10]

    Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip,

    M. Yao, O. Richter, G. Zhao, N. Qiao, Y. Xing, D. Wang, T. Hu, W. Fang, T. Demirci, M. De Marchi, L. Deng, T. Yan, C. Nielsen, S. Sheik, C. Wu, Y. Tian, B. Xu, and G. Li, “Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip,” Nature Communicatio...

  3. [11]

    Spiking deep residual networks,

    Y. Hu, H. Tang, and G. Pan, “Spiking deep residual networks,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 8, pp. 5200–5205, 2023

  4. [12]

    Enabling deep spiking neural networks with hybrid conversion and spike tim- ing dependent backpropagation,

    N. Rathi, G. Srinivasan, P . Panda, and K. Roy, “Enabling deep spiking neural networks with hybrid conversion and spike tim- ing dependent backpropagation,” in International Conference on Learning Representations, 2020

  5. [14]

    Progressive tandem learning for pattern recognition with deep spiking neural networks,

    J. Wu, C. Xu, X. Han, D. Zhou, M. Zhang, H. Li, and K. C. Tan, “Progressive tandem learning for pattern recognition with deep spiking neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, pp. 7824–7840, 2021

  6. [15]

    Masked spiking transformer,

    Z. Wang, Y. Fang, J. Cao, Q. Zhang, Z. Wang, and R. Xu, “Masked spiking transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 1761–1771

  7. [16]

    Spatio-temporal backpropagation for training high-performance spiking neural networks,

    Y. Wu, L. Deng, G. Li, J. Zhu, and L. Shi, “Spatio-temporal backpropagation for training high-performance spiking neural networks,” Frontiers in Neuroscience, vol. 12, p. 331, 2018

  8. [17]

    Training spiking neural networks using lessons from deep learning,

    J. K. Eshraghian, M. Ward, E. O. Neftci, X. Wang, G. Lenz, G. Dwivedi, M. Bennamoun, D. S. Jeong, and W. D. Lu, “Training spiking neural networks using lessons from deep learning,” Proceedings of the IEEE, vol. 111, no. 9, pp. 1016–1054, 2023

  9. [18]

    Deep residual learning in spiking neural networks,

    W. Fang, Z. Yu, Y. Chen, T. Huang, T. Masquelier, and Y. Tian, “Deep residual learning in spiking neural networks,” Advances in Neural Information Processing Systems , vol. 34, pp. 21 056–21 069, 2021

  10. [19]

    Going deeper with directly-trained larger spiking neural networks,

    H. Zheng, Y. Wu, L. Deng, Y. Hu, and G. Li, “Going deeper with directly-trained larger spiking neural networks,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 12, 2021, pp. 11 062–11 070

  11. [20]

    Vtsnn: a virtual temporal spiking neural net- work,

    X.-R. Qiu, Z.-R. Wang, Z. Luan, R.-J. Zhu, X. Wu, M.-L. Zhang, and L.-J. Deng, “Vtsnn: a virtual temporal spiking neural net- work,” Frontiers in Neuroscience, vol. 17, p. 1091097, 2023

  12. [21]

    Attention spiking neural networks,

    M. Yao, G. Zhao, H. Zhang, Y. Hu, L. Deng, Y. Tian, B. Xu, and G. Li, “Attention spiking neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 8, pp. 9393– 9410, 2023

  13. [22]

    Tensor decomposition based attention module for spiking neu- ral networks,

    H. Deng, R. Zhu, X. Qiu, Y. Duan, M. Zhang, and L.-J. Deng, “Tensor decomposition based attention module for spiking neu- ral networks,” Knowledge-Based Systems, vol. 295, p. 111780, 2024

  14. [23]

    Rsc- snn: Exploring the trade-off between adversarial robustness and accuracy in spiking neural networks via randomized smoothing coding,

    K. Wu, M. Yao, Y. Chou, X. Qiu, R. Yang, B. Xu, and G. Li, “Rsc- snn: Exploring the trade-off between adversarial robustness and accuracy in spiking neural networks via randomized smoothing coding,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, p...

  15. [24]

    Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection,

    X. Luo, M. Yao, Y. Chou, B. Xu, and G. Li, “Integer-valued training and spike-driven inference spiking neural network for high-performance and energy-efficient object detection,” arXiv preprint arXiv:2407.20708, 2024

  16. [25]

    Spikingjelly: An open- source machine learning infrastructure platform for spike-based intelligence,

    W. Fang, Y. Chen, J. Ding, Z. Yu, T. Masquelier, D. Chen, L. Huang, H. Zhou, G. Li, and Y. Tian, “Spikingjelly: An open- source machine learning infrastructure platform for spike-based intelligence,” Science Advances, vol. 9, no. 40, p. eadi1480, 2023

  17. [26]

    Im-loss: information maximization loss for spiking neural net- works,

    Y. Guo, Y. Chen, L. Zhang, X. Liu, Y. Wang, X. Huang, and Z. Ma, “Im-loss: information maximization loss for spiking neural net- works,” Advances in Neural Information Processing Systems, vol. 35, pp. 156–166, 2022

  18. [27]

    Rethinking the performance comparison between snns and anns,

    L. Deng, Y. Wu, X. Hu, L. Liang, Y. Ding, G. Li, G. Zhao, P . Li, and Y. Xie, “Rethinking the performance comparison between snns and anns,” Neural Networks, vol. 121, pp. 294–307, 2020

  19. [28]

    Spike-driven transformer v2: Meta spiking neural network ar- chitecture inspiring the design of next-generation neuromorphic chips,

    M. Yao, J. Hu, T. Hu, Y. Xu, Z. Zhou, Y. Tian, B. XU, and G. Li, “Spike-driven transformer v2: Meta spiking neural network ar- chitecture inspiring the design of next-generation neuromorphic chips,” in The Twelfth International Conference on Learning Repre- sentations, 2024

  20. [29]

    Spike- driven transformer,

    M. Yao, J. Hu, Z. Zhou, L. Yuan, Y. Tian, B. Xu, and G. Li, “Spike- driven transformer,” in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 64 043–64 058

  21. [30]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y. Li, P . Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009

  22. [31]

    Designing bert for convolutional networks: Sparse and hierarchical masked FOR REVIEW 16 modeling,

    K. Tian, Y. Jiang, C. Lin, L. Wang, Z. Yuan et al. , “Designing bert for convolutional networks: Sparse and hierarchical masked FOR REVIEW 16 modeling,” in The Eleventh International Conference on Learning Representations, 2022

  23. [32]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2009, pp. 248–255

  24. [33]

    Microsoft coco: Common objects in context,

    T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P . Perona, D. Ramanan, P . Doll´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision . Springer, 2014, pp. 740–755

  25. [34]

    Scene parsing through ade20k dataset,

    B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 633–641

  26. [35]

    Hardvs: Revisiting human activity recognition with dynamic vision sensors,

    X. Wang, Z. Wu, B. Jiang, Z. Bao, L. Zhu, G. Li, Y. Wang, and Y. Tian, “Hardvs: Revisiting human activity recognition with dynamic vision sensors,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 6, 2024, pp. 5615–5623

  27. [36]

    Online training through time for spiking neural networks,

    M. Xiao, Q. Meng, Z. Zhang, D. He, and Z. Lin, “Online training through time for spiking neural networks,” Advances in Neural Information Processing Systems, vol. 35, pp. 20 717–20 730, 2022

  28. [37]

    Towards memory-and time-efficient backpropagation for train- ing spiking neural networks,

    Q. Meng, M. Xiao, S. Yan, Y. Wang, Z. Lin, and Z.-Q. Luo, “Towards memory-and time-efficient backpropagation for train- ing spiking neural networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 6166–6176

  29. [38]

    Rethinking pretraining as a bridge from anns to snns,

    Y. Lin, Y. Hu, S. Ma, D. Yu, and G. Li, “Rethinking pretraining as a bridge from anns to snns,” IEEE Transactions on Neural Networks and Learning Systems, pp. 1–14, 2022

  30. [39]

    Memory-efficient reversible spiking neural networks,

    H. Zhang and Y. Zhang, “Memory-efficient reversible spiking neural networks,” in Proceedings of the AAAI Conference on Ar- tificial Intelligence, vol. 38, no. 15, 2024, pp. 16 759–16 767

  31. [40]

    High-performance temporal reversible spiking neural networks with o(l) training memory and o(1) inference cost,

    J. Hu, M. Yao, X. Qiu, Y. Chou, Y. Cai, N. Qiao, Y. Tian, B. Xu, and G. Li, “High-performance temporal reversible spiking neural networks with o(l) training memory and o(1) inference cost,” arXiv preprint arXiv:2405.16466, 2024

  32. [41]

    Towards ultra low latency spiking neural networks for vision and sequential tasks using temporal pruning,

    S. S. Chowdhury, N. Rathi, and K. Roy, “Towards ultra low latency spiking neural networks for vision and sequential tasks using temporal pruning,” in European Conference on Computer Vision. Springer, 2022, pp. 709–726

  33. [42]

    Seenn: Towards temporal spiking early exit neural networks,

    Y. Li, T. Geller, Y. Kim, and P . Panda, “Seenn: Towards temporal spiking early exit neural networks,” in Advances in Neural Infor- mation Processing Systems, vol. 36. Curran Associates, Inc., 2023, pp. 63 327–63 342

  34. [43]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778

  35. [44]

    Identity mappings in deep residual networks,

    K. He, X. Zhang, and S. Ren, “Identity mappings in deep residual networks,” in European conference on computer vision . Springer, 2016, pp. 630–650

  36. [45]

    Advancing spiking neural networks toward deep residual learning,

    Y. Hu, L. Deng, Y. Wu, M. Yao, and G. Li, “Advancing spiking neural networks toward deep residual learning,” IEEE Transac- tions on Neural Networks and Learning Systems , pp. 1–15, 2024

  37. [46]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017

  38. [47]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Repre- sentations, 2020

  39. [48]

    Spike transformer: Monocular depth estimation for spiking camera,

    J. Zhang, L. Tang, Z. Yu, J. Lu, and T. Huang, “Spike transformer: Monocular depth estimation for spiking camera,” in European Conference on Computer Vision. Springer, 2022, pp. 34–52

  40. [49]

    Com- plex dynamic neurons improved spiking transformer network for efficient automatic speech recognition,

    Q. Wang, T. Zhang, M. Han, Y. Wang, D. Zhang, and B. Xu, “Com- plex dynamic neurons improved spiking transformer network for efficient automatic speech recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 1, 2023, pp. 102–109

  41. [50]

    Spikformer: When spiking neural network meets transformer,

    Z. Zhou, Y. Zhu, C. He, Y. Wang, S. YAN, Y. Tian, and L. Yuan, “Spikformer: When spiking neural network meets transformer,” in The Eleventh International Conference on Learning Representations, 2023

  42. [51]

    Metaformer is actually what you need for vision,

    W. Yu, M. Luo, P . Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 10 819–10 829

  43. [52]

    MetaLA: Unified optimal linear approximation to softmax attention map,

    Y. Chou, M. Yao, K. Wang, Y. Pan, R.-J. Zhu, J. Wu, Y. Zhong, Y. Qiao, B. XU, and G. Li, “MetaLA: Unified optimal linear approximation to softmax attention map,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems ,

  44. [53]

    Bert: Pre- training of deep bidirectional transformers for language under- standing,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre- training of deep bidirectional transformers for language under- standing,” arXiv preprint arXiv:1810.04805, 2018

  45. [54]

    Beit: Bert pre-training of image transformers,

    H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: Bert pre-training of image transformers,” arXiv preprint arXiv:2106.08254, 2021

  46. [55]

    Image bert pre-training with online tokenizer,

    J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong, “Image bert pre-training with online tokenizer,” in International Conference on Learning Representations, 2021

  47. [56]

    Masked feature prediction for self-supervised visual pre- training,

    C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichten- hofer, “Masked feature prediction for self-supervised visual pre- training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 14 668–14 678

  48. [57]

    Mcmae: Masked convolution meets masked autoencoders,

    P . Gao, T. Ma, H. Li, Z. Lin, J. Dai, and Y. Qiao, “Mcmae: Masked convolution meets masked autoencoders,” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran Associates, Inc., 2022, pp...

  49. [58]

    Convnext v2: Co-designing and scaling convnets with masked autoencoders,

    S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, and S. Xie, “Convnext v2: Co-designing and scaling convnets with masked autoencoders,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2023, pp. 16 133– 16 142

  50. [59]

    Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks,

    E. O. Neftci, H. Mostafa, and F. Zenke, “Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based optimization to spiking neural networks,” IEEE Signal Processing Magazine, vol. 36, no. 6, pp. 51–63, 2019

  51. [60]

    Gerstner, W

    W. Gerstner, W. M. Kistler, R. Naud, and L. Paninski, Neuronal dynamics: From single neurons to networks and models of cognition . Cambridge University Press, 2014

  52. [61]

    Neuromorphic silicon neuron cir- cuits,

    G. Indiveri, B. Linares-Barranco, T. J. Hamilton, A. v. Schaik, R. Etienne-Cummings, T. Delbruck, S.-C. Liu, P . Dudek, P . H¨afliger, S. Renaud et al. , “Neuromorphic silicon neuron cir- cuits,” Frontiers in Neuroscience, vol. 5, p. 73, 2011

  53. [62]

    Parallel spiking neurons with high efficiency and ability to learn long-term dependencies,

    W. Fang, Z. Yu, Z. Zhou, D. Chen, Y. Chen, Z. Ma, T. Masquelier, and Y. Tian, “Parallel spiking neurons with high efficiency and ability to learn long-term dependencies,” in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 53 674–53 687

  54. [63]

    Spiking neural networks with im- proved inherent recurrence dynamics for sequential learning,

    W. Ponghiran and K. Roy, “Spiking neural networks with im- proved inherent recurrence dynamics for sequential learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 7, 2022, pp. 8001–8008

  55. [64]

    Is conventional snn really efficient? a perspective from network quantization,

    G. Shen, D. Zhao, T. Li, J. Li, and Y. Zeng, “Is conventional snn really efficient? a perspective from network quantization,” arXiv preprint arXiv:2311.10802, 2023

  56. [65]

    Liaf-net: Leaky integrate and analog fire network for lightweight and ef- ficient spatiotemporal information processing,

    Z. Wu, H. Zhang, Y. Lin, G. Li, M. Wang, and Y. Tang, “Liaf-net: Leaky integrate and analog fire network for lightweight and ef- ficient spatiotemporal information processing,” IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 11, pp. 6249– 6262, 2021

  57. [66]

    Direct training for spiking neural networks: Faster, larger, better,

    Y. Wu, L. Deng, G. Li, J. Zhu, Y. Xie, and L. Shi, “Direct training for spiking neural networks: Faster, larger, better,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 1311–1318

  58. [67]

    Diet-snn: A low-latency spiking neural network with direct input encoding and leakage and threshold optimization,

    N. Rathi and K. Roy, “Diet-snn: A low-latency spiking neural network with direct input encoding and leakage and threshold optimization,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 6, pp. 3174–3182, 2023

  59. [68]

    Event-driven learning for spiking neural networks,

    W. Wei, M. Zhang, J. Zhang, A. Belatreche, J. Wu, Z. Xu, X. Qiu, H. Chen, Y. Yang, and H. Li, “Event-driven learning for spiking neural networks,” arXiv preprint arXiv:2403.00270, 2024

  60. [69]

    Mobilenetv2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 4510–4520

  61. [70]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift,

    S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Interna- tional conference on machine learning. PMLR, 2015, pp. 448–456

  62. [71]

    Xception: Deep learning with depthwise separable convolutions,

    F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 1251–1258

  63. [72]

    Efficientvit: Enhanced linear at- tention for high-resolution low-computation visual recognition,

    H. Cai, C. Gan, and S. Han, “Efficientvit: Enhanced linear at- tention for high-resolution low-computation visual recognition,” arXiv preprint arXiv:2205.14756, 2022. FOR REVIEW 17

  64. [73]

    Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features,

    S. N. Wadekar and A. Chaurasia, “Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features,” arXiv preprint arXiv:2209.15159, 2022

  65. [74]

    Rethinking mobile block for efficient attention-based models,

    J. Zhang, X. Li, J. Li, L. Liu, Z. Xue, B. Zhang, Z. Jiang, T. Huang, Y. Wang, and C. Wang, “Rethinking mobile block for efficient attention-based models,” in Proceedings of IEEE/CVF International Conference on Computer Vision, 2023, pp. 1389–1400

  66. [75]

    Training data-efficient image transformers & distil- lation through attention,

    H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J ´egou, “Training data-efficient image transformers & distil- lation through attention,” in International Conference on Machine Learning. PMLR, 2021, pp. 10 347–10 357

  67. [76]

    How mask matters: Towards theoretical understandings of masked autoencoders,

    Q. Zhang, Y. Wang, and Y. Wang, “How mask matters: Towards theoretical understandings of masked autoencoders,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 127–27 139, 2022

  68. [77]

    Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,

    W. Wang, E. Xie, X. Li, D.-P . Fan, K. Song, D. Liang, T. Lu, P . Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 568–578

  69. [78]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10 012–10 022

  70. [79]

    Metaformer baselines for vision,

    W. Yu, C. Si, P . Zhou, M. Luo, Y. Zhou, J. Feng, S. Yan, and X. Wang, “Metaformer baselines for vision,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023

  71. [80]

    Tokens-to-token vit: Training vision transformers from scratch on imagenet,

    L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z.-H. Jiang, F. E. Tay, J. Feng, and S. Yan, “Tokens-to-token vit: Training vision transformers from scratch on imagenet,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 558–567

  72. [81]

    Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,

    X. Ding, X. Zhang, J. Han, and G. Ding, “Scaling up your kernels to 31x31: Revisiting large kernel design in cnns,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2022, pp. 11 963–11 975

  73. [82]

    Focal self-attention for local-global interactions in vision transformers,

    J. Yang, C. Li, P . Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao, “Focal self-attention for local-global interactions in vision transformers,” arXiv preprint arXiv:2107.00641, 2021

  74. [83]

    A convnet for the 2020s,

    Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 11 976–11 986

  75. [84]

    Bottleneck transformers for visual recognition,

    A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P . Abbeel, and A. Vaswani, “Bottleneck transformers for visual recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 16 519–16 529

  76. [85]

    Crossformer++: A versatile vision transformer hinging on cross-scale attention,

    W. Wang, W. Chen, Q. Qiu, L. Chen, B. Wu, B. Lin, X. He, and W. Liu, “Crossformer++: A versatile vision transformer hinging on cross-scale attention,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 5, pp. 3123–3136, 2024

  77. [86]

    Opti- mal ann-snn conversion for high-accuracy and ultra-low-latency spiking neural networks,

    T. Bu, W. Fang, J. Ding, P . DAI, Z. Yu, and T. Huang, “Opti- mal ann-snn conversion for high-accuracy and ultra-low-latency spiking neural networks,” in International Conference on Learning Representations, 2021

  78. [87]

    Fast-snn: Fast spiking neural network by converting quantized ann,

    Y. Hu, Q. Zheng, X. Jiang, and G. Pan, “Fast-snn: Fast spiking neural network by converting quantized ann,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 12, pp. 14 546–14 562, 2023

  79. [88]

    Toward high-accuracy and low-latency spiking neural networks with two-stage optimization,

    Z. Wang, Y. Zhang, S. Lian, X. Cui, R. Yan, and H. Tang, “Toward high-accuracy and low-latency spiking neural networks with two-stage optimization,” IEEE Transactions on Neural Networks and Learning Systems, pp. 1–15, 2023

  80. [89]

    Temporal efficient training of spiking neural network via gradient re-weighting,

    S. Deng, Y. Li, S. Zhang, and S. Gu, “Temporal efficient training of spiking neural network via gradient re-weighting,” in Interna- tional Conference on Learning Representations, 2022

  81. [90]

    Training high-performance low-latency spiking neural networks by differentiation on spike representation,

    Q. Meng, M. Xiao, S. Yan, Y. Wang, Z. Lin, and Z.-Q. Luo, “Training high-performance low-latency spiking neural networks by differentiation on spike representation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 444–12 453

  82. [91]

    Gated attention coding for training high-performance and efficient spik- ing neural networks,

    X. Qiu, R.-J. Zhu, Y. Chou, Z. Wang, L.-j. Deng, and G. Li, “Gated attention coding for training high-performance and efficient spik- ing neural networks,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 1, 2024, pp. 601–610

  83. [92]

    Spikformer v2: Join the high accuracy club on imagenet with an snn ticket,

    Z. Zhou, K. Che, W. Fang, K. Tian, Y. Zhu, S. Yan, Y. Tian, and L. Yuan, “Spikformer v2: Join the high accuracy club on imagenet with an snn ticket,” arXiv preprint arXiv:2401.02020, 2024

  84. [93]

    Cifar10-dvs: an event- stream dataset for object classification,

    H. Li, H. Liu, X. Ji, G. Li, and L. Shi, “Cifar10-dvs: an event- stream dataset for object classification,” Frontiers in Neuroscience, vol. 11, p. 309, 2017

  85. [94]

    Deformable detr: Deformable transformers for end-to-end object detection,

    X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in International Conference on Learning Representations, 2020

  86. [95]

    Spiking-yolo: Spiking neural network for energy-efficient object detection,

    S. Kim, S. Park, B. Na, and S. Yoon, “Spiking-yolo: Spiking neural network for energy-efficient object detection,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, pp. 11 270–11 277, 2020

  87. [96]

    Spike calibration: Fast and accurate conversion of spiking neural network for ob- ject detection and segmentation,

    Y. Li, X. He, Y. Dong, Q. Kong, and Y. Zeng, “Spike calibration: Fast and accurate conversion of spiking neural network for ob- ject detection and segmentation,” arXiv preprint arXiv:2207.02702, 2022

  88. [97]

    Direct training high-performance spiking neural networks for object recognition and detection,

    H. Zhang, Y. Li, B. He, X. Fan, Y. Wang, and Y. Zhang, “Direct training high-performance spiking neural networks for object recognition and detection,” Frontiers in Neuroscience , vol. 17, p. 1229951, 2023

  89. [98]

    Deep directly-trained spiking neural networks for object detection,

    Q. Su, Y. Chou, Y. Hu, J. Li, S. Mei, Z. Zhang, and G. Li, “Deep directly-trained spiking neural networks for object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 6555–6565

  90. [99]

    Internimage: Exploring large-scale vision foundation models with deformable convolutions,

    W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li et al. , “Internimage: Exploring large-scale vision foundation models with deformable convolutions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14...

  91. [100]

    Resnest: Split-attention networks,

    H. Zhang, C. Wu, Z. Zhang, Y. Zhu, H. Lin, Z. Zhang, Y. Sun, T. He, J. Mueller, R. Manmatha et al. , “Resnest: Split-attention networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2736–2746

  92. [101]

    Slowfast networks for video recognition,

    C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6202–6211

  93. [102]

    Action-net: Multipath excitation for action recognition,

    Z. Wang, Q. She, and A. Smolic, “Action-net: Multipath excitation for action recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 13 214– 13 223

  94. [103]

    Is space-time attention all you need for video understanding?

    G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in International Conference on Machine Learning. PMLR, 2021, p. 4

  95. [104]

    Inherent redundancy in spiking neural networks,

    M. Yao, J. Hu, G. Zhao, Y. Wang, Z. Zhang, B. Xu, and G. Li, “Inherent redundancy in spiking neural networks,” inProceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 16 924–16 934

  96. [105]

    The effective rank: A measure of effective dimensionality,

    O. Roy and M. Vetterli, “The effective rank: A measure of effective dimensionality,” in2007 15th European Signal Processing conference. IEEE, 2007, pp. 606–610

  97. [2024]

    Available: https://openreview.net/forum?id= Y8YVCOMEpz

    [Online]. Available: https://openreview.net/forum?id= Y8YVCOMEpz

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.