Pith. sign in

REVIEW 5 major objections 6 minor 53 references

Head-Tail-Aware KL Divergence in Knowledge Distillation for Spiking Neural Networks

T0 review · 5 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read HTA-KL, a knowledge-distillation loss for spiking neural networks, adaptively mixes forward and reverse KL divergence so the SNN student matches both the head and tail of the ANN teacher's output distribution, and beats prior distillation…

desk verdict A plausible KL reweighting for SNN distillation that overclaims its results — the idea is worth a revision, but not as written. read the letter →

arxiv 2504.20445 v2 pith:ERSDNC5G submitted 2025-04-29 cs.AI

classification cs.AI
keywords spikingneuralnetworksknowledgedistillationKLdivergenceforwardreversehead-tailweightingenergyefficiencylowtimestep
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Spiking neural networks promise energy-efficient inference but trail ANNs in accuracy. This paper argues that knowledge distillation from an ANN teacher to an SNN student fails because the usual KL loss overweights the teacher's high-probability predictions and ignores its low-probability tail. It proposes HTA-KL, which sorts the teacher's class probabilities, uses a cumulative-probability threshold to split them into head and tail regions, and adaptively weights a forward-KL term (head alignment) and a reverse-KL term (tail alignment) by the mismatch in each region. The paper reports that HTA-KL beats three prior SNN distillation methods on CIFAR-10, CIFAR-100, and Tiny ImageNet, often reaching comparable accuracy at fewer timesteps, and does so without changing the student architecture.

What carries the argument

The central object is the HTA-KL loss, defined by Eq. (16): $L_{HTA-KL} = \lambda_{head} L_{FKL} + \lambda_{tail} L_{RKL}$, with the weights computed from a cumulative-probability head/tail mask over the sorted teacher distribution (Eqs. 9\textendash15). The mask splits classes into head and tail at a cumulative threshold $\delta=0.5$, and the weights are the normalized head and tail absolute distances between teacher and student. This carries the argument because it converts the observation that forward KL aligns high-probability regions and reverse KL aligns low-probability regions into a concrete, adaptive loss for SNN training.

What would settle it

Train CIFAR-100 ResNet-19 SNN students with a grid of fixed head-tail ratios, including $\lambda=0.5$, and compare with HTA-KL's adaptive weights at timesteps 2, 4, and 6; if the best fixed ratio matches or beats the adaptive version, the adaptivity claim is falsified.

Watch

Extended reading notes

Core claim

HTA-KL recasts SNN knowledge distillation as the problem of aligning two regions of the teacher's output distribution at once. Given teacher and student softmax probabilities, the method sorts teacher probabilities in descending order, reorders the student's to match, and computes the absolute per-class distance $D_i$. A cumulative sum $S_i$ over the sorted teacher distribution, cut at $\delta=0.5$, marks the head (high-probability) and tail (low-probability) classes. The head and tail distances $d_{head}$ and $d_{tail}$ then set the weights $\lambda_{head} = d_{head}/(d_{head}+d_{tail})$ and $\lambda_{tail} = d_{tail}/(d_{head}+d_{tail})$, and the training loss is $\lambda_{head}$ times forward KL plus $\lambda_{tail}$ times reverse KL. The paper's central claim is that this dynamic weighting lets the student absorb both the teacher's confident predictions and its rare-class structure, closing the ANN-SNN accuracy gap more efficiently than fixed KL distillation, with lower spike firing rates and lower estimated energy consumption at short timesteps.

Load-bearing premise

The reported gains depend on the adaptive weighting itself being what helps; the paper never compares the adaptive head-tail weights with the best fixed ratio, so if a constant mixture of forward and reverse KL works just as well, the core claim weakens.

Editorial extensions

If this is right

  • On CIFAR-100, HTA-KL improves ResNet-19 accuracy over prior distillation baselines by 0.53, 0.39, and 0.85 percentage points at timesteps 2, 4, and 6.
  • On Tiny ImageNet, ResNet-20 with HTA-KL reaches 64.32 percent accuracy at timestep 2, matching methods that need timestep 4.
  • HTA-KL keeps spike firing rates moderate, for example 27.44 percent for ResNet-19 compared with 36.49 percent for LaSNN, while improving accuracy.
  • Estimated inference energy with HTA-KL is lower than KDSNN and LaSNN on ResNet-20, at 0.426734 mJ versus 0.440279 mJ and 0.463346 mJ.
  • The method adds no architectural change to the student, since the mask and weights come from the teacher and student output probabilities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the adaptive head-tail weighting should be testable against a fixed balanced mixture; the ablation in the paper only sweeps fixed ratios and never compares them with the adaptive weights, so a tuned constant might match HTA-KL.
  • Beyond the paper, the same cumulative-mask weighting could be applied to ANN-to-ANN distillation or to any student-teacher setup with long-tailed outputs, since the mask is computed from the teacher distribution alone.
  • Beyond the paper, the energy comparison counts only MAC and AC operations in 45 nm technology, so a fuller accounting that includes memory access and data movement would be needed to confirm real-device savings.
  • Beyond the paper, because the mask depends only on teacher probabilities, HTA-KL could be combined with other SNN training tricks (surrogate gradients, membrane normalization) without recomputing the head-tail split.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper proposes Head-Tail-Aware KL divergence (HTA-KL) for knowledge distillation from ANN teachers to SNN students. The method sorts the teacher's class probabilities, separates head and tail regions using a cumulative-probability threshold, and adaptively weights forward KL and reverse KL terms based on the teacher--student distance in each region. Experiments on CIFAR-10, CIFAR-100, and Tiny ImageNet claim that HTA-KL outperforms existing SNN distillation methods (KDSNN, LaSNN, BKDSNN) across architectures and timesteps, including at low timesteps, with improved energy efficiency. The paper also reports ablation, spike firing rate, and t-SNE analyses.

Significance. If the claimed results were supported, the paper would make a modest but useful contribution: it transfers the head/tail insight from language-model knowledge distillation to SNNs and proposes a concrete adaptive weighting scheme. The paper also provides a broad experimental comparison, an energy-consumption analysis, and firing-rate statistics, which are valuable in principle. However, the central empirical claims are internally inconsistent with the paper's own tables, and the proposed adaptive mechanism is not isolated in the ablation. As written, the contribution cannot be assessed as a reliable advance, so the significance is currently low.

major comments (5)
  1. [Section IV-B, Table I] The claim that HTA-KL 'consistently outperforms these methods across all architectures and timesteps' is contradicted by Table I. Specific counterexamples include CIFAR-10 ResNet-19 at T=1 (LaSNN 96.19% vs. HTA-KL 96.11%), CIFAR-10 ResNet-20 at T=4 (KDSNN 94.07% vs. HTA-KL 94.06%), CIFAR-10 VGG-16 at T=2 (BKDSNN 94.61% vs. HTA-KL 94.44%), and CIFAR-100 ResNet-19 at T=1 (BKDSNN 78.77% vs. HTA-KL 78.75%). The central superiority claim is therefore unsupported by the paper's own results.
  2. [Section IV-B, Table II] The claim that HTA-KL achieves strong performance with fewer timesteps is not established on Tiny ImageNet because architecture and timestep are confounded: HTA-KL is evaluated on ResNet-20 and VGG-16 at T=2, whereas the baselines are SEW ResNet-18/34 at T=4. Moreover, the ANN teacher at T=1 already achieves 65.72% (ResNet-20) and 65.94% (VGG-16), both above HTA-KL's T=2 results of 64.32% and 64.10%. This comparison does not demonstrate a timestep advantage over the baselines or over the teacher.
  3. [Section IV-E, Table III] The energy-efficiency claim that HTA-KL 'outperforming KDSNN and BKDSNN in energy efficiency' is contradicted by Table III. On ResNet-19, BKDSNN consumes 1.78773 mJ versus HTA-KL's 2.02173 mJ, and on VGG-16, KDSNN consumes 1.438471 mJ versus HTA-KL's 1.440955 mJ. The stated favorable trade-off is therefore not consistently supported by the data.
  4. [Section IV-C, Eq. (15)] The adaptive weighting mechanism is not validated as the source of the reported gains. Section IV-C sweeps a fixed head-tail ratio and reports that accuracy is best at a balanced ratio, but it never compares the adaptive weights from Eq. (15) against the best fixed ratio. Without that comparison, the paper's core novelty—adaptive region weighting—is untested, and the improvements could be achieved by a fixed balanced combination of FKL and RKL.
  5. [Section IV-A, Eqs. (2), (5), (13)] The values of the hyperparameters alpha in Eq. (5), temperature tau in Eq. (2), and threshold delta in Eq. (13) are not reported in the implementation details, and the ablations do not cover sensitivity to tau or delta. This makes the experiments non-reproducible and leaves open the possibility that the reported results depend on fine-tuned parameter choices.
minor comments (6)
  1. [Section IV-B] The sentence 'as shown in section IV' should refer to Table I explicitly; the current pointer is unhelpful because Section IV contains many tables and figures.
  2. [Table I and Fig. 3] The spelling of the baseline method is inconsistent: 'LaSNN' in the text and references but 'LASNN' in Table I and Fig. 3. Please use a single convention consistently.
  3. [Table II] The entry 'Spkfmer-6-512' appears to be a typo for 'Spikformer-6-512'.
  4. [Section III-B, Eqs. (11) and (15)] The definitions of d_head and d_tail are given only in prose; please write them with explicit summation ranges and state clearly that M_head and M_tail depend on the cumulative sum of the sorted teacher distribution.
  5. [Abstract and Section IV-B] The abstract says HTA-KL 'outperforms existing methods on most datasets', while Section IV-B claims it 'consistently outperforms these methods across all architectures and timesteps'. These statements are in tension, and both should be aligned with the actual table entries.
  6. [Fig. 2 caption] The caption says the figure shows accuracy for 'different ratios between the head and tail losses', but the axis labels and legend are not clear about what ratio is varied and for which model and dataset. Please label the axes and specify the experimental setting.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: HTA-KL is a constructive loss definition with no derivation step that reduces to its own inputs.

full rationale

I examined the derivation chain of HTA-KL. The method is defined by Equations (8)-(16): the teacher probabilities are sorted, a cumulative mask splits classes into head and tail using a threshold delta, per-class absolute distances between aligned teacher and student probabilities are computed, region weights lambda_head and lambda_tail are set as normalized distances, and the loss is defined as lambda_head * L_FKL + lambda_tail * L_RKL. This is a constructive loss definition, not a derivation of a predicted quantity from an input. The adaptive weights depend on the teacher and current student distributions, so they are neither fitted parameters nor predictions of an external result. The head/tail interpretation of FKL and RKL is explicitly attributed to prior work [15], and the paper's contribution is the SNN-specific combination with a cumulative-probability mask; no load-bearing self-citation or imported uniqueness theorem is invoked. The paper's experimental claims are problematic in other ways, as Table I contains counterexamples to the 'consistently outperforms' statement and the Tiny ImageNet comparison confounds architecture with timestep, but those are empirical-support and experimental-design issues, not circularity. No step in the paper reduces to its own inputs by construction, so the circularity score is 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The method rests on the cited head/tail behavior of FKL vs RKL, the hand-set threshold delta=0.5, and unstated temperature tau and alpha values. No new entities are introduced. The adaptive weights are computed from the current student-teacher distance, which is a training dynamic rather than a fitted constant, but the threshold delta is a hand-set hyperparameter.

free parameters (3)
  • delta = 0.5 (default)
    Eq. (13): classes are divided into head and tail at cumulative probability 0.5. No sensitivity analysis or justification for this threshold is provided.
  • alpha = not reported
    Eq. (5): L_SKD = (1-alpha) L_CE + alpha L_KL. The balance between cross-entropy and KD loss is not specified in the implementation details.
  • temperature tau = not reported
    Eqs. (2) and (8): the softmax temperature controls the peakiness of teacher and student distributions, directly affecting the head/tail split. No value is given.
assumptions (3)
  • domain assumption Forward KL aligns teacher head (high-probability) regions; reverse KL aligns tail (low-probability) regions.
    Taken from the cited AKL paper (Wu et al., 2025) and used to justify combining FKL and RKL in Eqs. (15)-(16).
  • domain assumption The cumulative probability threshold delta=0.5 correctly separates head from tail classes.
    Eq. (13) defines M_head by S_i < 0.5; no theoretical or empirical justification for this specific threshold is provided.
  • domain assumption Temporal averaging of student softmax over timesteps is a faithful distillation target.
    Eq. (3) uses Q_SNN_avg = (1/T) sum Q_SNN(t), a standard but unstated assumption in SNN knowledge distillation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Head-Tail-Aware KL Divergence in Knowledge Distillation for Spiking Neural Networks." pith.science (2026). https://pith.science/paper/ERSDNC5G

@misc{pith2026250420445,
  author       = {Pith},
  title        = {Pith review of: Head-Tail-Aware KL Divergence in Knowledge Distillation for Spiking Neural Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ERSDNC5G}},
  note         = {Machine review of arXiv:2504.20445}
}
read the original abstract

Spiking Neural Networks (SNNs) have emerged as a promising approach for energy-efficient and biologically plausible computation. However, due to limitations in existing training methods and inherent model constraints, SNNs often exhibit a performance gap when compared to Artificial Neural Networks (ANNs). Knowledge distillation (KD) has been explored as a technique to transfer knowledge from ANN teacher models to SNN student models to mitigate this gap. Traditional KD methods typically use Kullback-Leibler (KL) divergence to align output distributions. However, conventional KL-based approaches fail to fully exploit the unique characteristics of SNNs, as they tend to overemphasize high-probability predictions while neglecting low-probability ones, leading to suboptimal generalization. To address this, we propose Head-Tail Aware Kullback-Leibler (HTA-KL) divergence, a novel KD method for SNNs. HTA-KL introduces a cumulative probability-based mask to dynamically distinguish between high- and low-probability regions. It assigns adaptive weights to ensure balanced knowledge transfer, enhancing the overall performance. By integrating forward KL (FKL) and reverse KL (RKL) divergence, our method effectively align both head and tail regions of the distribution. We evaluate our methods on CIFAR-10, CIFAR-100 and Tiny ImageNet datasets. Our method outperforms existing methods on most datasets with fewer timesteps.

Figures

Figures reproduced from arXiv: 2504.20445 by the authors.

Figure 1
Figure 1. The framework of the proposed HTA-KL for SNN training. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Impact of the Head-Tail Ratio on Accuracy. The graph shows the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Spike firing rate comparison for different distillation methods [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: t-SNE Visualization of features learned by teacher ANN and different distillation methods. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 43 canonical work pages

  1. [1]

    Networks of spiking neurons: The third generation of neural network models,

    W. Maass, “Networks of spiking neurons: The third generation of neural network models,” Neural Networks , vol. 10, no. 9, pp. 1659–1671, Dec. 1997. [Online]. Available: https://www.sciencedirect.com/science/ article/pii/S0893608097000117

  2. [2]

    Spiking Neural Networks and Their Applications: A Review,

    K. Yamazaki, V .-K. V o-Ho, D. Bulsara, and N. Le, “Spiking Neural Networks and Their Applications: A Review,” Brain Sciences, vol. 12, no. 7, p. 863, Jul. 2022, number: 7 Publisher: Multidisciplinary Digital Publishing Institute. [Online]. Available: https://www.mdpi.com/ 2076-3425/12/7/863

  3. [3]

    A review of learning in biologically plausible spiking neural networks,

    A. Taherkhani, A. Belatreche, Y . Li, G. Cosma, L. P. Maguire, and T. M. McGinnity, “A review of learning in biologically plausible spiking neural networks,” Neural Networks , vol. 122, pp. 253–272, 2020. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0893608019303181

  4. [4]

    Relative Pose Estimation for Multi-Camera Systems from Point Correspondences with Scale Ratio,

    B. Guan and J. Zhao, “Relative Pose Estimation for Multi-Camera Systems from Point Correspondences with Scale Ratio,” in Proceedings of the 30th ACM International Conference on Multimedia . Lisboa Portugal: ACM, Oct. 2022, pp. 5036–5044. [Online]. Available: https://dl.acm.org/doi/10.1145/3503161.3547788

  5. [5]

    Optical flow- guided 6dof object pose tracking with an event camera,

    Z. Liu, B. Guan, Y . Shang, S. Liang, Z. Yu, and Q. Yu, “Optical flow- guided 6dof object pose tracking with an event camera,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 6501–6509

  6. [6]

    Line-based 6- dof object pose estimation and tracking with an event camera,

    Z. Liu, B. Guan, Y . Shang, Q. Yu, and L. Kneip, “Line-based 6- dof object pose estimation and tracking with an event camera,” IEEE Transactions on Image Processing , 2024

  7. [7]

    Stereo event-based, 6-dof pose tracking for uncooperative spacecraft,

    Z. Liu, B. Guan, Y . Shang, Y . Bian, P. Sun, and Q. Yu, “Stereo event-based, 6-dof pose tracking for uncooperative spacecraft,” IEEE Transactions on Geoscience and Remote Sensing , 2025

  8. [8]

    Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks,

    Y . Wu, L. Deng, G. Li, J. Zhu, and L. Shi, “Spatio-Temporal Backpropagation for Training High-Performance Spiking Neural Networks,” Frontiers in Neuroscience , vol. 12, p. 323875, 2018. [Online]. Available: https://www.frontiersin.org/articles/10.3389/fnins. 2018.00331

Show all 53 references
  1. [9]

    Surrogate Gradient Learning in Spiking Neural Networks: Bringing the Power of Gradient-Based Optimization to Spiking Neural Networks,

    E. O. Neftci, H. Mostafa, and F. Zenke, “Surrogate Gradient Learning in Spiking Neural Networks: Bringing the Power of Gradient-Based Optimization to Spiking Neural Networks,” IEEE Signal Processing Magazine , vol. 36, no. 6, pp. 51–63, Nov. 2019, conference Name: IEEE Signal ...

  2. [10]

    Knowledge Distillation from A Stronger Teacher,

    T. Huang, S. You, F. Wang, C. Qian, and C. Xu, “Knowledge Distillation from A Stronger Teacher,” Advances in Neural Information Processing Systems , vol. 35, pp. 33 716–33 727, Dec. 2022. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2022/hash/ da669dfd...

  3. [11]

    Decoupled Knowledge Distillation,

    B. Zhao, Q. Cui, R. Song, Y . Qiu, and J. Liang, “Decoupled Knowledge Distillation,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . New Orleans, LA, USA: IEEE, Jun. 2022, pp. 11 943–11 952. [Online]. Available: https: //ieeexplore.ieee.org/docu...

  4. [12]

    Constructing Deep Spiking Neural Networks From Artificial Neural Networks With Knowledge Distillation,

    Q. Xu, Y . Li, J. Shen, J. K. Liu, H. Tang, and G. Pan, “Constructing Deep Spiking Neural Networks From Artificial Neural Networks With Knowledge Distillation,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2023, pp. 7886–7895. [Online]. ...

  5. [13]

    LaSNN: Layer-wise ANN-to-SNN Distillation for Effective and Efficient Training in Deep Spiking Neural Networks,

    D. Hong, J. Shen, Y . Qi, and Y . Wang, “LaSNN: Layer-wise ANN-to-SNN Distillation for Effective and Efficient Training in Deep Spiking Neural Networks,” Apr. 2023, arXiv:2304.09101 [cs]. [Online]. Available: http://arxiv.org/abs/2304.09101

  6. [14]

    BKDSNN: Enhancing the Performance of Learning-Based Spiking Neural Networks Training with Blurred Knowledge Distillation,

    Z. Xu, K. You, Q. Guo, X. Wang, and Z. He, “BKDSNN: Enhancing the Performance of Learning-Based Spiking Neural Networks Training with Blurred Knowledge Distillation,” inComputer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds....

  7. [15]

    Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models,

    T. Wu, C. Tao, J. Wang, R. Yang, Z. Zhao, and N. Wong, “Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models,” in Proceedings of the 31st International Conference on Computational Linguistics , O. Rambow, L. Wanner, M. Apidianaki, H. Al-Kh...

  8. [16]

    A New ANN-SNN Conversion Method with High Accuracy, Low Latency and Good Robustness,

    B. Wang, J. Cao, J. Chen, S. Feng, and Y . Wang, “A New ANN-SNN Conversion Method with High Accuracy, Low Latency and Good Robustness,” in Proceedings of IJCAI , Macau, SAR China, 2023, pp. 3067–3075. [Online]. Available: https://www.ijcai.org/proceedings/ 2023/342

  9. [17]

    Advancing Spiking Neural Networks Toward Deep Residual Learning,

    Y . Hu, L. Deng, Y . Wu, M. Yao, and G. Li, “Advancing Spiking Neural Networks Toward Deep Residual Learning,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–15, 2024, conference Name: IEEE Transactions on Neural Networks and Learning Systems. [Online]. Avai...

  10. [18]

    DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks,

    T. Zhang, K. Yu, J. Zhang, and H. Wang, “DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Apr. 2025, pp. 1–5, iSSN: 2379-190X. [Onli...

  11. [19]

    Going Deeper With Directly-Trained Larger Spiking Neural Networks,

    H. Zheng, Y . Wu, L. Deng, Y . Hu, and G. Li, “Going Deeper With Directly-Trained Larger Spiking Neural Networks,” Proceedings of AAAI, vol. 35, pp. 11 062–11 070, May 2021, number: 12. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/17320

  12. [20]

    Membrane Potential Batch Normalization for Spiking Neural Networks,

    Y . Guo, Y . Zhang, Y . Chen, W. Peng, X. Liu, L. Zhang, X. Huang, and Z. Ma, “Membrane Potential Batch Normalization for Spiking Neural Networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct. 2023, pp. 19 420–19 430

  13. [21]

    FSTA-SNN:Frequency-Based Spatial-Temporal Attention Module for Spiking Neural Networks,

    K. Yu, T. Zhang, H. Wang, and Q. Xu, “FSTA-SNN:Frequency-Based Spatial-Temporal Attention Module for Spiking Neural Networks,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 21, pp. 22 227–22 235, Apr. 2025, number: 21. [Online]. Available: https:...

  14. [22]

    ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural Networks,

    J. Shen, Q. Xu, J. K. Liu, Y . Wang, G. Pan, and H. Tang, “ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural Networks,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 1, pp. 86–93, Jun. 2023, number: 1. [Online]. Available: h...

  15. [23]

    Deep Residual Learning in Spiking Neural Networks,

    W. Fang, Z. Yu, Y . Chen, T. Huang, T. Masquelier, and Y . Tian, “Deep Residual Learning in Spiking Neural Networks,” in Advances in Neural Information Processing Systems, vol. 34. Curran Associates, Inc., 2021, pp. 21 056–21 069. [Online]. Available: https://proceedings.neuri...

  16. [24]

    Multi-Level Firing with Spiking DS-ResNet: Enabling Better and Deeper Directly- Trained Spiking Neural Networks,

    L. Feng, Q. Liu, H. Tang, D. Ma, and G. Pan, “Multi-Level Firing with Spiking DS-ResNet: Enabling Better and Deeper Directly- Trained Spiking Neural Networks,” in Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence . Vienna, Austria: Inter...

  17. [25]

    RSNN: Recurrent Spiking Neural Networks for Dynamic Spatial-Temporal Information Processing,

    Q. Xu, X. Fang, Y . Li, J. Shen, D. Ma, Y . Xu, and G. Pan, “RSNN: Recurrent Spiking Neural Networks for Dynamic Spatial-Temporal Information Processing,” in Proceedings of the 32nd ACM International Conference on Multimedia , ser. MM ’24. New York, NY , USA: Association for C...

  18. [26]

    SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks,

    X. Shi, Z. Hao, and Z. Yu, “SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural Networks,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 5610–5619. [Online]. Available: https://openaccess.thecvf.com/content/CVP...

  19. [27]

    Spikformer: When Spiking Neural Network Meets Transformer,

    Z. Zhou, Y . Zhu, C. He, Y . Wang, S. Yan, Y . Tian, and L. Yuan, “Spikformer: When Spiking Neural Network Meets Transformer,” in ICLR, Sep. 2022. [Online]. Available: https: //openreview.net/forum?id=frE4fUwz h

  20. [28]

    Temporal spiking generative adversarial networks for heading direction decoding,

    J. Shen, K. Wang, W. Gao, J. K. Liu, Q. Xu, G. Pan, X. Chen, and H. Tang, “Temporal spiking generative adversarial networks for heading direction decoding,” Neural Networks , vol. 184, p. 106975, Apr. 2025. [Online]. Available: https://linkinghub.elsevier.com/retrieve/ pii/S08...

  21. [29]

    Spiking Token Mixer: An event-driven friendly Former structure for spiking neural networks,

    S. Deng, Y . Wu, K. Du, and S. Gu, “Spiking Token Mixer: An event-driven friendly Former structure for spiking neural networks,” Advances in Neural Information Processing Systems, vol. 37, pp. 128 825–128 846, Dec. 2024. [Online]. Available: https://proceedings.neurips.cc/pape...

  22. [30]

    Revisiting Knowledge Distillation via Label Smoothing Regularization,

    L. Yuan, F. E. Tay, G. Li, T. Wang, and J. Feng, “Revisiting Knowledge Distillation via Label Smoothing Regularization,” 2020, pp. 3903–3911. [Online]. Available: https://openaccess.thecvf.com/ content CVPR 2020/html/Yuan Revisiting Knowledge Distillation via Label Smoothing R...

  23. [31]

    Distilling Knowledge via Knowledge Review,

    P. Chen, S. Liu, H. Zhao, and J. Jia, “Distilling Knowledge via Knowledge Review,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Nashville, TN, USA: IEEE, Jun. 2021, pp. 5006–5015. [Online]. Available: https://ieeexplore.ieee. org/document/9578915/

  24. [32]

    Logit Standardization in Knowledge Distillation,

    S. Sun, W. Ren, J. Li, R. Wang, and X. Cao, “Logit Standardization in Knowledge Distillation,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 15 731–15 740. [Online]. Available: https://openaccess.thecvf.com/content/CVPR2024/html/Sun L...

  25. [33]

    Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks,

    K. Yu, C. Yu, T. Zhang, X. Zhao, S. Yang, H. Wang, Q. Zhang, and Q. Xu, “Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks,” Mar. 2025, arXiv:2503.03144 [cs]. [Online]. Available: http://arxiv.org/abs/2503. 03144

  26. [34]

    Distilling Spikes: Knowledge Distillation in Spiking Neural Networks,

    R. K. Kushawaha, S. Kumar, B. Banerjee, and R. Velmurugan, “Distilling Spikes: Knowledge Distillation in Spiking Neural Networks,” in 2020 25th International Conference on Pattern Recognition (ICPR) . Milan, Italy: IEEE, Jan. 2021, pp. 4536–4543. [Online]. Available: https://i...

  27. [35]

    Spike-Thrift: Towards Energy-Efficient Deep Spiking Neural Networks by Limiting Spiking Activity via Attention-Guided Compression,

    S. Kundu, G. Datta, M. Pedram, and P. A. Beerel, “Spike-Thrift: Towards Energy-Efficient Deep Spiking Neural Networks by Limiting Spiking Activity via Attention-Guided Compression,” 2021, pp. 3953–3962. [Online]. Available: https://openaccess.thecvf.com/content/W ACV2021/ html...

  28. [36]

    SpikeBERT: A Language Spikformer Trained with Two-Stage Knowledge Distillation from BERT,

    C. Lv, T. Li, J. Xu, C. Gu, Z. Ling, C. Zhang, X. Zheng, and X. Huang, “SpikeBERT: A Language Spikformer Trained with Two-Stage Knowledge Distillation from BERT,” Aug. 2023, arXiv:2308.15122 [cs]. [Online]. Available: http://arxiv.org/abs/2308.15122

  29. [37]

    Joint A-SNN: Joint training of artificial and spiking neural networks via self-Distillation and weight factorization,

    Y . Guo, W. Peng, Y . Chen, L. Zhang, X. Liu, X. Huang, and Z. Ma, “Joint A-SNN: Joint training of artificial and spiking neural networks via self-Distillation and weight factorization,” Pattern Recognition, vol. 142, p. 109639, Oct. 2023. [Online]. Available: https://www.scie...

  30. [38]

    Reversing Structural Pattern Learning with Biologically Inspired Knowledge Distillation for Spiking Neural Networks,

    Q. Xu, Y . Li, X. Fang, J. Shen, Q. Zhang, and G. Pan, “Reversing Structural Pattern Learning with Biologically Inspired Knowledge Distillation for Spiking Neural Networks,” in 2024 ACM MM , Jul

  31. [39]

    Self-architectural knowledge distillation for spiking neural networks,

    H. Qiu, M. Ning, Z. Song, W. Fang, Y . Chen, T. Sun, Z. Ma, L. Yuan, and Y . Tian, “Self-architectural knowledge distillation for spiking neural networks,” Neural Networks , vol. 178, p. 106475, Oct. 2024. [Online]. Available: https://www.sciencedirect.com/science/ article/pii...

  32. [40]

    RecDis-SNN: Rectifying Membrane Potential Distribution for Directly Training Spiking Neural Networks,

    Y . Guo, X. Tong, Y . Chen, L. Zhang, X. Liu, Z. Ma, and X. Huang, “RecDis-SNN: Rectifying Membrane Potential Distribution for Directly Training Spiking Neural Networks,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . New Orleans, LA, USA: IEEE...

  33. [41]

    Temporal Efficient Training of Spiking Neural Network via Gradient Re-weighting,

    S. Deng, Y . Li, S. Zhang, and S. Gu, “Temporal Efficient Training of Spiking Neural Network via Gradient Re-weighting,” in Proceedings of ICLR , Oct. 2021. [Online]. Available: https: //openreview.net/forum?id= XNtisL32jv

  34. [42]

    Learnable Surrogate Gradient for Direct Training Spiking Neural Networks,

    S. Lian, J. Shen, Q. Liu, Z. Wang, R. Yan, and H. Tang, “Learnable Surrogate Gradient for Direct Training Spiking Neural Networks,” in Proceedings of IJCAI , Macau, SAR China, Aug. 2023, pp. 3002–3010. [Online]. Available: https://www.ijcai.org/proceedings/2023/335

  35. [43]

    Tensor Decomposition Based Attention Module for Spiking Neural Networks,

    H. Deng, R. Zhu, X. Qiu, Y . Duan, M. Zhang, and L. Deng, “Tensor Decomposition Based Attention Module for Spiking Neural Networks,” Oct. 2023, arXiv:2310.14576 [cs]. [Online]. Available: http://arxiv.org/abs/2310.14576

  36. [44]

    IM- Loss: Information Maximization Loss for Spiking Neural Networks,

    Y . Guo, Y . Chen, L. Zhang, X. Liu, Y . Wang, X. Huang, and Z. Ma, “IM- Loss: Information Maximization Loss for Spiking Neural Networks,” Proceedings of NeurIPS , vol. 35, pp. 156–166, Dec. 2022. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2022/hash/...

  37. [45]

    IM-LIF: Improved Neuronal Dynamics With Attention Mechanism for Direct Training Deep Spiking Neural Network,

    S. Lian, J. Shen, Z. Wang, and H. Tang, “IM-LIF: Improved Neuronal Dynamics With Attention Mechanism for Direct Training Deep Spiking Neural Network,” IEEE Transactions on Emerging Topics in Computational Intelligence , pp. 1–11, 2024, conference Name: IEEE Transactions on Eme...

  38. [46]

    Spikingformer: Spike-driven Residual Learning for Transformer-based Spiking Neural Network,

    C. Zhou, L. Yu, Z. Zhou, Z. Ma, H. Zhang, H. Zhou, and Y . Tian, “Spikingformer: Spike-driven Residual Learning for Transformer-based Spiking Neural Network,” May 2023, arXiv:2304.11954. [Online]. Available: http://arxiv.org/abs/2304.11954

  39. [47]

    Cifar-10 (canadian institute for advanced research),

    A. Krizhevsky, V . Nair, and G. Hinton, “Cifar-10 (canadian institute for advanced research),” URL http://www. cs. toronto. edu/kriz/cifar. html , vol. 5, no. 4, p. 1, 2010

  40. [48]

    Tiny ImageNet Visual Recognition Challenge,

    Y . Le and X. Yang, “Tiny ImageNet Visual Recognition Challenge,” 2015

  41. [49]

    GLIF: A Unified Gated Leaky Integrate-and-Fire Neuron for Spiking Neural Networks,

    X. Yao, F. Li, Z. Mo, and J. Cheng, “GLIF: A Unified Gated Leaky Integrate-and-Fire Neuron for Spiking Neural Networks,” Proceedings of NeurIPS , vol. 35, pp. 32 160–32 171, Dec. 2022. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2022/hash/ cfa8440d500...

  42. [50]

    1.1 Computing’s energy problem (and what we can do about it),

    M. Horowitz, “1.1 Computing’s energy problem (and what we can do about it),” in 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC), Feb. 2014, pp. 10–14, iSSN: 2376-

  43. [51]

    A reconfigurable on-line learning spiking neuromorphic processor comprising 256 neurons and 128K synapses,

    N. Qiao, H. Mostafa, F. Corradi, M. Osswald, F. Stefanini, D. Sumislawska, and G. Indiveri, “A reconfigurable on-line learning spiking neuromorphic processor comprising 256 neurons and 128K synapses,” Frontiers in Neuroscience , vol. 9, Apr. 2015, publisher: Frontiers. [Online...

  44. [2024]

    Available: https://openreview.net/forum?id=r9X3P6qvAj

    [Online]. Available: https://openreview.net/forum?id=r9X3P6qvAj

  45. [8606]

    Available: https://ieeexplore.ieee.org/abstract/document/ 6757323

    [Online]. Available: https://ieeexplore.ieee.org/abstract/document/ 6757323

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.