Applying L2 normalization to the final hidden features of audio and visual spiking networks, followed by spiking MLP fusion, yields 98.6% on CIFAR10-AV and 97.2% on UrbanSound8K-AV, outperforming a transformer-based SNN baseline.
Go- ing deeper with image transformers,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.NE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Spiking Neural Network Feature Discrimination Boosts Modality Fusion
Applying L2 normalization to the final hidden features of audio and visual spiking networks, followed by spiking MLP fusion, yields 98.6% on CIFAR10-AV and 97.2% on UrbanSound8K-AV, outperforming a transformer-based SNN baseline.