MDST++ reaches state-of-the-art audio-visual zero-shot learning on three benchmarks by converting RGB frames to synthetic events and processing them with a spiking transformer that shrinks timesteps.
Enhancing unsupervised video representation learning by decoupling the scene and the motion,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
MDST++ reaches state-of-the-art audio-visual zero-shot learning on three benchmarks by converting RGB frames to synthetic events and processing them with a spiking transformer that shrinks timesteps.