Audio self-supervised learning with dual-softmax differential attention reports state-of-the-art numbers on AS-2M, AS20K, SPC-2, and ESC-50, but with tuning caveats and no code.
Masked autoencoders are scalable vision learners,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
Audio self-supervised learning with dual-softmax differential attention reports state-of-the-art numbers on AS-2M, AS20K, SPC-2, and ESC-50, but with tuning caveats and no code.