REVIEW 3 cited by
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Audio tagging is an important task of mapping audio samples to their corresponding categories. Recently endeavours that exploit transformer models in this field have achieved great success. However, the quadratic self-attention cost limits the scaling of audio transformer models and further constrains the development of more universal audio models. In this paper, we attempt to solve this problem by proposing Audio Mamba, a self-attention-free approach that captures long audio spectrogram dependency with state space models. Our experimental results on two audio-tagging datasets demonstrate the parameter efficiency of Audio Mamba, it achieves comparable results to SOTA audio spectrogram transformers with one third parameters.
Forward citations
Cited by 3 Pith papers
-
Recognizing Dementia from Neuropsychological Tests with State Space Models
Demenba applies Mamba/VMamba state space models to full-length neuropsychological interview audio and reports improved dementia classification AUC over an EfficientNet baseline, but the evaluation has significant sele...
-
Comparison of spectrogram scaling in multi-label Music Genre Recognition
On a custom 18k-song multi-label dataset, Mel-scaled spectrograms outperform standard spectrograms for music genre classification with transfer-learned ResNets.
-
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
A Transformer-Mamba model that adds a learned correction signal to degraded speech beats adapted active-noise-control baselines on denoising, dereverberation, and declipping in simulation.
Discussion (0). Continue with ORCID to comment.