REVIEW 4 cited by
XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformers and their variants have achieved great success in speech processing. However, their multi-head self-attention mechanism is computationally expensive. Therefore, one novel selective state space model, Mamba, has been proposed as an alternative. Building on its success in automatic speech recognition, we apply Mamba for spoofing attack detection. Mamba is well-suited for this task as it can capture the artifacts in spoofed speech signals by handling long-length sequences. However, Mamba's performance may suffer when it is trained with limited labeled data. To mitigate this, we propose combining a new structure of Mamba based on a dual-column architecture with self-supervised learning, using the pre-trained wav2vec 2.0 model. The experiments show that our proposed approach achieves competitive results and faster inference on the ASVspoof 2021 LA and DF datasets, and on the more challenging In-the-Wild dataset, it emerges as the strongest candidate for spoofing attack detection. The code has been publicly released in https://github.com/swagshaw/XLSR-Mamba.
Forward citations
Cited by 4 Pith papers
-
Tell me Habibi, is it Real or Fake?
ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.
-
WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes
WaveVerify embeds audio watermarks with a FiLM-based generator and extracts them with a Mixture-of-Experts detector, reporting zero bit error and high localization under common distortions.
-
Evaluating Fake Music Detection Performance Under Audio Augmentations
SONICS, a recent fake-music detector, suffers large accuracy drops under light audio augmentations and fails to generalize to unseen generative models.
-
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
An XLSR-Conformer system trained with Real Emphasis, Fake Dispersion, and multi-class N-pair loss reaches 95.6% in-domain and up to 44.8% out-of-domain source tracing accuracy on MLAAD, versus 83.4% and 26.5% for the ...
Discussion (0). Sign in to comment.