Pith. sign in

REVIEW 4 cited by

XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.10027 v2 pith:WWHWTVY7 submitted 2024-11-15 eess.AS cs.SD

classification eess.AScs.SD
keywords mambaattackdetectionmodelspeechspoofingbeendual-column
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers and their variants have achieved great success in speech processing. However, their multi-head self-attention mechanism is computationally expensive. Therefore, one novel selective state space model, Mamba, has been proposed as an alternative. Building on its success in automatic speech recognition, we apply Mamba for spoofing attack detection. Mamba is well-suited for this task as it can capture the artifacts in spoofed speech signals by handling long-length sequences. However, Mamba's performance may suffer when it is trained with limited labeled data. To mitigate this, we propose combining a new structure of Mamba based on a dual-column architecture with self-supervised learning, using the pre-trained wav2vec 2.0 model. The experiments show that our proposed approach achieves competitive results and faster inference on the ASVspoof 2021 LA and DF datasets, and on the more challenging In-the-Wild dataset, it emerges as the strongest candidate for spoofing attack detection. The code has been publicly released in https://github.com/swagshaw/XLSR-Mamba.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tell me Habibi, is it Real or Fake?

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.

  2. WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes

    cs.CR 2025-07 conditional novelty 6.0 of 10

    WaveVerify embeds audio watermarks with a FiLM-based generator and extracts them with a Mixture-of-Experts detector, reporting zero bit error and high localization under common distortions.

  3. Evaluating Fake Music Detection Performance Under Audio Augmentations

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SONICS, a recent fake-music detector, suffers large accuracy drops under light audio augmentations and fails to generalize to unseen generative models.

  4. Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion

    cs.SD 2025-06 conditional novelty 4.0 of 10

    An XLSR-Conformer system trained with Real Emphasis, Fake Dispersion, and multi-class N-pair loss reaches 95.6% in-domain and up to 44.8% out-of-domain source tracing accuracy on MLAAD, versus 83.4% and 26.5% for the ...

Pith tools