Pith. sign in

REVIEW 4 major objections 1 minor 68 references

Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative

T0 review · 4 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Fake-Mamba claims bidirectional Mamba with an XLSR front-end detects synthetic speech at 0.97% EER on ASVspoof 21 LA, 1.74% on 21 DF, and 5.85% on In-The-Wild, outperforming XLSR-Conformer and XLSR-Mamba while preserving real-time inference

desk verdict The submitted full text is an unrelated medical-imaging paper, so the abstract's EER claims for Fake-Mamba are unverifiable and the manuscript cannot be seriously reviewed as submitted. read the letter →

arxiv 2508.09294 v1 pith:7GWET7AC submitted 2025-08-12 eess.AS cs.AIcs.CLcs.LGcs.SYeess.SY

classification eess.AScs.AIcs.CLcs.LGcs.SYeess.SY
keywords speechdeepfakedetectionbidirectionalMambaXLSRfront-endself-attentionalternativeASVspoof2021LADFIn-The-Wildreal-timeinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that bidirectional Mamba, a state-space sequence model, can replace self-attention in speech deepfake detection. It reports that its Fake-Mamba system, pairing an XLSR front-end with three new Mamba encoders, achieves equal error rates of 0.97%, 1.74%, and 5.85% on ASVspoof 21 LA, ASVspoof 21 DF, and In-The-Wild, respectively, beating XLSR-Conformer and XLSR-Mamba. The authors claim this holds while keeping inference real time across utterance lengths. A sympathetic reader would care because it suggests a cheaper, faster alternative to attention-based detectors without losing accuracy.

What carries the argument

The central mechanism is the bidirectional Mamba encoder, a state-space sequence model that processes speech forward and backward, used as a drop-in alternative to self-attention. The paper proposes three variants; the named best performer is PN-BiMamba. Paired with the XLSR front-end, the encoder extracts both local and global artifacts that distinguish synthetic from natural speech.

What would settle it

Retrain XLSR-Conformer and XLSR-Mamba under identical training folds, optimizer settings, and evaluation scripts on ASVspoof 21 LA, then compare equal error rates with Fake-Mamba's reported 0.97%. If the gap closes or reverses, the central superiority claim fails.

Watch

Extended reading notes

Core claim

On the paper's own terms, bidirectional Mamba is competitive with or superior to self-attention for detecting synthetic speech. The system, Fake-Mamba, uses an XLSR pre-trained front-end for linguistic representations, then one of three proposed Mamba encoders—TransBiMamba, ConBiMamba, or PN-BiMamba—to model local and global artifacts. PN-BiMamba is the best at capturing subtle cues of synthetic speech. Evaluated on standard benchmarks, it reports state-of-the-art equal error rates and real-time inference across utterance lengths.

Load-bearing premise

The reported equal error rates beat the baselines only if XLSR-Conformer and XLSR-Mamba were retrained under the exact same data splits, evaluation protocol, and hyperparameter budget; the abstract does not spell out those controls.

Editorial extensions

If this is right

  • If the EER numbers hold under matched conditions, Mamba-based encoders become a strong candidate for real-time synthetic speech detection on long utterances.
  • The success of bidirectional Mamba suggests attention is not necessary for capturing both local and global artifacts in speech forensics; other audio classification tasks may adopt similar encoders.
  • The three encoder variants form a lightweight design space, and the best one, PN-BiMamba, could be plugged into other front-ends for forensic or anti-spoofing systems.
  • Real-time inference across utterance lengths implies deployability in streaming or interactive settings where latency constraints previously favored lighter architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported gains are relative to specific baselines; if XLSR-Conformer and XLSR-Mamba were not retrained with identical data splits, evaluation protocol, and compute budgets, the margins may shrink or reverse.
  • The architecture may generalize beyond the three tested benchmarks to other spoofing types or unseen generators, but that is not demonstrated here.
  • The XLSR front-end likely dominates parameter count and inference cost; the real-time claim may hold only for the encoder component, not the full pipeline.
  • If the Mamba encoder works on raw features rather than fine-tuned front-end outputs, it could transfer to other audio forensics tasks such as voice conversion or replay detection.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 1 minor

Summary. The submission, as identified by its abstract (arXiv:2508.09294), claims to present Fake-Mamba, a speech deepfake detector combining an XLSR front-end with bidirectional Mamba encoders (TransBiMamba, ConBiMamba, PN-BiMamba). The abstract reports EERs of 0.97%, 1.74%, and 5.85% on ASVspoof 21 LA, ASVspoof 21 DF, and In-The-Wild, respectively, with 'substantial relative gains' over XLSR-Conformer and XLSR-Mamba, and real-time inference. However, the full text supplied is an entirely different manuscript: 'Ethical Medical Image Synthesis' by Jin et al. (arXiv:2508.09293). The body contains no description of Fake-Mamba, no architectural details, no experimental setup, no EER tables, no baseline comparisons, and no runtime measurements. The central claims of the abstract are therefore completely unsupported by the accompanying document.

Significance. If the reported results were accompanied by a proper method description, experimental protocol, and reproducible evaluation, the work would be significant for the speech deepfake detection community: replacing self-attention with bidirectional Mamba while maintaining real-time performance on three standard benchmarks would be a useful contribution. As submitted, however, the significance cannot be assessed. The manuscript supplies no derivations, no ablation, no statistical tests, and no falsifiable details. The only verifiable content is a medical-image-synthesis ethics paper that is unrelated to the claimed topic. Thus, the potential significance of the underlying idea does not translate into significance of this submission.

major comments (4)
  1. [Full text (all sections)] The body of the manuscript is 'Ethical Medical Image Synthesis' by Jin, Sinha, Abhishek, and Hamarneh (arXiv:2508.09293), which contains no mention of Fake-Mamba, TransBiMamba, ConBiMamba, PN-BiMamba, XLSR, ASVspoof, In-The-Wild, EER, or speech deepfake detection. The abstract's quantitative claims — 0.97%, 1.74%, and 5.85% EER — are asserted with no supporting derivation or evaluation in the document. This is a load-bearing evidentiary failure: the central claim of the paper is unverifiable from the submitted text.
  2. [Abstract (comparison claims)] The abstract claims 'substantial relative gains' over XLSR-Conformer and XLSR-Mamba. No table, figure, or textual description reports how these baselines were configured, whether they were retrained under identical data splits and hyperparameter budgets, or which scoring metric was used. Without this information, the claimed improvements cannot be checked and may reflect protocol differences rather than architectural merit.
  3. [Abstract (real-time claim)] The abstract states that the framework 'maintains real-time inference across utterance lengths.' No latency measurements, hardware specifications, batch sizes, or utterance-length sweeps are provided anywhere in the manuscript. This claim is therefore unsupported and not reproducible from the submitted document.
  4. [Abstract (method description)] The core innovation is described only as three named encoders: TransBiMamba, ConBiMamba, and PN-BiMamba. The manuscript contains no equations, no architectural diagrams, no pseudocode, and no ablations isolating the contribution of each encoder. The central research question — whether bidirectional Mamba can serve as a competitive alternative to self-attention — cannot be evaluated because the proposed architecture is never defined.
minor comments (1)
  1. [General] The GitHub repository link in the abstract cannot be assessed from the manuscript; no code snapshot, license, or documentation is included. If this submission is the result of a metadata/upload error, the authors should be asked to provide the corrected full text corresponding to the abstract.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation observable: Fake-Mamba's claimed results have no supporting text in the supplied manuscript, so there is no derivation chain to be circular.

full rationale

The supplied full text is arXiv:2508.09293, 'Ethical Medical Image Synthesis' by Jin et al., which shares no content with the Fake-Mamba abstract: there is no description of TransBiMamba, ConBiMamba, PN-BiMamba, the XLSR front-end, the ASVspoof/In-The-Wild protocols, EER tables, or real-time benchmarks. The abstract of arXiv:2508.09294 asserts the central results (0.97%, 1.74%, 5.85% EER) and the SOTA comparison, but no method, equations, fitted parameters, or evaluation details are present to walk. An absent derivation cannot be circular: there is no equation that reduces to an input, no fitted value renamed as a prediction, and no load-bearing self-citation. The manuscript-text mismatch is a serious verifiability/evidentiary failure, and the quantitative claims should be treated as unsupported by this document, but it is not a circularity defect under the stated criteria. Hence the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

Only the abstract is available because the provided full text is a different paper. The central claim appears to rest on the XLSR front-end and Mamba architecture choices, but no specific free parameters or invented entities are described. Two domain assumptions are listed based directly on wording in the abstract.

assumptions (2)
  • domain assumption The XLSR front-end provides rich linguistic representations that encode subtle cues of synthetic speech.
    The abstract states 'Leveraging XLSR's rich linguistic representations, PN-BiMamba can effectively capture the subtle cues of synthetic speech.' This is an assumption about the utility of the pretrained front-end for the deepfake detection task.
  • domain assumption Bidirectional Mamba can capture both local and global artifacts in speech better than self-attention.
    The abstract frames the core innovation as replacing self-attention with bidirectional Mamba 'to capture both local and global artifacts.' The claimed advantage of the method rests on this architectural assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative." pith.science (2026). https://pith.science/paper/7GWET7AC

@misc{pith2026250809294,
  author       = {Pith},
  title        = {Pith review of: Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7GWET7AC}},
  note         = {Machine review of arXiv:2508.09294}
}
read the original abstract

Advances in speech synthesis intensify security threats, motivating real-time deepfake detection research. We investigate whether bidirectional Mamba can serve as a competitive alternative to Self-Attention in detecting synthetic speech. Our solution, Fake-Mamba, integrates an XLSR front-end with bidirectional Mamba to capture both local and global artifacts. Our core innovation introduces three efficient encoders: TransBiMamba, ConBiMamba, and PN-BiMamba. Leveraging XLSR's rich linguistic representations, PN-BiMamba can effectively capture the subtle cues of synthetic speech. Evaluated on ASVspoof 21 LA, 21 DF, and In-The-Wild benchmarks, Fake-Mamba achieves 0.97%, 1.74%, and 5.85% EER, respectively, representing substantial relative gains over SOTA models XLSR-Conformer and XLSR-Mamba. The framework maintains real-time inference across utterance lengths, demonstrating strong generalization and practical viability. The code is available at https://github.com/xuanxixi/Fake-Mamba.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

68 extracted references · 60 canonical work pages

  1. [1]

    Y. Jeon, Y. Kim, and G. G. Lee, ``Enhancing zero-shot multi-speaker tts with negated speaker representations,'' in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, 2024, pp. 18\,336--18\,344

  2. [2]

    Hsu, Y.-C

    C.-J. Hsu, Y.-C. Lin, C.-C. Lin, W.-C. Chen, H. L. Chung, C.-A. Li, Y.-C. Chen, C.-Y. Yu, M.-J. Lee, C.-C. Chen, R.-H. Huang, H. yi Lee, and D.-S. Shiu, ``Breezyvoice: Adapting tts for taiwanese mandarin with enhanced polyphone disambiguation -- challenges and insights,'' 2025. [Online]. Available: https://arxiv.org/abs/2501.17790

  3. [3]

    Y. K. Kan, K. Xu, H. Li et al., ``Voicedefense: Protecting automatic speaker verification models against black-box adversarial attacks,'' in Proc. Interspeech 2024, 2024, pp. 517--521

  4. [5]

    X. Xuan, K. kui Sin, Y. Zhou, and C. Kit, ``Translaw: Benchmarking large language models in multi-agent simulation of the collaborative translation,'' 2025. [Online]. Available: https://arxiv.org/abs/2507.00875

  5. [6]

    Zhang and C

    W. Zhang and C. Luo, ``Ge-gnn: Gated edge-augmented graph neural network for fraud detection,'' IEEE Transactions on Big Data, vol. 11, no. 4, pp. 1664--1676, 2025

  6. [7]

    B. Ding, R. Han, Z. Ma, and X. Xuan, ``Crowd density estimation based on multi-level attention maps,'' in 2021 IEEE 5th Information Technology,Networking,Electronic and Automation Control Conference (ITNEC), vol. 5, 2021, pp. 1759--1765

  7. [8]

    Du, I.-M

    J. Du, I.-M. Lin, I.-H. Chiu, X. Chen, H. Wu, W. Ren, Y. Tsao, H.-Y. Lee, and J.-S. R. Jang, ``Dfadd: The diffusion and flow-matching based audio deepfake dataset,'' in 2024 IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 921--928

  8. [9]

    J. Du, X. Chen, H. Wu, L. Zhang, I.-M. Lin, I.-H. Chiu, W. Ren, Y. Tseng, Y. Tsao, J.-S. R. Jang, and H. yi Lee, ``Codecfake-omni: A large-scale codec-based deepfake speech dataset,'' CoRR, vol. abs/2501.08238, January 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2501.08238

Show all 68 references
  1. [10]

    Gulati, J

    A. Gulati, J. Qin, C. C. Chiu et al., ``Conformer: Convolution-augmented transformer for speech recognition,'' pp. 5036--5040, 2020

  2. [11]

    X. Xuan, R. Han, and J. Gao, ``Conformer-based speaker recognition model for real-time multi-scenarios,'' Computer Engineering and Applications, vol. 60, no. 7, pp. 147--156, 2024

  3. [12]

    H. Shin, J. Heo, J. Kim et al., ``Hm-conformer: A conformer-based audio deepfake detection system with hierarchical pooling and multi-level classification token aggregation methods,'' in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing ...

  4. [13]

    D. T. Truong, R. Tao, T. Nguyen et al., ``Temporal-channel modeling in multi-head self-attention for synthetic speech detection,'' in Proceedings of Interspeech 2024, 2024, pp. 537--541

  5. [14]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin, ``Attention is all you need,'' in Advances in Neural Information Processing Systems, vol. 30. 1em plus 0.5em minus 0.4em NeurIPS, 2017. [Online]. Available: https://arxiv.org/abs...

  6. [15]

    J. Yang, R. K. Das, and H. Li, ``Significance of subband features for synthetic speech detection,'' IEEE Transactions on Information Forensics and Security, vol. 15, pp. 2160--2170, 2019

  7. [16]

    Sriskandaraja, V

    K. Sriskandaraja, V. Sethu, P. N. Le et al., ``Investigation of sub-band discriminative information between spoofed and genuine speech,'' in Interspeech, 2016, pp. 1710--1714

  8. [17]

    Zhang, D

    W. Zhang, D. Xu, X. Xuan, L. Jiang, G. Yao, R. Han, X. Lang, and C. Luo, ``Addressing noise and stochasticity in fraud detection for service networks,'' 2025. [Online]. Available: https://arxiv.org/abs/2505.00946

  9. [18]

    Zhang, D

    W. Zhang, D. Xu, G. Yao, X. Lin, R. Guan, C. Du, R. Han, X. Xuan, and C. Luo, ``Frect: Frequency-augmented convolutional transformer for robust time series anomaly detection,'' in Advanced Intelligent Computing Technology and Applications, D.-S. Huang, W. Chen, Y. Pan, and H. ...

  10. [19]

    Zhang and C

    W. Zhang and C. Luo, ``Decomposition-based multi-scale transformer framework for time series anomaly detection,'' Neural Networks, vol. 187, p. 107399, 2025. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0893608025002783

  11. [20]

    Z. Lin, J. Wang, R. Li, F. Shen, and X. Xuan, ``Primek-net: Multi-scale spectral learning via group prime-kernel convolutional neural networks for single channel speech enhancement,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processin...

  12. [21]

    Wu, H.-L

    H. Wu, H.-L. Chung, Y.-C. Lin, Y.-K. Wu, X. Chen, Y.-C. Pai, H.-H. Wang, K.-W. Chang, A. Liu, and H.-y. Lee, ``Codec- SUPERB : An in-depth analysis of sound codec models,'' in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V. Sri...

  13. [22]

    Ren, Y.-C

    W. Ren, Y.-C. Lin, H.-C. Chou, H. Wu, Y.-C. Wu, C.-C. Lee, H.-Y. Lee, H.-M. Wang, and Y. Tsao, ``Emo-codec: An in-depth look at emotion preservation capacity of legacy and neural codec models with subjective and objective evaluations,'' in 2024 Asia Pacific Signal and Informat...

  14. [23]

    Gu and T

    A. Gu and T. Dao, ``Mamba: Linear-time sequence modeling with selective state spaces,'' in First Conference on Language Modeling, 2024. [Online]. Available: https://openreview.net/forum?id=tEYskw1VY2

  15. [24]

    H. Zhao, M. Zhang, W. Zhao, and et al., `` Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference ,'' in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 10, 2025, pp. 10\,421--10\,429

  16. [25]

    B. Lenz, O. Lieber, A. Arazi, and et al., `` Jamba: Hybrid Transformer-Mamba Language Models ,'' in The Thirteenth International Conference on Learning Representations, 2025

  17. [26]

    Waleffe, W

    R. Waleffe, W. Byeon, D. Riach et al., ``An empirical study of mamba-based language models,'' arXiv preprint arXiv:2406.07887, 2024

  18. [27]

    L. Zhu, B. Liao, Q. Zhang et al., ``Vision mamba: Efficient visual representation learning with bidirectional state space model,'' in Forty-first International Conference on Machine Learning, 2024

  19. [28]

    Z. Wang, F. Kong, S. Feng, and et al., `` Is Mamba Effective for Time Series Forecasting? '' Neurocomputing, vol. 619, p. 129178, 2025

  20. [29]

    Q. Li, J. Qin, D. Cui, and et al., `` CMMamba: Channel Mixing Mamba for Time Series Forecasting ,'' Journal of Big Data, vol. 11, no. 1, p. 153, 2024

  21. [30]

    Yamagishi, X

    J. Yamagishi, X. Wang, M. Todisco et al., ``Asvspoof 2021: Accelerating progress in spoofed and deepfake speech detection,'' in ASVspoof 2021 Workshop - Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021

  22. [31]

    2783--2787

    Nicolas Müller and Pavel Czempin and Franziska Diekmann and Adam Froghyar and Konstantin Böttinger , `` Does Audio Deepfake Detection Generalize? '' in Interspeech 2022 , 2022 , pp. 2783--2787

  23. [32]

    M. H. Erol, A. Senocak, J. Feng, and et al., `` Audio Mamba: Bidirectional State Space Model for Audio Representation Learning ,'' IEEE Signal Processing Letters, 2024

  24. [33]

    Yadav and Z.-H

    S. Yadav and Z.-H. Tan, `` Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations ,'' in Proceedings of the 25th International Conference on Speech Communication and Technology (Interspeech 2024), 2024, pp. 552--556

  25. [34]

    Shams, S

    S. Shams, S. S. Dindar, X. Jiang, and et al., `` SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model ,'' in 2024 IEEE Spoken Language Technology Workshop (SLT). 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 1053--1059

  26. [35]

    Lee and J

    D. Lee and J. W. Choi, `` DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0.4em IEEE, 2025, pp. 1--5

  27. [36]

    Zhang and S

    T. Zhang and S. Ruan, `` VM-ASR: A Lightweight Dual-Stream U-Net Model for Efficient Audio Super-Resolution ,'' IEEE Transactions on Audio, Speech and Language Processing, 2025

  28. [37]

    Gao and N

    X. Gao and N. F. Chen, `` Speech-Mamba: Long-Context Speech Recognition with Selective State Space Models ,'' in 2024 IEEE Spoken Language Technology Workshop (SLT). 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 1--8

  29. [38]

    Chao, W.-H

    R. Chao, W.-H. Cheng, M. La Quatra, and et al., ``An investigation of incorporating mamba for speech enhancement,'' in 2024 IEEE Spoken Language Technology Workshop (SLT). 1em plus 0.5em minus 0.4em IEEE, 2024, pp. 302--308

  30. [39]

    W. Ren, H. Wu et al., ``Leveraging joint spectral and spatial learning with mamba for multichannel speech enhancement,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1--5

  31. [40]

    Jiang, C

    X. Jiang, C. Han, and N. Mesgarani, `` Dual-Path Mamba: Short and Long-Term Bidirectional Selective Structured State Space Models for Speech Separation ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em m...

  32. [41]

    T. H. Avenstrup, B. Elek, I. L. M \'a di, and et al., `` SepMamba: State-Space Models for Speaker Separation Using Mamba ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0.4em IEEE, 2025, pp. 1--5

  33. [42]

    Plaquet, N

    A. Plaquet, N. Tawara, M. Delcroix, and et al., `` Mamba-Based Segmentation Model for Speaker Diarization ,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0.4em IEEE, 2025, pp. 1--5

  34. [43]

    C. Fan, Y. Gao, Z. Pan, J. Zhang, H. Zhang, J. Zhang, and Z. Lv, ``Improved feature extraction network for neuro-oriented target speaker extraction,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1--5

  35. [44]

    X. Xuan, J. Dong, and T. Xuan, ``Research on front-end of asv system based on mel spectrum in noise scenario,'' in 2022 IEEE 10th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), vol. 10, 2022, pp. 2638--2642

  36. [45]

    Xuan and R

    X. Xuan and R. Han, ``Research on acoustic feature extractor for automatic speaker verification systerm,'' in 2022 IEEE 10th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), vol. 10, 2022, pp. 2628--2633

  37. [46]

    X. Xuan, R. Jin, T. Xuan, G. Du, and K. Xuan, ``Multi-scene robust speaker verification system built on improved ecapa-tdnn,'' in 2022 IEEE 6th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC ), 2022, pp. 1689--1693

  38. [47]

    X. Xuan, R. Han, and B. Ding, ``Research on speaker identification models based on cnn and additive angular margin loss,'' in 2021 2nd International Conference on Electronics, Communications and Information Technology (CECIT), 2021, pp. 1046--1050

  39. [48]

    X. Xuan, Z. Zhu, and C. Kit, ``Efficient real-time multi-scenario speaker recognition with mel-spectrogram-based hybrid tdnn for edge system,'' in INTERSPEECH 2024-Young Female* Researchers in Speech Workshop (YFRSW 2024), 2024

  40. [49]

    Y. Chen, J. Yi, J. Xue et al., ``Rawbmamba: End-to-end bidirectional state space model for audio deepfake detection,'' pp. 2720--2724, 2024

  41. [50]

    Xiao and R

    Y. Xiao and R. K. Das, ``Xlsr-mamba: A dual-column bidirectional state space model for spoofing attack detection,'' IEEE Signal Processing Letters, vol. 32, pp. 1276--1280, 2025

  42. [51]

    J. D. Hamilton, ``State-space models,'' Handbook of Econometrics, vol. 4, pp. 3039--3080, 1986

  43. [52]

    2278--2282

    Arun Babu and Changhan Wang and Andros Tjandra and Kushal Lakhotia and Qiantong Xu and Naman Goyal and Kritika Singh and Patrick von Platen and Yatharth Saraf and Juan Pino and Alexei Baevski and Alexis Conneau and Michael Auli , `` XLS-R: Self-supervised Cross-lingual Speech ...

  44. [53]

    Baevski, Y

    A. Baevski, Y. Zhou, A. Mohamed, and et al., ``wav2vec 2.0: A framework for self-supervised learning of speech representations,'' Advances in Neural Information Processing Systems, vol. 33, pp. 12\,449--12\,460, 2020

  45. [54]

    X. Xuan, Y. Xiao, R. K. Das, and T. Kinnunen, ``Multilingual source tracing of speech deepfakes: A first benchmark,'' arXiv preprint arXiv:2508.04143, 2025

  46. [55]

    Wang, Z.-C

    S.-H. Wang, Z.-C. Chen, J. Shi, M.-T. Chuang, G.-T. Lin, K.-P. Huang, D. Harwath, S.-W. Li, and H. yi Lee, ``How to learn a new language? an efficient solution for self-supervised learning models unseen languages adaption in low-resource scenario,'' 2025. [Online]. Available: ...

  47. [56]

    Lin, Y.-C

    H.-C. Lin, Y.-C. Lin et al., ``Improving speech emotion recognition in under-resourced languages via speech-to-speech translation with bootstrapping data selection,'' in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025,...

  48. [57]

    W., ``Investigating self-supervised front ends for speech spoofing countermeasures,'' in The Speaker and Language Recognition Workshop (Odyssey 2022), 2022, pp

    X. W., ``Investigating self-supervised front ends for speech spoofing countermeasures,'' in The Speaker and Language Recognition Workshop (Odyssey 2022), 2022, pp. 112--119

  49. [58]

    Zhang, Q

    X. Zhang, Q. Zhang, H. Liu, T. Xiao, X. Qian, B. Ahmed, E. Ambikairajah, H. Li, and J. Epps, ``Mamba in speech: Towards an alternative to self-attention,'' IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 1933--1948, 2025

  50. [59]

    Rosello, A

    E. Rosello, A. Gomez-Alanis, A. M. Gomez, and A. Peinado, ``A conformer-based classifier for variable-length utterance processing in anti-spoofing,'' in Interspeech 2023, 2023, pp. 5281--5285

  51. [60]

    Zhang, S

    Q. Zhang, S. Wen, and T. Hu, `` Audio deepfake detection with self-supervised XLS-R and SLS classifier ,'' in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 6765--6773

  52. [61]

    Todisco, X

    M. Todisco, X. Wang et al., `` ASVspoof 2019: Future horizons in spoofed and fake audio detection ,'' pp. 1008--1012, 2019

  53. [62]

    H. Tak, M. Kamble, J. Patino et al., ``Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,'' in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1em plus 0.5em minus 0...

  54. [63]

    Wang, W.-N

    C. Wang, W.-N. Hsu, Y. Adi, A. Polyak, A. Lee, P.-J. Chen, J. Gu, and J. Pino, ``fairseq s 2: A scalable and integrable speech synthesis toolkit,'' in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, H. Adel and S. ...

  55. [64]

    Y. Gao, C. Herold, Z. Yang, and H. Ney, ``Revisiting checkpoint averaging for neural machine translation,'' in Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, Y. He, H. Ji, S. Li, Y. Liu, and C.-H. Chang, Eds. 1em plus 0.5em minus 0.4em Association...

  56. [65]

    X. Xuan, R. Han, S. Ji, and B. Ding, ``Research on clothing image classification models based on cnn and transfer learning,'' in 2021 IEEE 5th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC), vol. 5, 2021, pp. 1461--1466

  57. [66]

    Sholokhov, M

    A. Sholokhov, M. Sahidullah, and T. Kinnunen, ``Semi-supervised speech activity detection with an application to automatic speaker verification,'' Computer Speech & Language, vol. 47, pp. 132--156, 2018

  58. [67]

    Arora, W

    S. Arora, W. Hu, and P. K. Kothari, ``An analysis of the t-sne algorithm for data visualization,'' in Conference on Learning Theory. 1em plus 0.5em minus 0.4em PMLR, 2018, pp. 1455--1462

  59. [68]

    A. Cui, C. Zhao, X. Deng, G. Jiang, Y. Yang, G. Yao, R. Han, W. Zhang, and X. Xuan, ``Unlocking the full potential of separable convolutions on tensor cores,'' in International Conference on Intelligent Computing. 1em plus 0.5em minus 0.4em Springer, 2025, pp. 39--50

  60. [69]

    fairseq S 2: A Scalable and Integrable Speech Synthesis Toolkit

    11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEco...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.