Pith. sign in

REVIEW 2 major objections 46 references

Diversity of attack methods outweighs raw data volume for robust speech anti-spoofing models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-27 19:35 UTC pith:ZPIKOQ5L

load-bearing objection Diversity beats scale for cross-dataset generalization in anti-spoofing, but the experiments leave room for dataset-specific confounds. the 2 major comments →

arxiv 2606.08038 v1 pith:ZPIKOQ5L submitted 2026-06-06 cs.SD

Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis

classification cs.SD
keywords speech anti-spoofingdataset scaledata diversitymodel generalizationcross-dataset evaluationspoofing attacksaudio securityoverfitting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tests the common assumption that bigger anti-spoofing datasets automatically produce better models. Experiments show that simply scaling up data from the same generation techniques brings little improvement and can reduce performance on unseen data due to overfitting. In contrast, a smaller training set that mixes multiple distinct attack types achieves stronger results when tested across different datasets. This leads to the conclusion that dataset design should emphasize variety in how spoofing signals are created rather than collecting ever-larger volumes under limited methods.

Core claim

Larger training sets created by expanding scale under fixed generation methods yield negligible gains and can degrade cross-domain generalization through overfitting, whereas a smaller composite training set that incorporates diverse attack types significantly outperforms those larger limited-diversity sets in cross-dataset evaluations.

What carries the argument

Controlled comparison that decouples scale from diversity by varying data volume while holding generation methods constant, then introducing diversity through dataset mixing, evaluated via cross-dataset generalization.

Load-bearing premise

The selected representative datasets and standard model architectures reveal general effects of scale versus diversity without being dominated by dataset-specific artifacts or evaluation choices.

What would settle it

A controlled test showing that a single-method dataset scaled to several times the size of the composite diverse set matches or exceeds the composite's cross-dataset accuracy would falsify the claim that diversity outweighs scale.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Excessive scaling under fixed spoofing methods leads to overfitting and weaker performance on new domains.
  • Mixing datasets with different attack generation methods improves generalization more effectively than increasing volume.
  • Dataset construction for anti-spoofing should prioritize coverage of varied generation techniques over total sample count.
  • Models trained on smaller but diverse sets can achieve better cross-dataset results than those trained on larger homogeneous sets.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same scale-versus-diversity trade-off may appear in other audio classification or security tasks that rely on synthetic or adversarial examples.
  • Practical dataset curation would benefit from explicit metrics that quantify attack-method variety rather than relying on size alone.
  • This finding suggests testing whether current large public datasets already contain enough hidden diversity or whether new collection efforts are needed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The paper challenges the assumption that larger speech anti-spoofing datasets always yield better generalization. Through experiments on representative datasets, it reports two findings: (1) excessive scaling under fixed generation methods yields negligible or negative returns on cross-domain performance due to overfitting; (2) a smaller composite training set with diverse attacks significantly outperforms larger but less diverse datasets in cross-dataset evaluations. The conclusion advocates prioritizing diversity of generation methods over scale in future dataset construction.

Significance. If the empirical findings hold after controlling for confounds, the work would meaningfully shift priorities in anti-spoofing dataset design away from indiscriminate scaling toward deliberate diversification of attack generation methods, potentially enabling more robust models with reduced data collection costs.

major comments (2)
  1. [Abstract] Abstract: the claim that 'a smaller composite training set featuring diverse attacks significantly outperforms larger-scale datasets with limited diversity in cross-dataset evaluations' is load-bearing for the central thesis, yet the abstract (and implied methods) provides no details on statistical significance testing, variance across runs, or controls that would isolate diversity from alignment between specific attack types in the composite and the evaluation distributions.
  2. [Abstract] Abstract and implied methods: the weakest link is the assumption that the chosen representative datasets and standard model architectures reveal general scale-versus-diversity effects. No ablations are described that vary model architectures, add further datasets, or explicitly control for dataset-specific artifacts (e.g., recording conditions, attack implementation details) that could explain the cross-dataset gains without invoking diversity per se.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address the major comments point by point below, indicating planned revisions where appropriate.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the claim that 'a smaller composite training set featuring diverse attacks significantly outperforms larger-scale datasets with limited diversity in cross-dataset evaluations' is load-bearing for the central thesis, yet the abstract (and implied methods) provides no details on statistical significance testing, variance across runs, or controls that would isolate diversity from alignment between specific attack types in the composite and the evaluation distributions.

    Authors: We agree the abstract would benefit from greater transparency on these points. The full manuscript already averages results over multiple random seeds and reports standard deviations to indicate run-to-run variance. We will revise the abstract to reference this practice and add a brief note on statistical testing (e.g., paired t-tests between key conditions). On isolating diversity from attack-type alignment, the composite set draws attacks from distinct generation pipelines that do not match the evaluation distributions' specific implementations; cross-dataset evaluation is the primary control. We will expand the methods/discussion section with an explicit paragraph on this design choice and any residual confounds. revision: yes

  2. Referee: [Abstract] Abstract and implied methods: the weakest link is the assumption that the chosen representative datasets and standard model architectures reveal general scale-versus-diversity effects. No ablations are described that vary model architectures, add further datasets, or explicitly control for dataset-specific artifacts (e.g., recording conditions, attack implementation details) that could explain the cross-dataset gains without invoking diversity per se.

    Authors: The experiments employ representative public datasets from the ASVspoof series and a widely used model architecture to align with community standards. We acknowledge that additional architecture ablations and extra datasets would strengthen claims of generality; these were outside the original scope due to computational cost. We will add a limitations subsection that (1) justifies the dataset and model selections, (2) discusses potential recording/implementation artifacts, and (3) explains why the observed cross-dataset gains are more plausibly attributed to attack diversity than to those artifacts. If revision timeline permits, we will include one additional architecture for a targeted ablation. revision: partial

Circularity Check

0 steps flagged

No circularity: purely empirical comparisons

full rationale

The paper reports experimental results on dataset scale versus diversity in speech anti-spoofing without any equations, derivations, fitted parameters renamed as predictions, or self-citation chains. Central claims rest on direct cross-dataset performance measurements rather than reducing to inputs by construction. This is the expected outcome for an empirical study with no mathematical derivation chain.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The work is an empirical ML study that relies on standard supervised training and cross-dataset evaluation practices rather than introducing new free parameters, axioms, or postulated entities.

axioms (1)
  • domain assumption Standard machine-learning assumptions that training and test distributions are sufficiently related for cross-dataset evaluation to be meaningful and that model training follows conventional supervised procedures.
    Invoked implicitly when interpreting cross-dataset performance as evidence of generalization.

pith-pipeline@v0.9.1-grok · 5684 in / 1261 out tokens · 25463 ms · 2026-06-27T19:35:10.591066+00:00 · methodology

0 comments
read the original abstract

The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains unclear whether indiscriminate scaling commensurately improves model generalization. This study challenges the "scale-first" paradigm by decoupling the impacts of training data scale versus diversity. Through experiments on representative datasets, we report two key findings: (1) Larger is not always better. Expanding data scale excessively under fixed generation methods yields negligible returns and may even degrade cross-domain generalization due to overfitting.(2) Diversity outweighs scale. A smaller composite training set featuring diverse attacks significantly outperforms larger-scale datasets with limited diversity in cross-dataset evaluations. We conclude that future dataset construction should prioritize the diversity of generation methods over scale to effectively enhance model generalization.

Figures

Figures reproduced from arXiv: 2606.08038 by Daixian Li, Guanxiang Feng, Jiajun Liu, Jun Xue, Yanzhen Ren, Yi Chai, Yihuan Huang, Zhuolin Yi.

Figure 1
Figure 1. Figure 1: Motivations for this work. (a) presents our survey results on the scale and diversity of training sets from repre￾sentative speech anti-spoofing datasets over the past decade. The vertical axis denotes the total duration of training data, and the color map indicates the number of generation methods covered. For datasets without an explicit data split, statistics are computed over the entire dataset. (b) il… view at source ↗
Figure 2
Figure 2. Figure 2: illustrates the detailed pipelines of our exploratory experiments. In the first experiment, training sets are randomly sampled at different proportions to compare the effects of vary￾ing data scales on model performance. Subsequently, to inves￾tigate the impact of generation methods diversity, we extract a specific number of samples from several distinct datasets to con￾struct a high-diversity composite tr… view at source ↗
Figure 3
Figure 3. Figure 3: Evaluation EER (%, ↓) of models trained on varying proportions of the (a) Speechfake-BD and (b) ASVspoof5 training sets. The first row in each heatmap shows in-domain test results, while the remaining rows report cross-dataset generalization performance. The best results are highlighted in bold. fixed, simply increasing the amount of training data is not an ef￾fective strategy for improving spoofing detect… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

46 extracted references · 8 canonical work pages · 3 internal anchors

  1. [1]

    Introduction With the rapid development of speech generation technologies such as TTS and VC, generated speech has achieved remark- able improvements in both naturalness and controllability[1]. While these advances have facilitated industrial applications and real-world deployments, they have also introduced increas- ingly severe security threats in areas...

  2. [2]

    Fake Speech Detection Fake speech detection aims to distinguish generated speech from real speech

    Related Works 2.1. Fake Speech Detection Fake speech detection aims to distinguish generated speech from real speech. Early research primarily focused on de- signing discriminative features that capture the differences be- tween the two. Spectral coefficients are utilized by exten- sive studies[11, 12, 13], while other speech attributes have also been exp...

  3. [3]

    Experimental Setups In this section, two complementary sets of experiments are de- signed to investigate how training data scale and diversity affect the performance of fake speech detection models. While there is no standardized definition of diversity in speech anti-spoofing datasets, it is broadly recognized as a multifaceted concept in- volving the di...

  4. [4]

    We conduct a systematic analysis of the results and derive the following two key findings

    Results and Analysis This section presents the experimental results of our exploration on training data scale and diversity across multiple training and test sets, with model performance evaluated using the Equal Er- ror Rate (EER) metric. We conduct a systematic analysis of the results and derive the following two key findings. Finding 1: Model performan...

  5. [5]

    Discussion Although this study analyzes the impact of training data scale and diversity on the generalization performance of speech anti- spoofing models, several aspects remain insufficiently explored and merit further investigation. Beyond training data factors, model capacity is also an im- portant determinant of generalization performance, as models w...

  6. [6]

    Motivated by this observation, we conduct two exploratory experiments to examine how training data scale and diversity affect model performance

    Conclusion In this paper, we systematically review the evolution of speech anti-spoofing datasets over the past decade and observe that the scale of corresponding training data has exhibited an exponen- tial growth trend. Motivated by this observation, we conduct two exploratory experiments to examine how training data scale and diversity affect model per...

  7. [7]

    Acknowledgements This work is supported by the Natural Science Foundation of China (NSFC) under the grant NO.62572358, 62372334

  8. [8]

    The research design, experiments, analysis, and conclusions were conducted and verified entirely by the authors

    Generative AI Use Disclosure Generative AI tools (GPT-4o) were employed exclusively for language refinement and grammatical corrections in this manuscript. The research design, experiments, analysis, and conclusions were conducted and verified entirely by the authors. All authors are responsible and accountable for the work and the content of this paper

  9. [9]

    Towards con- trollable speech synthesis in the era of large language models: A systematic survey,

    T. Xie, Y . Rong, P. Zhang, W. Wang, and L. Liu, “Towards con- trollable speech synthesis in the era of large language models: A systematic survey,” inProceedings of the 2025 Conference on Em- pirical Methods in Natural Language Processing, 2025, pp. 764– 791

  10. [10]

    A survey on speech deep- fake detection,

    M. Li, Y . Ahmadiadli, and X.-P. Zhang, “A survey on speech deep- fake detection,”ACM Computing Surveys, vol. 57, no. 7, pp. 1–38, 2025

  11. [11]

    Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,

    Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilc ¸i, M. Sahidullah, and A. Sizov, “Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,” in Interspeech 2015. ISCA, Sep. 2015, pp. 2037–2041

  12. [12]

    Asvspoof 2019: A large-scale public database of synthe- sized, converted and replayed speech,

    X. Wang, J. Yamagishi, M. Todisco, H. Delgado, A. Nautsch, N. Evans, M. Sahidullah, V . Vestman, T. Kinnunen, K. A. Lee et al., “Asvspoof 2019: A large-scale public database of synthe- sized, converted and replayed speech,”Computer Speech & Lan- guage, vol. 64, p. 101114, 2020

  13. [13]

    ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,

    J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K. A. Lee, T. Kinnunen, N. Evans, and H. Delgado, “ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,” in2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, pp. 47–54

  14. [14]

    ASVspoof 5: crowdsourced speech data, deepfakes, and adversarial attacks at scale,

    X. Wang, H. Delgado, H. Tak, J.-w. Jung, H.-j. Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. H. Kinnunen, N. Evans, K. A. Lee, and J. Yamagishi, “ASVspoof 5: crowdsourced speech data, deepfakes, and adversarial attacks at scale,” inThe Auto- matic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024). ISCA, Aug. 2024, pp. 1–8

  15. [15]

    Speechfake: A large-scale multilingual speech deepfake dataset incorporating cutting-edge generation methods,

    W. Huang, Y . Gu, Z. Wang, H. Zhu, and Y . Qian, “Speechfake: A large-scale multilingual speech deepfake dataset incorporating cutting-edge generation methods,” inProceedings of the 63rd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 9985–9998

  16. [16]

    Post-training for deepfake speech detection,

    W. Ge, X. Wang, X. Liu, and J. Yamagishi, “Post-training for deepfake speech detection,”arXiv preprint arXiv:2506.21090, 2025

  17. [17]

    Spoofceleb: Speech deepfake detection and sasv in the wild,

    J.-w. Jung, Y . Wu, X. Wang, J.-H. Kim, S. Maiti, Y . Matsunaga, H.-j. Shim, J. Tian, N. Evans, J. S. Chung, W. Zhang, S. Um, S. Takamichi, and S. Watanabe, “Spoofceleb: Speech deepfake detection and sasv in the wild,”IEEE Open Journal of Signal Pro- cessing, vol. 6, pp. 68–77, 2025

  18. [18]

    A data-centric approach to generalizable speech deepfake detection,

    W. Huang, Y . Mao, and Y . Qian, “A data-centric approach to generalizable speech deepfake detection,”arXiv preprint arXiv:2512.18210, 2025

  19. [19]

    Deep residual neu- ral networks for audio spoofing detection,

    M. Alzantot, Z. Wang, and M. B. Srivastava, “Deep residual neu- ral networks for audio spoofing detection,” inInterspeech 2019, 2019, pp. 1078–1082

  20. [20]

    Deepfake audio detection via mfcc fea- tures using machine learning,

    A. Hamza, A. R. R. Javed, F. Iqbal, N. Kryvinska, A. S. Almadhor, Z. Jalil, and R. Borghol, “Deepfake audio detection via mfcc fea- tures using machine learning,”IEEE Access, vol. 10, pp. 134 018– 134 028, 2022

  21. [21]

    Is synthetic voice detection research going into the right direction?

    S. Borz `ı, O. Giudice, F. Stanco, and D. Allegra, “Is synthetic voice detection research going into the right direction?” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion Workshops (CVPRW), 2022, pp. 71–80

  22. [22]

    Audio deepfake detection based on a combination of f0 information and real plus imaginary spectrogram features,

    J. Xue, C. Fan, Z. Lv, J. Tao, J. Yi, C. Zheng, Z. Wen, M. Yuan, and S. Shao, “Audio deepfake detection based on a combination of f0 information and real plus imaginary spectrogram features,” inProceedings of the 1st international workshop on deepfake de- tection for audio multimedia, 2022, pp. 19–26

  23. [23]

    Learning from yourself: A self-distillation method for fake speech detection,

    J. Xue, C. Fan, J. Yi, C. Wang, Z. Wen, D. Zhang, and Z. Lv, “Learning from yourself: A self-distillation method for fake speech detection,” inICASSP 2023-2023 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5

  24. [24]

    Deepfake speech detection through emotion recognition: A semantic approach,

    E. Conti, D. Salvi, C. Borrelli, B. Hosler, P. Bestagini, F. An- tonacci, A. Sarti, M. C. Stamm, and S. Tubaro, “Deepfake speech detection through emotion recognition: A semantic approach,” in ICASSP 2022 - 2022 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2022, pp. 8962– 8966

  25. [25]

    Bts-e: Au- dio deepfake detection using breathing-talking-silence encoder,

    T.-P. Doan, L. Nguyen-Vu, S. Jung, and K. Hong, “Bts-e: Au- dio deepfake detection using breathing-talking-silence encoder,” inICASSP 2023 - 2023 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5

  26. [26]

    End-to-end anti-spoofing with rawnet2,

    H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with rawnet2,” inICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 6369–6373

  27. [27]

    End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detec- tion,

    H. Tak, J. weon Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detec- tion,” in2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, pp. 1–8

  28. [28]

    Aasist: Audio anti-spoofing using in- tegrated spectro-temporal graph attention networks,

    J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, “Aasist: Audio anti-spoofing using in- tegrated spectro-temporal graph attention networks,” inICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 6367–6371

  29. [29]

    Automatic speaker verification spoofing and deep- fake detection using wav2vec 2.0 and data augmentation,

    H. Tak, M. Todisco, X. Wang, J. weon Jung, J. Yamagishi, and N. Evans, “Automatic speaker verification spoofing and deep- fake detection using wav2vec 2.0 and data augmentation,” inThe Speaker and Language Recognition Workshop (Odyssey 2022), 2022, pp. 112–119

  30. [30]

    Multi-level ssl feature gating for audio deepfake detection,

    H. M. Tran, D. Lolive, A. Sini, A. Delhay, P.-F. Marteau, and D. Guennec, “Multi-level ssl feature gating for audio deepfake detection,” inProceedings of the 33rd ACM International Confer- ence on Multimedia, 2025, pp. 11 766–11 775

  31. [31]

    Fmfcc-a: A challeng- ing mandarin dataset for synthetic speech detection,

    Z. Zhang, Y . Gu, X. Yi, and X. Zhao, “Fmfcc-a: A challeng- ing mandarin dataset for synthetic speech detection,” inDigi- tal Forensics and Watermarking: 20th International Workshop, IWDW 2021, Beijing, China, November 20-22, 2021, Revised Se- lected Papers. Berlin, Heidelberg: Springer-Verlag, 2021, pp. 117–131

  32. [32]

    Add 2022: the first audio deep synthe- sis detection challenge,

    J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, C. Wang, T. Wang, Z. Tian, Y . Bai, C. Fanet al., “Add 2022: the first audio deep synthe- sis detection challenge,” inICASSP 2022-2022 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 9216–9220

  33. [33]

    Add 2023: the second audio deepfake detection challenge,

    J. Yi, J. Tao, R. Fu, X. Yan, C. Wang, T. Wang, C. Y . Zhang, X. Zhang, Y . Zhao, Y . Renet al., “Add 2023: the second audio deepfake detection challenge,”arXiv preprint arXiv:2305.13774, 2023

  34. [34]

    Diffssd: A diffusion-based dataset for speech forensics,

    K. Bhagtani, A. K. S. Yadav, P. Bestagini, and E. J. Delp, “Diffssd: A diffusion-based dataset for speech forensics,” inICASSP 2025- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  35. [35]

    Diffuse or confuse: A diffu- sion deepfake speech dataset,

    A. Firc, K. Malinka, and P. Han´aˇcek, “Diffuse or confuse: A diffu- sion deepfake speech dataset,” in2024 International Conference of the Biometrics Special Interest Group (BIOSIG). IEEE, Sep. 2024, pp. 1–7

  36. [36]

    Codecfake: Enhancing anti-spoofing models against deepfake audios from codec-based speech synthesis systems,

    H. Wu, Y . Tseng, and H.-y. Lee, “Codecfake: Enhancing anti-spoofing models against deepfake audios from codec-based speech synthesis systems,” inInterspeech 2024. ISCA, Sep. 2024, pp. 1770–1774

  37. [37]

    The codecfake dataset and countermeasures for the universally detection of deepfake audio,

    Y . Xie, Y . Lu, R. Fu, Z. Wen, Z. Wang, J. Tao, X. Qi, X. Wang, Y . Liu, H. Cheng, L. Ye, and Y . Sun, “The codecfake dataset and countermeasures for the universally detection of deepfake audio,” IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 386–400, 2025

  38. [38]

    Collecting, curating, and annotating good qual- ity speech deepfake dataset for famous figures: Process and chal- lenges,

    H. Ali, S. Subramani, R. Varahamurthy, N. Adupa, L. Bollinani, and H. Malik, “Collecting, curating, and annotating good qual- ity speech deepfake dataset for famous figures: Process and chal- lenges,” inInterspeech 2025, Aug. 2025, pp. 3928–3932

  39. [39]

    Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

    J. Xue, T. Zhang, Z. Yi, Y . Huang, Y . Chai, Y . Zhang, and Y . Ren, “Profiling the voice: Speaker-specific phoneme fingerprinting for speech deepfake detection,”arXiv preprint arXiv:2605.17737, 2026

  40. [40]

    Fake speech wild: Detecting deepfake speech on social media platform,

    Y . Xie, R. Fu, X. Wang, Z. Wang, Y . Li, Z. Wen, H. Cheng, and L. Ye, “Fake speech wild: Detecting deepfake speech on social media platform,”arXiv preprint arXiv:2508.10559, 2025

  41. [41]

    Does audio deepfake detection generalize?

    N. M ¨uller, P. Czempin, F. Diekmann, A. Froghyar, and K. B¨ottinger, “Does audio deepfake detection generalize?” inIn- terspeech 2022. ISCA, Sep. 2022, pp. 2783–2787

  42. [42]

    RTCFake: Speech Deepfake Detection in Real-Time Communication

    J. Xue, Z. Yi, Y . Huang, Y . Ren, Y . Chen, C. Fan, Z. Su, Y . Zhang, and B. Cai, “Rtcfake: Speech deepfake detection in real-time communication,”arXiv preprint arXiv:2604.23742, 2026

  43. [43]

    V oicewukong: Benchmarking deepfake voice detection,

    Z. Yan, Y . Zhao, and H. Wang, “V oicewukong: Benchmarking deepfake voice detection,” in34th USENIX Security Symposium (USENIX Security 25), 2025, pp. 4561–4580

  44. [44]

    Mlaad: The multi- language audio anti-spoofing dataset,

    N. M. M ¨uller, P. Kawa, W. H. Choong, E. Casanova, E. G ¨olge, T. M¨uller, P. Syga, P. Sperl, and K. B¨ottinger, “Mlaad: The multi- language audio anti-spoofing dataset,” in2024 International Joint Conference on Neural Networks (IJCNN). IEEE, 2024, pp. 1–7

  45. [45]

    Cross- domain audio deepfake detection: Dataset and analysis,

    Y . Li, M. Zhang, M. Ren, M. Ma, D. Wei, and H. Yang, “Cross- domain audio deepfake detection: Dataset and analysis,” Sep. 2024, arXiv:2404.04904 [cs]

  46. [46]

    Raw- boost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,

    H. Tak, M. Kamble, J. Patino, M. Todisco, and N. Evans, “Raw- boost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,” inICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 6382–6386