REVIEW 2 major objections 4 minor 44 references
Non-speech intervals are the dominant shortcut that deepfake audio detectors learn from standard spoofing corpora.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 04:30 UTC pith:I7EMU7MA
load-bearing objection Clean diagnostic that turns the known non-speech artifact problem into a measurable Cd-vs-Ci distinction; the Z-preservation premise is the only real soft spot and it is already flagged by the authors. the 2 major comments →
An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Non-speech interventions produce the largest performance shifts of any tested category, confirming that non-speech intervals form a confounded shortcut dependency: the training protocol creates a spurious association between non-speech structure and the spoof label, and the learned representation leaks that structure rather than remaining independent of it given the intrinsic generative artifacts.
What carries the argument
The directed graphical model that partitions the waveform into intrinsic artifacts Z, idiosyncratic pipeline artifacts Cd, and exogenous channel factors Ci, together with the causal-sufficiency condition that the model representation must satisfy ˆZ ⊥⊥ (Cd, Ci) | Z; controlled acoustic interventions then test whether performance collapses when Cd is altered while Z is left untouched.
Load-bearing premise
The framework assumes that post-hoc waveform edits such as padding silence or noise never change the true generative fingerprint of the synthesizer, only the idiosyncratic or channel factors around it.
What would settle it
If an intervention that is claimed to touch only non-speech structure also systematically destroys or fabricates known vocoder-phase or spectral artifacts inside the speech regions, and the same performance collapse is observed even for models known to ignore non-speech, then the diagnostic mapping from performance drop to shortcut reliance is invalid.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an intervention-based diagnostic framework for identifying shortcut learning in audio deepfake (spoofing) countermeasures. It models the data-generating process as a DAG that separates intrinsic synthesis artifacts Z from idiosyncratic pipeline artifacts Cd and exogenous channel effects Ci, then defines confounded shortcut dependency via two conditions: a training-distribution association Cd ̸⊥⊥ S and representational leakage ˆZ ̸⊥⊥ Cd | Z. The framework is operationalised with corpus-level JSD analysis of eleven acoustic descriptors plus controlled waveform interventions (non-speech padding, spectral band-cuts/downsampling, energy/noise/peak-norm) applied to evaluation data. Five training configurations of XLS-R-300M + RawGAT-ST are evaluated on ASVspoof 2019/2021 LA and ASVspoof 5. Relative DCF degradation δm,p and embedding geometry show that non-speech interventions produce the largest shifts (δ > 60 for RB-19), while codec/channel effects degrade models more uniformly, supporting the claim that non-speech structure is a dominant confounded shortcut.
Significance. If the diagnostic logic holds, the work supplies a reusable, falsifiable protocol that the anti-spoofing community currently lacks: a formal criterion distinguishing shortcut exploitation from ordinary domain shift, together with concrete acoustic interventions and a sensitivity metric. The empirical core is solid—five configurations, three evaluation regimes, full intervention table, embedding analysis, and a clean codec contrast—and the non-speech finding is consistent with prior ASVspoof protocol design choices. The DAG formalisation and the explicit Z-preservation premise give the results a clearer interpretive status than purely empirical perturbation studies. Strengths include the transparent reporting of both relative and absolute costs, the contrastive Ci test, and the demonstration that simply enlarging the training corpus does not remove the confound when the new data preserve the same Cd–S association.
major comments (2)
- [Section 2.4, Table 2, Fig. 6] Section 2.4 (and the diagnostic claim built on it) rests on the premise that post-hoc waveform interventions modify only Cd (or mask observability of Z) and never alter the intrinsic generative fingerprint Z itself. The qualitative argument is clear, yet the manuscript supplies no direct empirical check that synthesis-specific traces remain intact under the non-speech interventions that drive the headline result (Table 2, RB-19 δm,p > 60; Fig. 6). Spectral band-cut S1 already produces large δ under the Both target and is interpreted by the authors as Z-masking; an analogous verification (e.g., a known vocoder-phase or high-frequency artifact measure before/after padding) would substantially strengthen the inference that the non-speech performance collapse is pure shortcut reliance rather than inadvertent destruction of Z.
- [Table 2, Eq. (4)] Table 2 reports point estimates of relative DCF degradation without uncertainty (bootstrap intervals, multiple random seeds, or even standard errors). Several of the decisive cells for RB-19 are extreme (δ > 60) while others are near zero or negative; without a measure of variability it is difficult to judge whether the ranking of intervention categories is stable or sensitive to threshold choice and finite-sample effects. Adding even a modest uncertainty quantification would make the sensitivity profiles more conclusive.
minor comments (4)
- [Figure 3] Figure 3 caption and surrounding text refer to “arrows: shift from adding AS5 data,” but the visual encoding of circle size (distribution shift) versus arrow direction is dense; a short legend or colour key would improve readability.
- [Figures 1–2] Notation for the internal representation switches between ˆZ and Zrn / ˆZ in the architecture diagram (Fig. 2) and the causal graph (Fig. 1); a single consistent symbol would reduce cognitive load.
- [Section 3.1] The custom DA pipeline is described as “inspired by strategies from recent ASVspoof 5 submissions” but the precise probability schedule for each stage is not tabulated; a short supplementary table would aid reproducibility.
- [Throughout] Typographical inconsistencies appear in author affiliations and a few reference entries (e.g., “V oIP”, “mad tx”); a light copy-edit pass would clean these.
Circularity Check
No significant circularity: the DAG is an interpretive scaffold; measured DCF/EER/embedding shifts are external empirical quantities independent of the definitions.
full rationale
The paper's derivation chain is: (1) posit a DAG of the data-generating process distinguishing Z (intrinsic) from Cd (idiosyncratic) and Ci (exogenous); (2) define confounded shortcut dependency via two conditions (spurious Cd–S association under Dtrain plus representational leakage ˆZ ̸⊥⊥ Cd | Z); (3) design post-hoc waveform interventions claimed to alter only Cd (or mask Z) while leaving Z fixed; (4) measure relative DCF degradation δm,p and embedding cosine shifts on real ASVspoof models/data. None of these steps reduces a claimed prediction or first-principles result to its own inputs by construction. The DAG and independence statements are definitional scaffolding, not fitted parameters or uniqueness theorems. The quantitative results (Table 2 δm,p > 60 under non-speech padding for RB-19; Fig. 6 inter-class collapse; JSD corpus analysis; codec contrast) are ordinary empirical measurements on held-out evaluation sets. Self-citations are to standard ASVspoof protocols, RawBoost, Silero VAD, and prior artifact papers; none is load-bearing for the central claim that non-speech is a dominant shortcut. The Z-preservation premise is an assumption (flagged by the reader), not a circular reduction. Score 0 is therefore warranted.
Axiom & Free-Parameter Ledger
free parameters (4)
- non-speech padding duration =
4.0 s
- peak-normalization targets =
0.65 / 0.45
- AWGN SNR levels =
20/10/5 dB
- training hyper-parameters =
lr=1e-6, wd=1e-4, w=(0.1,0.9)
axioms (3)
- domain assumption Intrinsic artifacts Z are imprinted at generation time and cannot be altered by any post-hoc waveform perturbation.
- ad hoc to paper The data-generating process factors as the DAG of Fig. 1 (M → S, M → Z, M → Cd, Ci → x, etc.).
- domain assumption Causal sufficiency ˆZ ⊥⊥ (Cd, Ci) | Z is the ideal that a robust countermeasure should satisfy.
invented entities (2)
-
Idiosyncratic artifacts Cd
no independent evidence
-
Intrinsic artifacts Z
no independent evidence
read the original abstract
While deepfake audio detection systems achieve high performance in controlled benchmarks, their reliability often diminishes in the wild. Prior work shows that dataset-specific artifacts contribute to this gap. Yet, systematic tools to identify which acoustic properties a model exploits as shortcuts remain limited. We propose an intervention-based diagnostic framework, grounded in a directed graphical model, that formally distinguishes confound-driven shortcut dependencies from legitimate domain shift. We operationalise this through controlled acoustic perturbations targeting non-speech structure, spectral content, and signal energy, complemented by corpus-level distributional analysis. Evaluating XLS-R-300M with RawGAT-ST across ASVspoof challenges datasets, we quantify model sensitivity to specific intervention types. Results reveal that non-speech interventions produce the largest performance shifts, confirming non-speech intervals as a dominant shortcut.
Figures
Reference graph
Works this paper leans on
-
[1]
shortcut learn- ing
Introduction Detecting synthetic speech has become a critical challenge, and deep learning models have seemingly risen to the occa- sion, achieving remarkably low error rates on standard bench- marks [1, 2]. However, this success is often an illusion. When these highly accurate models are tested in the wild, facing new datasets, different scenarios, or un...
2017
-
[2]
Problem Formulation 2.1. Task Definition and Notation Letx∈R T denote a raw speech waveform ofTsamples and let S∈ {0,1}indicate the true nature of the utterance, whereS=1 denotes bonafide andS=0spoofed. A hybrid spoofing coun- termeasure learns a mappingf θ :R T →[0,1], decomposed as fθ =h ψ ◦g ϕ,whereg ϕ is a self-supervised front-end that pro- duces a s...
Pith/arXiv arXiv 2026
-
[3]
Architectural Configuration The hybrid countermeasure follows thef θ =h ψ ◦g ϕ for- mulation (Section 2.1, Fig
Experimental Setup 3.1. Architectural Configuration The hybrid countermeasure follows thef θ =h ψ ◦g ϕ for- mulation (Section 2.1, Fig. 2). The front-endg ϕ is XLS- R-300M [12], held frozen or fine-tuned. The classifierh ψ follows the RawGAT-ST architecture [29], producing a time- independent embedding ˆZ∈R 160 mapped to two class logits. Training and Eva...
2019
-
[4]
We first establish the gen- eralisation gap to frame our diagnostic question
Results and Analysis This section presents the empirical application of the diagnostic framework developed in Section 2. We first establish the gen- eralisation gap to frame our diagnostic question. Next, we test for shortcut dependency through perturbation-based interven- tions and representational analysis. Finally, we analyse codec and channel-driven d...
2019
-
[5]
Conclusion In this paper, we have proposed an intervention-based diag- nostic framework that formally distinguishes shortcut learning from legitimate domain shift in spoofing countermeasures. By grounding our analysis in a directed acyclic graph, we derived the necessary conditions to identify when a model abandons the true generative footprint (Z) to exp...
-
[6]
Acknowledgements This work has received funding from MCIN/AEI/10.13039/501100011033 under Grant PID2024- 155948OB-C53
-
[7]
Learn from real: reality defender’s submission to ASVspoof5 Challenge,
Y . Zhu, C. Goel, S. Koppisetti, T. Tran, A. Kumar, and G. Bharaj, “Learn from real: reality defender’s submission to ASVspoof5 Challenge,” inThe Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 116– 123
2024
-
[8]
Temporal variability and multi-viewed self-supervised representations to tackle the ASVspoof5 Deepfake Challenge,
Y . Xie, X. Wang, Z. Wang, R. Fu, W. Zhengqi, H. Cheng, and L. Ye, “Temporal variability and multi-viewed self-supervised representations to tackle the ASVspoof5 Deepfake Challenge,” inThe Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 101–108
2024
-
[9]
Does Audio Deepfake Detection Generalize?
N. M ¨uller, P. Czempin, F. Diekmann, A. Froghyar, and K. B ¨ottinger, “Does Audio Deepfake Detection Generalize?” in Interspeech 2022, 2022, pp. 2783–2787
2022
-
[10]
Asvspoof 2021: Towards spoofed and deepfake speech de- tection in the wild,
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kin- nunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, and K. A. Lee, “Asvspoof 2021: Towards spoofed and deepfake speech de- tection in the wild,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2507–2522, 2023
2021
-
[11]
Whispeak speech deepfake detec- tion systems for the ASVspoof5 Challenge,
P. Falez and T. Marteau, “Whispeak speech deepfake detec- tion systems for the ASVspoof5 Challenge,” inThe Auto- matic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 32–35
2024
-
[12]
Asvspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,
X. Wang, H. Delgado, H. Tak, J. weon Jung, H. jin Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, and J. Yamagishi, “Asvspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,” 2024. [Online]. Available: https: //arxiv.org/abs/2408.08739
Pith/arXiv arXiv 2024
-
[13]
BUT systems and analyses for the ASVspoof 5 Challenge,
J. Rohdin, L. Zhang, P. Old ˇrich, V . Stanˇek, D. Mihola, J. Peng, T. Stafylakis, D. Beveraki, A. Silnova, J. Brukner, and L. Bur- get, “BUT systems and analyses for the ASVspoof 5 Challenge,” inThe Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 24–31
2024
-
[14]
WavLM model ensemble for audio deepfake detection,
D. Combei, A. Stan, D. Oneata, and H. Cucu, “WavLM model ensemble for audio deepfake detection,” inThe Auto- matic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 170–175
2024
-
[15]
Exploring generalization to unseen au- dio data for spoofing: insights from SSL models,
A. Kulkarni, H. M. Tran, A. Kulkarni, S. Dowerah, D. Lo- live, and M. M. Doss, “Exploring generalization to unseen au- dio data for spoofing: insights from SSL models,” inThe Auto- matic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 86–93
2024
-
[16]
wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,” inAdvances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 12 449–12 460. [Online]. Available: https://proceedings.neur...
2020
-
[17]
Wavlm: Large-scale self-supervised pre-training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, p. 1505–1518, Oct. 2022. [Online]....
-
[18]
Xls-r: Self-supervised cross- lingual speech representation learning at scale,
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y . Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “Xls-r: Self-supervised cross- lingual speech representation learning at scale,” 2021. [Online]. Available: https://arxiv.org/abs/2111.09296
Pith/arXiv arXiv 2021
-
[19]
Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,
J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, “Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” inarXiv preprint arXiv:2110.01200, 2021
Pith/arXiv arXiv 2021
-
[20]
AASIST3: KAN-enhanced AASIST speech deepfake detection using SSL features and additional regularization for the ASVspoof 2024 Challenge,
K. Borodin, V . Kudryavtsev, D. Korzh, A. Efimenko, G. Mkrtchian, M. Gorodnichev, and O. Y . Rogov, “AASIST3: KAN-enhanced AASIST speech deepfake detection using SSL features and additional regularization for the ASVspoof 2024 Challenge,” inThe Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 48–55
2024
-
[21]
ASVspoof 5 Challenge: advanced ResNet architectures for robust voice spoofing detec- tion,
A.-T. Dao, M. Rouvier, and D. Matrouf, “ASVspoof 5 Challenge: advanced ResNet architectures for robust voice spoofing detec- tion,” inThe Automatic Speaker Verification Spoofing Counter- measures Workshop (ASVspoof 2024), 2024, pp. 163–169
2024
-
[22]
Enhancing spoofing detection in ASVspoof 5 Workshop 2024: fusion of WavLM- ResNet18-SA for optimal performance against speech deepfakes ,
P.-C. Chan, W.-Y . Chen, and J.-C. Wang, “Enhancing spoofing detection in ASVspoof 5 Workshop 2024: fusion of WavLM- ResNet18-SA for optimal performance against speech deepfakes ,” inThe Automatic Speaker Verification Spoofing Countermea- sures Workshop (ASVspoof 2024), 2024, pp. 158–162
2024
-
[23]
Shortcut learning in deep neural networks,
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,”Nature Machine Intelligence, vol. 2, no. 11, p. 665–673, Nov. 2020. [Online]. Available: http: //dx.doi.org/10.1038/s42256-020-00257-z
-
[24]
ASVspoof 2017 Version 2.0: meta- data analysis and baseline enhancements ,
H. Delgado, M. Todisco, M. Sahidullah, N. Evans, T. Kinnunen, K. A. Lee, and J. Yamagishi, “ASVspoof 2017 Version 2.0: meta- data analysis and baseline enhancements ,” inThe Speaker and Language Recognition Workshop (Odyssey 2018), 2018, pp. 296– 303
2017
-
[25]
Dataset artefacts in anti-spoofing systems: A case study on the asvspoof 2017 bench- mark,
B. Chettri, E. Benetos, and B. L. T. Sturm, “Dataset artefacts in anti-spoofing systems: A case study on the asvspoof 2017 bench- mark,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 3018–3028, 2020
2017
-
[26]
Speech is Silver, Silence is Golden: What do ASVspoof-trained Models Really Learn?
N. M ¨uller, F. Dieckmann, P. Czempin, R. Canals, K. B ¨ottinger, and J. Williams, “Speech is Silver, Silence is Golden: What do ASVspoof-trained Models Really Learn?” in2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, pp. 55–60
2021
-
[27]
Asvspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech,
X. Wang, H. Delgado, H. Tak, J. weon Jung, H. jin Shim, M. Todisco, I. Kukanov, X. Liu, M. Sahidullah, T. Kinnunen, N. Evans, K. A. Lee, J. Yamagishi, M. Jeong, G. Zhu, Y . Zang, Y . Zhang, S. Maiti, F. Lux, N. M¨uller, W. Zhang, C. Sun, S. Hou, S. Lyu, S. Le Maguer, C. Gong, H. Guo, L. Chen, and V . Singh, “Asvspoof 5: Design, collection and validation o...
2026
-
[28]
Exploring Self-supervised Embeddings and Syn- thetic Data Augmentation for Robust Audio Deepfake Detection,
J. M. Mart ´ın-Do˜nas, A. ´Alvarez, E. Rosello, A. M. Gomez, and A. M. Peinado, “Exploring Self-supervised Embeddings and Syn- thetic Data Augmentation for Robust Audio Deepfake Detection,” inInterspeech 2024, 2024, pp. 2085–2089
2024
-
[29]
Intema system description for the ASVspoof5 Challenge: power weighted score fusion,
A. Aliyev and A. Kondratev, “Intema system description for the ASVspoof5 Challenge: power weighted score fusion,” inThe Automatic Speaker Verification Spoofing Countermeasures Work- shop (ASVspoof 2024), 2024, pp. 152–157
2024
-
[30]
A study of guided masking data augmentation for deepfake speech detection,
D.-T. Truong, Y . Wang, K. A. Lee, M. Li, H. Nishizaki, and E. S. Chng, “A study of guided masking data augmentation for deepfake speech detection,” inThe Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 176–180
2024
-
[31]
USTC-KXDIGIT system description for ASVspoof5 Challenge,
Y . Chen, H. Wu, N. Jiang, X. Xia, Q. Gu, Y . Hao, P. Cai, Y . Guan, J. Wang, W.-L. Xie, L. Fang, S. Fang, Y . Song, W. Guo, L. Liu, and M. Xu, “USTC-KXDIGIT system description for ASVspoof5 Challenge,” inThe Automatic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 109–115
2024
-
[32]
Robust audio deepfake de- tection: exploring front-/back-end combinations and data aug- mentation strategies for the ASVspoof5 Challenge,
K. Sch ¨afer, J.-E. Choi, and M. Neu, “Robust audio deepfake de- tection: exploring front-/back-end combinations and data aug- mentation strategies for the ASVspoof5 Challenge,” inThe Auto- matic Speaker Verification Spoofing Countermeasures Workshop (ASVspoof 2024), 2024, pp. 56–63
2024
-
[33]
How to Construct Perfect and Worse-than-Coin-Flip Spoofing Countermeasures: A Word of Warning on Shortcut Learning,
H. jin Shim, R. Gonzalez Hautam ¨aki, M. Sahidullah, and T. Kin- nunen, “How to Construct Perfect and Worse-than-Coin-Flip Spoofing Countermeasures: A Word of Warning on Shortcut Learning,” inInterspeech 2023, 2023, pp. 785–789
2023
-
[34]
Shortcut learning in binary classifier black boxes: Applications to voice anti-spoofing and biometrics,
M. Sahidullah, H.-j. Shim, R. G. Hautam ¨aki, and T. H. Kinnunen, “Shortcut learning in binary classifier black boxes: Applications to voice anti-spoofing and biometrics,”IEEE Journal of Selected Topics in Signal Processing, 2025
2025
-
[35]
End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detec- tion,
H. Tak, J. weon Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detec- tion,” inProc. 2021 Edition of the Automatic Speaker Verification and Spoofing Countermeasures Challenge, 2021, pp. 1–8
2021
-
[36]
Asvspoof 2019: Future horizons in spoofed and fake audio detection,
M. Todisco, X. Wang, V . Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. H. Kinnunen, and K. A. Lee, “Asvspoof 2019: Future horizons in spoofed and fake audio detection,” inInterspeech 2019, 2019, pp. 1008–1012
2019
-
[37]
Bishop,Pattern recognition and machine learning
C. Bishop,Pattern recognition and machine learning. Springer New York, 2006, vol. 4. [Online]. Available: http://scholar.google.com/scholar.bib?q=info:jYxggZ6Ag1YJ: scholar.google.com/&output=citation&hl=en&as sdt=0,5& as vis=1&ct=citation&cd=0
2006
-
[38]
Pearl,Causality: Models, Reasoning and Inference, 2nd ed
J. Pearl,Causality: Models, Reasoning and Inference, 2nd ed. USA: Cambridge University Press, 2009
2009
-
[39]
Raw- boost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,
H. Tak, M. Kamble, J. Patino, M. Todisco, and N. Evans, “Raw- boost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing,” inIEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022
2022
-
[40]
A study on data augmentation of reverberant speech for robust speech recognition,
T. Ko, V . Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” inProc. ICASSP, 2017, pp. 5220–5224
2017
-
[41]
Wiener filter and deep neural networks: A well-balanced pair for speech enhancement,
D. Ribas, A. Miguel, A. Ortega, and E. Lleida, “Wiener filter and deep neural networks: A well-balanced pair for speech enhancement,”Applied Sciences, vol. 12, no. 18, 2022. [Online]. Available: https://www.mdpi.com/2076-3417/12/18/9000
2022
-
[42]
Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier,
S. Team, “Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier,” https:// github.com/snakers4/silero-vad, 2024
2024
-
[43]
Divergence measures based on the shannon entropy,
J. Lin, “Divergence measures based on the shannon entropy,” IEEE Transactions on Information Theory, vol. 37, no. 1, pp. 145– 151, 1991
1991
-
[44]
Tandem assessment of spoofing coun- termeasures and automatic speaker verification: Fundamentals,
T. Kinnunen, H. Delgado, N. Evans, K. A. Lee, V . Vestman, A. Nautsch, M. Todisco, X. Wang, M. Sahidullah, J. Yamag- ishi, and D. A. Reynolds, “Tandem assessment of spoofing coun- termeasures and automatic speaker verification: Fundamentals,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, vol. 28, pp. 2195–2210, 2020
2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.