REVIEW 4 major objections 4 minor 80 references
Improving Generalization for AI-Synthesized Voice Detection
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A disentanglement framework that trains on shared vocoder artifacts cuts unseen-vocoder detection error by up to 7.59 percentage points.
desk verdict A promising empirical recipe whose central mechanism is undermined by a sign error in the mutual information loss, yet the engineering is worth referee time. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the disentanglement-and-flattening pipeline. An encoder backbone is split into content and artifact branches: the artifact branch yields domain-specific features, which identify the vocoder, and domain-agnostic features, which flag synthetic voice regardless of generator. Cross-reconstruction through an AdaIN decoder forces the two artifact types to be separable, a contrastive loss organizes the feature space, and a mutual information term built on the Donsker–Varadhan lower bound aligns the domain-agnostic features with the content distribution. Sharpness-Aware Minimization perturbs weights toward higher loss before each gradient step, flattening the landscape; at inference only the domain-agnostic classification head is used.
What would settle it
Measure, on a held-out test set, the estimated mutual information between the learned domain-agnostic features and the vocoder identity, and the estimated mutual information between those features and the spoken content. The paper's mechanism predicts that after training the artifact features separate real from synthetic while remaining independent of which vocoder made them; if the features stay strongly tied to vocoder identity, or if editing only the mutual information loss leaves the cross-domain equal error rate essentially unchanged on a new-vocoder split, the central claim is falsified.
Extended reading notes
Core claim
The paper's central claim is that domain-agnostic artifact features, the traces left by speech synthesis that are common to many vocoders, can be extracted by disentanglement and then used directly for classification, giving a detector that generalizes across generators. It establishes that a content encoder and an artifact encoder with two classification heads (one for vocoder identity, one for real-versus-synthetic), a cross-reconstruction decoder, contrastive learning, and a mutual information term together produce these features, and that Sharpness-Aware Minimization then flattens the loss landscape and keeps the model out of sharp minima. The reported experiments on LibriSeVoc, ASVspoof2019, WaveFake, and FakeAVCeleb support the claim, with the largest gains on unseen vocoders.
Load-bearing premise
The central bet is that making the shared fake-voice features statistically more dependent on the spoken content turns them into universal fake-voice features; if that step actually does the opposite, the framework's claimed mechanism for cross-domain generalization is not supported.
Editorial extensions
If this is right
- Detectors trained with this pipeline keep working when new vocoder families appear, because the detector keys on shared artifacts rather than on the six generators it saw.
- The component ablation shows each piece, reconstruction, classification heads, contrastive loss, mutual information, and sharpness-aware minimization, contributes; removing the mutual information term alone costs about 6.85 equal error rate points on unseen FakeAVCeleb audio.
- Training-data diversity matters: using more vocoders in the training set monotonically lowers average equal error rate on both seen and unseen test vocoders.
- The approach also improves cross-dataset detection: trained on LibriSeVoc and tested on WaveFake, mean unseen-vocoder equal error rate drops from 34.06% to 23.38%.
Reading between the lines
- The mutual information objective as written maximizes dependence between content and artifact features, which is the opposite of disentangling them; the reported cross-domain gains could therefore come from sharpness-aware minimization, contrastive learning, or multi-task classification rather than from the claimed alignment. An isolated test would replace the mutual information term with a decorr
- If flatness is the real driver, then applying sharpness-aware minimization alone to simpler baselines should recover a meaningful fraction of the 7.59-point gain without any disentanglement; that is a direct test the paper does not report.
- The same recipe, shared artifact extraction plus flat-minimum optimization, is naturally transferable to other deepfake media where content and manipulation artifacts also mix; one could test it on cross-dataset face-swap or audio-visual deepfake benchmarks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a disentanglement framework for AI-synthesized voice detection aimed at improving generalization to unseen vocoders. The method trains a RawNet2-based encoder that separates content features, domain-specific artifact features, and domain-agnostic artifact features, guided by a multi-task classification loss, a contrastive loss, a reconstruction loss, a mutual information loss, and sharpness-aware minimization (SAM). The authors evaluate on LibriSeVoc, ASVspoof2019, WaveFake, and FakeAVCeleb, reporting improved equal error rates (EER) in both intra-domain and cross-domain settings compared to several baselines, and they release code. The central claim is that the domain-agnostic artifact features, made 'universally applicable' via the mutual information loss, and the flattened loss landscape from SAM jointly improve cross-domain generalization.
Significance. If the reported results hold, the framework would be a meaningful step toward cross-domain audio deepfake detection, which is a recognized weakness of current detectors. The paper's strengths include evaluation on multiple external benchmarks, a component-wise ablation study, and public code. However, the stated mechanism for the mutual information loss is technically questionable, and the empirical claims are based on single-run point estimates without uncertainty quantification. These issues are central because the proposed mechanism for domain-agnostic features is the paper's main novelty, and the numerical gains are the main evidence for it. With corrections to the MI formulation and stronger statistical validation, the work could be a solid contribution; in its current form, the supporting evidence is not fully convincing.
major comments (4)
- [Mutual Information Loss and Eq. (1)] The sign and interpretation of the mutual information loss are internally inconsistent. The paper states that maximizing MI(c; a^g) 'aligns domain-agnostic features with the content feature distribution,' but Eq. (1) subtracts λ4 L_MI from the total loss, so gradient descent on L maximizes the Donsker-Varadhan lower bound L_MI. Maximizing mutual information increases statistical dependence between c and a^g, which is the opposite of disentangling content from artifact features. The DV lower bound does not match marginal distributions; it amplifies per-sample dependence, as seen in Algorithm 2 where E_joint pulls c_i and a^g_i together for the same input. This undermines the claimed mechanism by which a^g becomes 'universally applicable' and vocoder-agnostic. The authors should either minimize MI(c; a^g) to enforce independence, or provide a different, correct justification for why maximizing MI yields domain-agnostic features, along with direct evidence about what a^g encodes. The ablation gain of VD over VC in Table 4 may then be attributable to an unintended regularizer rather than the described alignment, so the central claim is not yet supported.
- [Algorithm 2 (MI pseudocode)] The PyTorch-style pseudocode in Algorithm 2 is dimensionally unclear and not reproducible as written. The comment says c is 'B x n0 x dim' and a is 'B x n1 x dim', but u = torch.mm(a, c.t()) operates on 2D tensors; the subsequent reshape to 'B x B x n0 x n1' and the use of mean(2) are not consistent with typical 3D feature tensors. It is also unclear what the 'joint' and 'margin' scores represent after masking, and whether the average is over the batch, the feature dimensions, or both. Because the mutual information loss is a load-bearing component of the framework, the implementation must be specified unambiguously so that the reported results can be reproduced and the behavior of the loss term can be independently checked.
- [Experimental Results, Tables 1-4] All reported EER numbers appear to be single-run point estimates with no error bars, confidence intervals, or significance tests for the improvements over baselines. The abstract highlights improvements of 5.12% and 7.59% in EER, but without run-to-run variance it is impossible to assess whether these differences are statistically meaningful, particularly for a model with several interacting loss terms and hyperparameters. The authors should report mean and standard deviation over multiple random seeds, or at least provide significance tests for the main comparisons against Sun et al. and RawNet2. This is load-bearing because the central claim of outperforming state-of-the-art methods rests entirely on these point estimates.
- [Ablation Study, Table 4 and Figure 4] The paper claims that the mutual information module 'greatly improves performance in cross-domain evaluation,' but the evidence is mixed. In Table 4, adding MI (VD vs. VC) improves unseen ASP EER from 26.83 to 23.23 and unseen FakeAVCeleb from 27.64 to 20.79, yet worsens seen WF EER from 24.72 to 25.62. Given the MI sign issue described above, the authors should show what the learned a^g actually encodes, for example by measuring vocoder classification accuracy from a^g or by quantifying how much content information remains in a^g. The qualitative UMAP in Figure 4 is suggestive but not quantitative. Without direct evidence that a^g is both vocoder-invariant and content-independent, the claimed disentanglement effect is not established.
minor comments (4)
- [Throughout] There are many typographical and spacing errors, such as 'conetent' in Algorithm 2, 'V oice' in several headings, and inconsistent use of 'Vo ice' and 'voice.' These should be corrected in a revised version.
- [Eq. (1) and Section Mutual Information Loss] The notation for the mutual information loss is confusing: Eq. (1) subtracts L_MI, but the text refers to it as a 'loss' and claims it 'aligns' distributions. It would be clearer to call it a regularization term with an explicit sign convention, and to define whether the reported hyperparameter λ4 controls the magnitude of maximization or minimization.
- [Implementation Details] The hyperparameters for the method are given (λ1=0.1, λ2=0.3, λ3=0.05, λ4=0.03, b=3, γ=0.07), but the corresponding hyperparameter tuning procedure is not described. It is also unclear how the baselines were tuned for their own hyperparameters; a brief explanation would help fairness of comparison.
- [Appendix, Algorithm 1] Algorithm 1's update step computes ϵ* from ∇θL and then updates θ using the gradient at θ+ϵ*, but the line 'Update θ: θl+1 ← θl − β∇θL|θl+ϵ*' overloads ∇θL; this should be written more explicitly to avoid ambiguity.
Circularity Check
No significant circularity: the central generalization claims are evaluated against external baselines and held-out vocoder datasets; the only self-citation supplies a contrastive-learning technique, not the target result.
full rationale
The claimed derivation chain is self-contained and externally checked. The disentanglement objective (Eq. (1)) combines a classification loss, a contrastive loss, a reconstruction loss, and a mutual-information term; all are defined on the training data and labels, and their contributions are isolated in the ablation study (Table 4). The mutual-information term uses the Donsker-Varadhan lower bound from Belghazi et al. (2018), an external method, and its effect is measured, not assumed. The optimization-level component is SAM (Foret et al. 2020), an external technique adapted to this task. The paper's headline improvements (5.12% intra-domain and 7.59% cross-domain EER reductions) are comparisons against external baselines (LCNN, RawNet2, WavLM, XLS-R, Sun et al.) on external benchmarks (LibriSeVoc, ASVspoof2019, WaveFake, FakeAVCeleb), including held-out vocoders. The only self-citation is Lin et al. 2024, cited as 'Inspired by' for the contrastive loss; that prior work contributes a technique, not the paper's predicted outcome, and no uniqueness theorem or fitted parameter is imported from it. A reviewer-level concern that maximizing MI(c; ag) may not produce the stated 'alignment' with content features is a mechanistic correctness issue, not a circularity: the loss is not defined in terms of the test metric or the target result. Accordingly, no load-bearing step reduces to its own input.
Assumptions & free parameters
free parameters (6)
- lambda_1 =
0.1
- lambda_2 =
0.3
- lambda_3 =
0.05
- lambda_4 =
0.03
- contrastive margin b =
3
- SAM perturbation gamma =
0.07
assumptions (5)
- standard math Donsker-Varadhan representation provides a valid lower bound on mutual information that can be estimated and maximized with a neural network.
- domain assumption Vocoder identity is a sufficient domain label; all generalization-relevant variation is captured by the vocoder type in the training set.
- ad hoc to paper Maximizing MI(c; ag) aligns domain-agnostic features with the content feature distribution and makes them universally applicable.
- standard math SAM's first-order approximation and dual norm solution for the adversarial perturbation are valid for flattening the loss landscape.
- domain assumption Common domain-agnostic artifacts exist across different vocoders and are learnable from a finite set of seen vocoders.
invented entities (3)
-
domain-agnostic artifact feature space (a^g)
-
domain-specific artifact feature space (a^s)
-
content feature distribution as a benchmark
Cite this review
Pith. "Pith review of Improving Generalization for AI-Synthesized Voice Detection." pith.science (2026). https://pith.science/paper/ZFG6575K
@misc{pith2026241219279,
author = {Pith},
title = {Pith review of: Improving Generalization for AI-Synthesized Voice Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZFG6575K}},
note = {Machine review of arXiv:2412.19279}
}
read the original abstract
AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Aloufi, R.; Haddadi, H.; and Boyle, D. 2020. Privacy-preserving voice analysis via disentangled representations. In Proceedings of the 2020 ACM SIGSAC Conference on Cloud Computing Security Workshop, 1--14
work page 2020
-
[2]
ASVspoof. 2023. Automatic Speaker Verification and Spoofing Countermeasures Challenge. In https://www.asvspoof.org/
work page 2023
-
[3]
Babu, A.; Wang, C.; Tjandra, A.; Lakhotia, K.; Xu, Q.; Goyal, N.; Singh, K.; von Platen, P.; Saraf, Y.; Pino, J.; et al. 2021. XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale. arXiv e-prints, arXiv--2111
work page 2021
-
[4]
Barrington, S.; Barua, R.; Koorma, G.; and Farid, H. 2023. Single and Multi-Speaker Cloned Voice Detection: From Perceptual to Learned Features. IEEE International Workshop on Information Forensics and Security
work page 2023
-
[5]
I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D
Belghazi, M. I.; Baratin, A.; Rajeshwar, S.; Ozair, S.; Bengio, Y.; Courville, A.; and Hjelm, D. 2018. Mutual information neural estimation. In International conference on machine learning, 531--540. PMLR
work page 2018
-
[6]
Champion, P.; Jouvet, D.; and Larcher, A. 2022. Are disentangled representations all you need to build speaker anonymization systems?
work page 2022
-
[7]
Chen, M.; Zhou, Y.; Huang, H.; and Hain, T. 2022 a . Efficient Non-Autoregressive GAN Voice Conversion using VQWav2vec Features and Dynamic Convolution. arXiv:2203.17172
work page Pith review arXiv 2022
-
[8]
Chen, N.; Zhang, Y.; Zen, H.; Weiss, R. J.; Norouzi, M.; and Chan, W. 2020 a . WaveGrad: Estimating Gradients for Waveform Generation. In International Conference on Learning Representations
work page 2020
Show all 80 references
-
[9]
Chen, S.; Wang, C.; Chen, Z.; Wu, Y.; Liu, S.; Chen, Z.; Li, J.; Kanda, N.; Yoshioka, T.; Xiao, X.; et al. 2022 b . Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 16(6): 1505--1518
2022
-
[10]
Chen, T.; Kumar, A.; Nagarsheth, P.; Sivaraman, G.; and Khoury, E. 2020 b . Generalization of Audio Deepfake Detection . In Proc. The Speaker and Language Recognition Workshop (Odyssey 2020), 132--137
2020
-
[11]
Chettri, B.; Stoller, D.; Morfi, V.; Ram \'i rez, M. A. M.; Benetos, E.; and Sturm, B. L. 2019. Ensemble Models for Spoofing Detection in Automatic Speaker Verification. In Interspeech
2019
-
[12]
Ding, S.; Zhang, Y.; and Duan, Z. 2023. SAMO: Speaker Attractor Multi-Center One-Class Learning For Voice Anti-Spoofing. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5
2023
-
[13]
Donahue, J.; Dieleman, S.; Binkowski, M.; Elsen, E.; and Simonyan, K. 2021. End-to-End Adversarial Text-to-Speech. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net
2021
-
[14]
D.; and Varadhan, S
Donsker, M. D.; and Varadhan, S. S. 1983. Asymptotic evaluation of certain Markov process expectations for large time. IV. Communications on pure and applied mathematics, 36(2): 183--212
1983
-
[15]
Forbes. 2019. A Voice Deepfake Was Used To Scam A CEO Out Of \ 243,000. In https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice-deepfake-was-used-to-scam-a-ceo-out-of-243000/?sh=432e93d12241
2019
-
[16]
Foret, P.; Kleiner, A.; Mobahi, H.; and Neyshabur, B. 2020. Sharpness-aware Minimization for Efficiently Improving Generalization. In International Conference on Learning Representations
2020
-
[17]
Frank, J.; and Sch \"o nherr, L. 2021. WaveFake: A Data Set to Facilitate Audio Deepfake Detection. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
2021
-
[18]
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision, 1026--1034
2015
-
[19]
He, R.; Wu, X.; Sun, Z.; and Tan, T. 2017. Learning invariant deep representation for nir-vis face recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31
2017
-
[20]
He, R.; Wu, X.; Sun, Z.; and Tan, T. 2018. Wasserstein CNN: Learning invariant features for NIR-VIS face recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(7): 1761--1773
2018
-
[21]
D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y
Hjelm, R. D.; Fedorov, A.; Lavoie-Marchildon, S.; Grewal, K.; Bachman, P.; Trischler, A.; and Bengio, Y. 2018. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670
2018 arXiv
-
[22]
H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A
Hsu, W.-N.; Bolte, B.; Tsai, Y.-H. H.; Lakhotia, K.; Salakhutdinov, R.; and Mohamed, A. 2021. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE/ACM Trans. Audio, Speech and Lang. Proc., 29: 3451–3460
2021
-
[23]
Huang, X.; and Belongie, S. 2017. Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision, 1501--1510
2017
-
[24]
Jia, Y.; Zhang, Y.; Weiss, R.; Wang, Q.; Shen, J.; Ren, F.; Nguyen, P.; Pang, R.; Lopez Moreno, I.; Wu, Y.; et al. 2018. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. Advances in neural information processing systems, 31
2018
-
[25]
Joyce, J. M. 2011. Kullback-leibler divergence. In International Encyclopedia of Statistical Science, 720--722
2011
-
[26]
Kalchbrenner, N.; Elsen, E.; Simonyan, K.; Noury, S.; Casagrande, N.; Lockhart, E.; Stimberg, F.; Oord, A.; Dieleman, S.; and Kavukcuoglu, K. 2018. Efficient neural audio synthesis. In International Conference on Machine Learning, 2410--2419. PMLR
2018
-
[27]
Khalid, H.; Tariq, S.; Kim, M.; and Woo, S. S. 2021. FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
2021
-
[28]
Kim, J.; Kong, J.; and Son, J. 2021. Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learn...
2021
-
[29]
Kim, S.; Lee, S.-G.; Song, J.; Kim, J.; and Yoon, S. 2019. F lo W ave N et : A Generative Flow for Raw Audio. In Chaudhuri, K.; and Salakhutdinov, R., eds., Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Resea...
2019
-
[30]
P.; and Ba, J
Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980
2014 arXiv
-
[31]
B.; and Atwal, G
Kinney, J. B.; and Atwal, G. S. 2014. Equitability, mutual information, and the maximal information coefficient. Proceedings of the National Academy of Sciences, 111(9): 3354--3359
2014
-
[32]
Kong, J.; Kim, J.; and Bae, J. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33: 17022--17033
2020
-
[33]
Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; and Catanzaro, B. 2020. DiffWave: A Versatile Diffusion Model for Audio Synthesis. In International Conference on Learning Representations
2020
-
[34]
Z.; Sotelo, J.; De Brebisson, A.; Bengio, Y.; and Courville, A
Kumar, K.; Kumar, R.; De Boissiere, T.; Gestin, L.; Teoh, W. Z.; Sotelo, J.; De Brebisson, A.; Bengio, Y.; and Courville, A. C. 2019. Melgan: Generative adversarial networks for conditional waveform synthesis. Advances in neural information processing systems, 32
2019
-
[35]
Lavrentyeva, G.; Tseren, A.; Volkova, M.; Gorlanov, A.; Kozlov, A.; and Novoselov, S. 2019. STC antispoofing systems for the AsVspoof2019 challenge. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 1033--1037
2019
-
[36]
Lee, S.-H.; Kim, J.-H.; Lee, K.-E.; and Lee, S.-W. 2022. FRE-GAN 2: Fast and Efficient Frequency-Consistent Audio Synthesis. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6192--6196
2022
-
[37]
A.; Zare, A
Li, Y. A.; Zare, A. A.; and Mesgarani, N. 2021. StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion. In Interspeech
2021
-
[38]
Lian, J.; Zhang, C.; and Yu, D. 2022. Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6572--6576
2022
-
[39]
Lin, L.; He, X.; Ju, Y.; Wang, X.; Ding, F.; and Hu, S. 2024. Preserving fairness generalization in deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16815--16825
2024
-
[40]
Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023 a . A udio LDM : Text-to-Audio Generation with Latent Diffusion Models. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., Proceedings of the 4...
2023
-
[41]
A.; Zhang, H.; and Dang, J
Liu, X.; Liu, M.; Wang, L.; Lee, K. A.; Zhang, H.; and Dang, J. 2023 b . Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5
2023
-
[42]
Long, Z.; Zheng, Y.; Yu, M.; and Xin, J. 2022. Enhancing Zero-Shot Many to Many Voice Conversion via Self-Attention VAE with Structurally Regularized Layers. In 2022 5th International Conference on Artificial Intelligence for Industries (AI4I), 59--63
2022
-
[43]
Lorenzo-Trueba, J.; Drugman, T.; Latorre, J.; Merritt, T.; Putrycz, B.; Barra-Chicote, R.; Moinet, A.; and Aggarwal, V. 2019. Towards Achieving Robust Universal Neural Vocoding . In Proc. Interspeech 2019, 181--185
2019
-
[44]
Luong, M.; and Tran, V. A. 2021. Many-to-many voice conversion based feature disentanglement using variational autoencoder. In Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association
2021
-
[45]
McInnes , L.; Healy , J.; and Melville , J. 2018. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction . ArXiv e-prints
2018
-
[46]
Miao, C.; Shuang, L.; Liu, Z.; Minchuan, C.; Ma, J.; Wang, S.; and Xiao, J. 2021. EfficientTTS: An Efficient and High-Quality Text-to-Speech Architecture. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Pro...
2021
-
[47]
Müller, N.; Czempin, P.; Diekmann, F.; Froghyar, A.; and Böttinger, K. 2022. Does Audio Deepfake Detection Generalize? In Proc. Interspeech 2022, 2783--2787
2022
-
[48]
Oord, A. v. d.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499
2016 arXiv
-
[49]
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748
2018 arXiv
-
[50]
Paul, D.; Pantazis, Y.; and Stylianou, Y. 2020. Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions . In Proc. Interspeech 2020, 235--239
2020
-
[51]
Peng, K.; Ping, W.; Song, Z.; and Zhao, K. 2020. Non-Autoregressive Neural Text-to-Speech. In III, H. D.; and Singh, A., eds., Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, 7586--7598. PMLR
2020
-
[52]
Ping, W.; Peng, K.; and Chen, J. 2019. ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net
2019
-
[53]
Prenger, R.; Valle, R.; and Catanzaro, B. 2019. Waveglow: A Flow-based Generative Network for Speech Synthesis. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 3617--3621
2019
-
[54]
Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T. 2021. FastSpeech 2: Fast and High-Quality End-to-End Text to Speech. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net
2021
-
[55]
Salvi, D.; Bestagini, P.; and Tubaro, S. 2023. Reliability Estimation for Synthetic Speech Detection. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5
2023
-
[56]
Schneider, S.; Baevski, A.; Collobert, R.; and Auli, M. 2019. wav2vec: Unsupervised Pre-training for Speech Recognition. In Interspeech
2019
-
[57]
F.; Kastner, K.; Courville, A
Sotelo, J.; Mehri, S.; Kumar, K.; Santos, J. F.; Kastner, K.; Courville, A. C.; and Bengio, Y. 2017. Char2Wav: End-to-End Speech Synthesis. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings . O...
2017
-
[58]
Sun, C.; Jia, S.; Hou, S.; and Lyu, S. 2023. AI-Synthesized Voice Detection Using Neural Vocoder Artifacts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshop, 904--912
2023
-
[59]
Tak, H.; Patino, J.; Nautsch, A.; Evans, N. W. D.; and Todisco, M. 2020. An explainability study of the constant Q cepstral coefficient spoofing countermeasure for automatic speaker verification. In The Speaker and Language Recognition Workshop
2020
-
[60]
Tak, H.; Patino, J.; Todisco, M.; Nautsch, A.; Evans, N.; and Larcher, A. 2021. End-to-end anti-spoofing with RawNet2 . In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6369--6373. IEEE
2021
-
[61]
K.; and Liu, T
Tan, X.; Qin, T.; Soong, F. K.; and Liu, T. 2021. A Survey on Neural Speech Synthesis. CoRR, abs/2106.15561
2021 arXiv
-
[62]
A.; Sahidullah, M.; Evans, N.; Kinnunen, T.; and Yamagishi, J
Todisco, M.; Delgado, H.; Lee, K. A.; Sahidullah, M.; Evans, N.; Kinnunen, T.; and Yamagishi, J. 2018. Integrated Presentation Attack Detection and Automatic Speaker Verification: Common Features and Gaussian Back-end Fusion . In Proc. Interspeech 2018, 77--81
2018
-
[63]
Veaux, C.; Yamagishi, J.; MacDonald, K.; et al. 2016. Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
2016
-
[64]
Y.; Zhang, S.; and Chen, X
Wang, C.; Yi, J.; Tao, J.; Zhang, C. Y.; Zhang, S.; and Chen, X. 2023. Detection of Cross-Dataset Fake Audio Based on Prosodic and Pronunciation Features . In Proc. INTERSPEECH 2023, 3844--3848
2023
-
[65]
Wang, X.; Chen, H.; Tang, S.; Wu, Z.; and Zhu, W. 2022. Disentangled representation learning. arXiv preprint arXiv:2211.11695
2022 arXiv
-
[66]
Wang, X.; and Yamagishi, J. 2023. Can large-scale vocoded spoofed data improve speech spoofing countermeasure with a self-supervised front end? arXiv preprint arXiv:2309.06014
2023 arXiv
-
[67]
J.; Skerry-Ryan, R.; Battenberg, E.; Mariooryad, S.; and Kingma, D
Weiss, R. J.; Skerry-Ryan, R.; Battenberg, E.; Mariooryad, S.; and Kingma, D. P. 2021. Wave-Tacotron: Spectrogram-Free End-to-End Text-to-Speech Synthesis. In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5679--5683
2021
-
[68]
S.; Lee, B.-J.; jin Yu, H.; and Evans, N
weon Jung, J.; Heo, H.-S.; Tak, H.; jin Shim, H.; Chung, J. S.; Lee, B.-J.; jin Yu, H.; and Evans, N. W. D. 2021. AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and S...
2021
-
[69]
Xie, Y.; Cheng, H.; Wang, Y.; and Ye, L. 2023. Domain Generalization Via Aggregation and Separation for Audio Deepfake Detection. IEEE Transactions on Information Forensics and Security
2023
-
[70]
Yadav, A. K. S.; Bhagtani, K.; Xiang, Z.; Bestagini, P.; Tubaro, S.; and Delp, E. J. 2023. Dsvae: Interpretable disentangled representation for synthetic speech detection. arXiv preprint arXiv:2304.03323
2023 arXiv
-
[71]
A.; Kinnunen, T.; Evans, N.; et al
Yamagishi, J.; Wang, X.; Todisco, M.; Sahidullah, M.; Patino, J.; Nautsch, A.; Liu, X.; Lee, K. A.; Kinnunen, T.; Evans, N.; et al. 2021. ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection. arXiv preprint arXiv:2109.00537
2021 arXiv
-
[72]
Yamamoto, R.; Song, E.; and Kim, J.-M. 2020. Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 61...
2020
-
[73]
Yan, Z.; Zhang, Y.; Fan, Y.; and Wu, B. 2023. Ucf: Uncovering common features for generalizable deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 22412--22423
2023
-
[74]
Yu, C.; Lu, H.; Hu, N.; Yu, M.; Weng, C.; Xu, K.; Liu, P.; Tuo, D.; Kang, S.; Lei, G.; Su, D.; and Yu, D. 2020. DurIAN: Duration Informed Attention Network for Speech Synthesis . In Proc. Interspeech 2020, 2027--2031
2020
-
[75]
Zhai, B.; Gao, T.; Xue, F.; Rothchild, D.; Wu, B.; Gonzalez, J.; and Keutzer, K. 2020. SqueezeWave: Extremely Lightweight Vocoders for On-device Speech Synthesis. ArXiv, abs/2001.05685
2020 arXiv
-
[76]
Zhang, X.; Yi, J.; Tao, J.; Wang, C.; and Zhang, C. Y. 2023. Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection. In International conference on machine learning. PMLR
2023
-
[77]
Zhang, Y.; Wang, W.; and Zhang, P. 2021. The Effect of Silence and Dual-Band Fusion in Anti-Spoofing System . In Proc. Interspeech 2021, 4279--4283
2021
-
[78]
Zhang, Y.-J.; Pan, S.; He, L.; and Ling, Z.-H. 2019. Learning latent representations for style control and transfer in end-to-end speech synthesis. In IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE
2019
-
[79]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[80]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.