REVIEW 3 major objections 6 minor 80 references
Gender bias in audio deepfake detectors is set by who you train on, and no threshold fix can close the core error gap.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 14:47 UTC pith:5G6US73P
load-bearing objection Controlled composition experiment that cleanly shows bias direction tracks the minority gender and that Oracle calibration cannot move the EER gap; speaker-overlap is a real but non-load-bearing limit. the 3 major comments →
What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Training gender composition strongly predicts which gender is disadvantaged at test time, and the Equal Error Rate gap between genders is invariant to post-hoc threshold calibration—including Oracle calibration with full evaluation labels—remaining fixed at 1.317 pp across nine compositions. Feature type modulates magnitude: WavLM-Base+ produces gaps 3.0–4.3× larger than LogSpectrogram on identical training sets, and balanced data nearly eliminates LogSpectrogram bias while leaving WavLM bias largely intact.
What carries the argument
Nine controlled gender-composition training sets (female-only through male-only, plus partial mixes) on custom per-attack ASVspoof5 splits, paired with six threshold strategies including Oracle, used to separate composition effects from post-hoc operating-point choice and to show that EER gap tracks score-distribution disparity rather than threshold placement.
Load-bearing premise
The custom 70/10/20 splits are treated as cleanly isolating gender-composition effects even though they are not guaranteed to keep speakers out of both train and test, so some of the measured bias may reflect speaker reuse rather than composition alone.
What would settle it
Retrain the same attack-specific models on speaker-disjoint gender-composition splits with known speaker IDs and check whether the underrepresented-gender disadvantage and the fixed 1.317 pp EER gap under Oracle calibration still appear at the same magnitude.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies gender bias in audio deepfake detection on ASVspoof5 under a controlled custom 70/10/20 per-attack, per-gender re-split. It trains 384 attack-specific ResNet18 models (288 WavLM-Base+ across nine gender-composition configurations; 96 LogSpectrogram under three baselines) and evaluates six group-fairness metrics plus six post-hoc threshold calibrations, including an Oracle with evaluation labels. Main claims: (i) training gender composition strongly predicts bias direction—the underrepresented gender is worse at test time in 31/32 attacks; (ii) WavLM-Base+ yields EER gaps 3.0–4.3× larger than LogSpectrogram under matched training, and balanced training nearly removes LogSpec disparity but leaves WavLM largely biased; (iii) all six calibrations, including Oracle, leave mean ΔEER fixed at 1.317 pp, so distribution-level EER gaps require training-time fixes; (iv) attack type and metric choice change fairness rankings, so no single metric suffices.
Significance. If the results hold, this is a useful, large-scale controlled demonstration that gender fairness in spoofing countermeasures is primarily a data- and representation-level problem rather than a pure post-processing problem. Strengths include the systematic composition design (including the F+75%M vs M+25%F internal check), scale (384 models, 32 attacks), Poisson bootstrap with Benjamini–Hochberg correction, explicit multi-metric reporting, and practical deployment guidelines. The LogSpec vs WavLM contrast under identical composition is informative for the SSL-frontend literature. The work is incremental relative to the authors’ prior ASVspoof5 fairness papers [13], [24] but fills a clear gap by treating composition as the independent variable.
major comments (3)
- §III-F1–F3 and Table X / abstract: Per-gender EER is defined by finding, for each gender independently, the threshold where that gender’s FAR equals FRR. Under that definition, ΔEER is a property of the two score distributions (ROC shapes) and is invariant to any post-hoc choice of deployment thresholds, including Oracle. Reporting that all six strategies “leave the EER gap unchanged at 1.317 pp” is therefore largely definitional, not an empirical discovery. The substantive point—that thresholding cannot reshape group ROCs—is correct and worth making, but the abstract and §IV-C should reframe this as a definitional consequence of distribution-level EER and put empirical weight on the metric trade-offs (FPR gap, EOpD, SPD, etc.) and the composition→bias results, not on the numerical invariance of ΔEER under Oracle.
- Contribution 2 and §IV-B1 (Table VI): The claim that “combined training reduces the mean EER gap from 1.042 pp to 0.273 pp compared to within-attack balancing at a similar gender ratio” is not supported as stated. F+50%M (gap 1.042 pp) is not a similar gender ratio to Combined (100% F + 100% M). The nearer comparison is F+75%M (0.438 pp) or M+75%F (0.627 pp). Combined also uses more minority-gender data (and thus more total data) than any partial mix, so the extra drop to 0.273 pp confounds balance with sample size. The attribution to a special “population-level pooling” benefit beyond within-attack balancing should be revised or supported by a matched-N or matched-ratio control.
- §III-B and §V: Custom splits are stratified by attack and gender but not speaker-disjoint; evaluation speakers are expected to appear in training and overlap cannot be quantified. This does not overturn the directional composition→bias or the definitional ΔEER-invariance claims under fixed split construction, but it does limit interpretation of absolute EERs, cross-attack gaps, and any implication of speaker-independent fairness. Absolute gap magnitudes and cross-attack TED (already flagged as unstable) should be caveated more sharply in the abstract and conclusions so readers do not treat them as speaker-generalization results.
minor comments (6)
- §I and abstract: “we evaluated” should be “we evaluate” for tense consistency with the rest of the abstract.
- Table I / Fig. 1: Official ASVspoof5 partition labels (A01–A08 train, etc.) are shown alongside the custom re-split; a single sentence in the caption stating that official partitions are only the source pool would reduce confusion.
- §IV-B2: Cross-attack TED values (e.g., F: 0.115 → 54.785) are correctly called numerically unstable; consider moving them to an appendix or suppressing them in the main heatmap so they do not dominate Fig. 4.
- §III-G: N=1000 Poisson bootstrap resamples is a thin lower bound for claims at p≪0.001; stating this more prominently (as the authors already note for future work) would help readers calibrate confidence.
- Notation: SPD, EOpD, EOD, PPD, TED, FPRgap are defined clearly in §III-F2; adding a one-line “metric → deployment use-case” table early would help non-fairness readers navigate §IV-C1.
- References [13] and [24] are the authors’ own closely related work; ensure the novelty paragraph states explicitly what is new (composition as controlled IV; multi-composition calibration failure modes) versus reused metrics/protocol.
Circularity Check
Empirical composition-and-calibration study; no circular derivation chain
full rationale
The paper's load-bearing claims are experimental measurements: nine controlled gender-composition training sets, 384 attack-specific ResNet18 models (LogSpectrogram and WavLM-Base+), per-gender EER and six group-fairness metrics, and six post-hoc threshold strategies including Oracle. Bias direction tracking the underrepresented gender, the 3.0–4.3× larger WavLM gaps, and the invariance of ΔEER = 1.317 pp under every calibration (including Oracle) are read off the resulting score distributions and ROC curves; they are not obtained by fitting a parameter and re-labeling it as a prediction, nor by defining one quantity in terms of another. Self-citations [13] and [24] supply prior fixed-composition fairness numbers and the earlier balanced-training FPR-gap result that the present work extends; they are not used as uniqueness theorems or as the sole justification for the composition-variation or Oracle-EER findings. The known fact that a threshold cannot reshape a ROC (hence cannot change per-gender EER) is stated and then verified empirically, not smuggled in as a circular premise. Speaker non-disjointness of the custom splits is an external-validity limitation, not a circular step. No self-definitional, fitted-as-prediction, uniqueness-import, or ansatz-via-citation reductions appear.
Axiom & Free-Parameter Ledger
free parameters (3)
- learning_rate_and_weight_decay =
3e-5 / 1e-4
- practical_bias_threshold_0.5pp =
0.5 pp
- bootstrap_N_and_alpha =
N=1000, alpha=0.05
axioms (4)
- domain assumption Official ASVspoof5 binary gender labels correctly partition speakers into female and male for fairness evaluation.
- domain assumption Absolute EER gap is a valid primary indicator of distribution-level gender disparity that threshold calibration cannot change.
- ad hoc to paper Attack-specific models isolate attack-dependent bias without requiring speaker-independent generalization.
- standard math Poisson bootstrap on confusion-matrix cells adequately approximates utterance-level uncertainty for the reported p-values.
read the original abstract
Audio deepfake detection models determine whether speech is genuine or artificially generated, but high overall accuracy can mask substantial performance disparities across demographic groups. In this work, we investigate gender bias in audio deepfake detection using the ASVspoof5 dataset. We use ASVspoof5 under a controlled custom split designed to isolate gender-composition effects. We train attack-specific models on nine training sets with different gender compositions, ranging from female-only to male-only. We use a ResNet18 classifier with LogSpectrogram and WavLM-Base+ features, and we evaluated six post-hoc threshold calibration methods. Experimental results show that training data composition strongly predicts bias direction, with the underrepresented gender performing worse at test time. WavLM-Base+ features are shown to produce gender performance gaps 3.0 to 4.3 times larger than LogSpectrogram under identical training conditions, and balanced training is found to reduce LogSpectrogram bias but leave WavLM bias largely intact. Moreover, all six calibration strategies, including Oracle calibration with full test-set label access, leave the Equal Error Rate gap unchanged at 1.317 pp, confirming that threshold adjustment cannot correct underlying score distribution disparities. Overall, these findings suggest that gender fairness in audio deepfake detection must be addressed at training time, as post-hoc methods can only partially mitigate the resulting disparities
Figures
Reference graph
Works this paper leans on
-
[1]
Warning: Humans cannot reliably detect speech deepfakes,
K. T. Mai, S. Bray, T. Davies, and L. D. Griffin, “Warning: Humans cannot reliably detect speech deepfakes,”PLOS ONE, vol. 18, no. 8, p. e0285333, 2023
2023
-
[2]
Audio deepfake detection using deep learning,
O. A. Shaaban and R. Yildirim, “Audio deepfake detection using deep learning,”Engineering Reports, vol. 7, no. 3, p. e70087, 2025
2025
-
[3]
A comprehensive survey of deepfake generation and detection techniques in audio-visual media,
I. Khan, K. Khan, and A. Ahmad, “A comprehensive survey of deepfake generation and detection techniques in audio-visual media,”ICCK Journal of Image Analysis and Processing, vol. 1, no. 2, pp. 73–95, 2025
2025
-
[4]
Audio deepfake detection: A survey,
J. Yi, C. Wang, J. Tao, X. Zhang, C. Y . Zhang, and Y . Zhao, “Audio deepfake detection: A survey,”arXiv preprint arXiv:2308.14970, 2023
Pith/arXiv arXiv 2023
-
[5]
Vulnerabilities of audio-based biometric authentication systems against deepfake speech synthesis,
M. Hong, D. Jiang, Z. Xie, W. Zhao, G. Wang, and C. J. Zhang, “Vulnerabilities of audio-based biometric authentication systems against deepfake speech synthesis,”arXiv preprint arXiv:2601.02914, 2026
arXiv 2026
-
[6]
Audio deepfake detection: What has been achieved and what lies ahead,
B. Zhang, H. Cui, V . Nguyen, and M. Whitty, “Audio deepfake detection: What has been achieved and what lies ahead,”Sensors, vol. 25, no. 7, p. 1989, 2025
1989
-
[7]
ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,
X. Wang, H. Delgado, H. Tak, J. W. Jung, H. J. Shim, M. Todisco, et al., “ASVspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,”arXiv preprint arXiv:2408.08739, 2024
Pith/arXiv arXiv 2024
-
[8]
End-to-end anti-spoofing with RawNet2,
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with RawNet2,” inProc. ICASSP 2021 – IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 6369–6373
2021
-
[9]
AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,
J. W. Jung, H. S. Heo, H. Tak, H. J. Shim, J. S. Chung, B. J. Lee, et al., “AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” inProc. ICASSP 2022 – IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 6367–6371
2022
-
[10]
Bias in data-driven artificial intelligence systems: An introductory survey,
E. Ntoutsi, P. Fafalios, U. Gadiraju, V . Iosifidis, W. Nejdl, M. E. Vidal, et al., “Bias in data-driven artificial intelligence systems: An introductory survey,”Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, vol. 10, no. 3, p. e1356, 2020
2020
-
[11]
Fair voice biometrics: Impact of demographic imbalance on group fairness in speaker recogni- tion,
G. Fenu, M. Marras, G. Medda, and G. Meloni, “Fair voice biometrics: Impact of demographic imbalance on group fairness in speaker recogni- tion,” inProc. Interspeech, 2021, pp. 1892–1896
2021
-
[12]
Real-time detection of AI-generated speech for deepfake voice conversion,
J. J. Bird and A. Lotfi, “Real-time detection of AI-generated speech for deepfake voice conversion,”arXiv preprint arXiv:2308.12734, 2023
Pith/arXiv arXiv 2023
-
[13]
Gender fairness in audio deepfake detection: Performance and disparity analysis,
A. Fursule, S. Kshirsagar, and A. R. Avila, “Gender fairness in audio deepfake detection: Performance and disparity analysis,” inProc. 2026 IEEE Conference on Artificial Intelligence (CAI), 2026, pp. 2116–2121
2026
-
[14]
Fairness without demographics in repeated loss minimization,
T. Hashimoto, M. Srivastava, H. Namkoong, and P. Liang, “Fairness without demographics in repeated loss minimization,” inProc. Interna- tional Conference on Machine Learning, 2018, pp. 1929–1938
2018
-
[15]
AFSS: Artifact-focused self-synthesis for mitigat- ing bias in audio deepfake detection,
H. S. Nguyen-Le, H. C. Nguyen-Thanh, N. A. Le-Khac, D. T. Nguyen, and H. H. Nguyen-Le, “AFSS: Artifact-focused self-synthesis for mitigat- ing bias in audio deepfake detection,”arXiv preprint arXiv:2603.26856, 2026
arXiv 2026
-
[16]
GBDF: Gender balanced deepfake dataset towards fair deepfake detection,
A. V . Nadimpalli and A. Rattani, “GBDF: Gender balanced deepfake dataset towards fair deepfake detection,” inProc. International Confer- ence on Pattern Recognition, 2022, pp. 320–337
2022
-
[17]
A survey on bias and fairness in machine learning,
N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,”ACM Computing Surveys, vol. 54, no. 6, pp. 1–35, 2021
2021
-
[18]
Improving fairness in deepfake detection,
Y . Ju, S. Hu, S. Jia, G. H. Chen, and S. Lyu, “Improving fairness in deepfake detection,” inProc. IEEE/CVF Winter Conference on Applica- tions of Computer Vision, 2024, pp. 4655–4665
2024
-
[19]
Preserving fairness generalization in deepfake detection,
L. Lin, X. He, Y . Ju, X. Wang, F. Ding, and S. Hu, “Preserving fairness generalization in deepfake detection,” inProc. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 16815–16825
2024
-
[20]
FairSSD: Understanding bias in synthetic speech detectors,
A. K. S. Yadav, K. Bhagtani, D. Salvi, P. Bestagini, and E. J. Delp, “FairSSD: Understanding bias in synthetic speech detectors,” inProc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 4418–4428
2024
-
[21]
Equality of opportunity in supervised learning,
M. Hardt, E. Price, and N. Srebro, “Equality of opportunity in supervised learning,” inAdvances in Neural Information Processing Systems, vol. 29, 2016
2016
-
[22]
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments,
A. Chouldechova, “Fair prediction with disparate impact: A study of bias in recidivism prediction instruments,”Big Data, vol. 5, no. 2, pp. 153– 163, 2017
2017
-
[23]
Phonetic analysis of real and synthetic speech using HuBERT embeddings: Perspectives for deepfake detection,
D. E. Temmar, A. Hamadene, V . Nallaguntla, A. Fursule, M. S. Allili, S. Kshirsagar, and A. R. Avila, “Phonetic analysis of real and synthetic speech using HuBERT embeddings: Perspectives for deepfake detection,” inProc. 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), 2025, pp. 86–91
2025
-
[24]
Towards trustworthy audio deepfake detection: A systematic framework for diagnosing and mitigating gender bias,
A. Fursule, S. Kshirsagar, and A. R. Avila, “Towards trustworthy audio deepfake detection: A systematic framework for diagnosing and mitigating gender bias,” inProc. IEEE International Conference on Systems, Man, and Cybernetics (SMC), 2026
2026
-
[25]
PhonemeDF: A synthetic speech dataset for audio deepfake detection and naturalness evaluation,
V . Nallaguntla, A. Fursule, S. Kshirsagar, and A. R. Avila, “PhonemeDF: A synthetic speech dataset for audio deepfake detection and naturalness evaluation,”arXiv preprint arXiv:2603.15037, 2026
arXiv 2026
-
[26]
Investigating the impact of speech enhancement on audio deepfake detection in noisy environments,
S. Kshirsagar and A. R. Avila, “Investigating the impact of speech enhancement on audio deepfake detection in noisy environments,”arXiv preprint arXiv:2603.14767, 2026
arXiv 2026
-
[27]
An examination of fairness of AI models for deepfake detection,
L. Trinh and Y . Liu, “An examination of fairness of AI models for deepfake detection,”arXiv preprint arXiv:2105.00558, 2021
Pith/arXiv arXiv 2021
-
[28]
Analyzing fairness in deepfake detection with massively annotated databases,
Y . Xu, P. Terh ¨orst, M. Pedersen, and K. Raja, “Analyzing fairness in deepfake detection with massively annotated databases,”IEEE Transac- tions on Technology and Society, vol. 5, no. 1, pp. 93–106, 2024
2024
-
[29]
Bias in automated speaker recognition,
W. T. Hutiri and A. Y . Ding, “Bias in automated speaker recognition,” inProc. ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2022, pp. 230–247
2022
-
[30]
SCDF: A Speaker Char- acteristics DeepFake Speech Dataset for Bias Analysis,
V . Stan ˇek, K. Srna, A. Firc, and K. Malinka, “SCDF: A Speaker Char- acteristics DeepFake Speech Dataset for Bias Analysis,”arXiv preprint arXiv:2508.07944, 2025
Pith/arXiv arXiv 2025
-
[31]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProc. IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778
2016
-
[32]
Easy, interpretable, effective: openSMILE for voice deepfake detection,
O. Pascu, D. Oneat ¸ ˘a, H. Cucu, and N. M ¨uller, “Easy, interpretable, effective: openSMILE for voice deepfake detection,” inProc. ICASSP 2025 – IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–5
2025
-
[33]
Natural- Speech: End-to-end text-to-speech synthesis with human-level quality,
X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y . Liu, et al., “Natural- Speech: End-to-end text-to-speech synthesis with human-level quality,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 6, pp. 4234–4245, 2024
2024
-
[34]
ASVspoof: The automatic speaker verification spoofing and countermeasures challenge,
Z. Wu, J. Yamagishi, T. Kinnunen, C. Hanilc ¸i, M. Sahidullah, A. Sizov, et al., “ASVspoof: The automatic speaker verification spoofing and countermeasures challenge,”IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 4, pp. 588–604, 2017
2017
-
[35]
ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection,
J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, et al., “ASVspoof 2021: Accelerating progress in spoofed and deepfake speech detection,”arXiv preprint arXiv:2109.00537, 2021
Pith/arXiv arXiv 2021
-
[36]
ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild,
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, et al., “ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2507–2522, 2023
2021
-
[37]
R. Peri, K. Somandepalli, and S. Narayanan, “To train or not to train adversarially: A study of bias mitigation strategies for speaker recognition,”arXiv preprint arXiv:2203.09122, 2022
Pith/arXiv arXiv 2022
-
[38]
WavLM: Large-scale self-supervised pre-training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, et al., “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[39]
Controlling the false discovery rate: A practical and powerful approach to multiple testing,
Y . Benjamini and Y . Hochberg, “Controlling the false discovery rate: A practical and powerful approach to multiple testing,”Journal of the Royal Statistical Society: Series B (Methodological), vol. 57, no. 1, pp. 289–300, 1995
1995
-
[40]
An intervention-based framework for shortcut diagnosis in spoofing countermeasures,
S. Rubio, P. Bello, D. Ribas, A. Miguel, E. Lleida, and A. Ortega, “An intervention-based framework for shortcut diagnosis in spoofing countermeasures,” inProc. Odyssey 2026: The Speaker and Language Recognition Workshop, 2026, pp. 1–8
2026
-
[41]
Can SSL frontend generalize to all-type audio spoofing?
A. Das, Y . El Kheir, F. R. Guttierez, T. Polzehl, and S. M¨oller, “Can SSL frontend generalize to all-type audio spoofing?” inProc. Odyssey 2026: The Speaker and Language Recognition Workshop, 2026, pp. 277–283
2026
-
[42]
Decoupled weight decay regularization,
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017
Pith/arXiv arXiv 2017
-
[43]
Creating non-parametric bootstrap samples using Poisson frequencies,
J. A. Hanley and B. MacGibbon, “Creating non-parametric bootstrap samples using Poisson frequencies,”Computer Methods and Programs in Biomedicine, vol. 83, no. 1, pp. 57–62, 2006
2006
-
[44]
The control of the false discovery rate in multiple testing under dependency,
Y . Benjamini and D. Yekutieli, “The control of the false discovery rate in multiple testing under dependency,”The Annals of Statistics, vol. 29, no. 4, pp. 1165–1188, 2001. 15
2001
-
[45]
Efron and R
B. Efron and R. J. Tibshirani,An Introduction to the Bootstrap. New York, NY , USA: Chapman and Hall, 1993
1993
-
[46]
Creating new language and voice components for the updated MaryTTS text-to-speech synthesis platform,
I. Steiner and S. Le Maguer, “Creating new language and voice components for the updated MaryTTS text-to-speech synthesis platform,” inProceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan, May 2018
2018
-
[47]
ZMM-TTS: Zero-shot multilingual and multispeaker speech synthesis conditioned on self-supervised discrete speech representations,
C. Gong, X. Wang, E. Cooper, D. Wells, L. Wang, J. Dang, and J. Ya- magishi, “ZMM-TTS: Zero-shot multilingual and multispeaker speech synthesis conditioned on self-supervised discrete speech representations,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 4036–4051, 2024
2024
-
[48]
YourTTS: Towards zero-shot multi-speaker TTS and zero- shot voice conversion for everyone,
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G ¨olge, and M. A. Ponti, “YourTTS: Towards zero-shot multi-speaker TTS and zero- shot voice conversion for everyone,” inProceedings of the International Conference on Machine Learning (ICML), pp. 2709–2720, PMLR, Jun. 2022
2022
-
[49]
XTTS: A massively multilingual zero-shot text-to-speech model,
E. Casanova, K. Davis, E. G ¨olge, G. G ¨oknar, I. Gulea, L. Hart, and J. Weber, “XTTS: A massively multilingual zero-shot text-to-speech model,”arXiv preprint arXiv:2406.04904, 2024
Pith/arXiv arXiv 2024
-
[50]
Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,” inAdvances in Neural Information Processing Systems, vol. 33, pp. 8067–8077, 2020
2020
-
[51]
Grad- TTS: A diffusion probabilistic model for text-to-speech,
V . Popov, I. V ovk, V . Gogoryan, T. Sadekova, and M. Kudinov, “Grad- TTS: A diffusion probabilistic model for text-to-speech,” inProceedings of the International Conference on Machine Learning (ICML), pp. 8599– 8608, PMLR, Jul. 2021
2021
-
[52]
BigVGAN: A universal neural vocoder with large-scale training,
S. G. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “BigVGAN: A universal neural vocoder with large-scale training,”arXiv preprint arXiv:2206.04658, 2022
Pith/arXiv arXiv 2022
-
[53]
Exact prosody cloning in zero-shot multispeaker text-to-speech,
F. Lux, J. Koch, and N. T. Vu, “Exact prosody cloning in zero-shot multispeaker text-to-speech,” inProceedings of the 2022 IEEE Spoken Language Technology Workshop (SLT), pp. 962–969, IEEE, Jan. 2023
2022
-
[54]
FastPitch: Parallel text-to-speech with pitch prediction,
A. Ła ´ncucki, “FastPitch: Parallel text-to-speech with pitch prediction,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6588–6592, IEEE, Jun. 2021
2021
-
[55]
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” inProceedings of the International Conference on Machine Learning (ICML), pp. 5530–5540, PMLR, Jul. 2021
2021
-
[56]
Low-resource multilingual and zero- shot multispeaker TTS,
F. Lux, J. Koch, and N. T. Vu, “Low-resource multilingual and zero- shot multispeaker TTS,” inProceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 741–751, Nov. 2022
2022
-
[57]
Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme,
V . Popov, I. V ovk, V . Gogoryan, T. Sadekova, M. Kudinov, and J. Wei, “Diffusion-based voice conversion with fast maximum likelihood sam- pling scheme,”arXiv preprint arXiv:2109.13821, 2021
Pith/arXiv arXiv 2021
-
[58]
HiFi-GAN: Generative adversarial net- works for efficient and high-fidelity speech synthesis,
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial net- works for efficient and high-fidelity speech synthesis,” inAdvances in Neural Information Processing Systems, vol. 33, pp. 17022–17033, 2020
2020
-
[59]
Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y . Zhang, Y . Wang, R. Skerry-Ryan, R. A. Saurous, Y . Agiomyrgiannakis, and Y . Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” inProceedings of the IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4779– 47...
2018
-
[60]
Y . A. Li, A. Zare, and N. Mesgarani, “StarGANv2-VC: A diverse, unsu- pervised, non-parallel framework for natural-sounding voice conversion,” arXiv preprint arXiv:2107.10394, 2021
Pith/arXiv arXiv 2021
-
[61]
V oice conversion using speech-to-speech neuro-style transfer,
E. A. AlBadawy and S. Lyu, “V oice conversion using speech-to-speech neuro-style transfer,” inProc. Interspeech, 2020, pp. 4726–4730
2020
-
[62]
Self-supervised speech representation learning: A review,
A. Mohamed, H.-Y . Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Mangu, T. N. Sainath, and S. Watanabe, “Self-supervised speech representation learning: A review,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1179–1210, 2022
2022
-
[63]
Toward noise-aware audio deepfake detection: Survey, SNR-benchmarks, and practical recipes,
U. Sen, A. Luqman, and A. Chattopadhyay, “Toward noise-aware audio deepfake detection: Survey, SNR-benchmarks, and practical recipes,” arXiv preprint arXiv:2512.13744, 2025
arXiv 2025
-
[64]
C3-DINO: Joint contrastive and non-contrastive self-supervised learning for speaker verification,
C. Zhang and D. Yu, “C3-DINO: Joint contrastive and non-contrastive self-supervised learning for speaker verification,”IEEE Journal of Se- lected Topics in Signal Processing, vol. 16, no. 6, pp. 1273–1283, 2022
2022
-
[65]
Context and transcripts improve detection of deepfake audios of public figures,
C. Gao, M. Postiglione, J. Baldwin, N. Denisenko, I. Gortner, L. Fosdick, and V . S. Subrahmanian, “Context and transcripts improve detection of deepfake audios of public figures,”arXiv preprint arXiv:2601.13464, 2026
arXiv 2026
-
[66]
Fine-tuning self-supervised learning models for end-to-end pronunciation scoring,
A. I. Zahran, A. A. Fahmy, K. T. Wassif, and H. Bayomi, “Fine-tuning self-supervised learning models for end-to-end pronunciation scoring,” IEEE Access, vol. 11, pp. 112650–112663, 2023
2023
-
[67]
Using optimal f-measure and random resampling in gene ontology enrichment calculations,
W. Ge, Z. Fazal, and E. Jakobsson, “Using optimal f-measure and random resampling in gene ontology enrichment calculations,”Frontiers in Applied Mathematics and Statistics, vol. 5, p. 20, 2019
2019
-
[68]
Modified FDR controlling proce- dure for multi-stage analyses,
C. Tuglus and M. J. van der Laan, “Modified FDR controlling proce- dure for multi-stage analyses,”Statistical Applications in Genetics and Molecular Biology, vol. 8, no. 1, Art. 12, 2009
2009
-
[69]
C. Hanilc ¸i, M. Sahidullah, and T. Kinnunen, “Cyclostationarity analysis as a complement to self-supervised representations for speech deepfake detection,”arXiv preprint arXiv:2603.03921, 2026
arXiv 2026
-
[70]
Training-free cross- lingual dysarthria severity assessment via phonological subspace analysis in self-supervised speech representations,
B. Muller, A. A. Ortiz Barra ˜n´on, and L. Roberts, “Training-free cross- lingual dysarthria severity assessment via phonological subspace analysis in self-supervised speech representations,”medRxiv, 2026
2026
-
[71]
Bootstrap confidence regions for the intensity of a Poisson point process,
A. Cowling, P. Hall, and M. J. Phillips, “Bootstrap confidence regions for the intensity of a Poisson point process,”Journal of the American Statistical Association, vol. 91, no. 436, pp. 1516–1524, 1996
1996
-
[72]
PyTorch: An imperative style, high- performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An imperative style, high- performance deep learning library,” inAdvances in Neural Information Processing...
2019
-
[73]
Fairness through awareness,
C. Dwork, M. Hardt, T. Pitassi, O. Reingold, and R. Zemel, “Fairness through awareness,” inProceedings of the 3rd Innovations in Theoretical Computer Science Conference, 2012, pp. 214–226
2012
-
[74]
Phoneme-Level Deep- fake Detection Across Emotional Conditions Using Self-Supervised Em- beddings,
V . Nallaguntla, S. Kshirsagar, and A. R. Avila, “Phoneme-Level Deep- fake Detection Across Emotional Conditions Using Self-Supervised Em- beddings,”arXiv preprint arXiv:2605.03079, 2026
Pith/arXiv arXiv 2026
-
[75]
A review on fairness in machine learning,
D. Pessach and E. Shmueli, “A review on fairness in machine learning,” ACM Computing Surveys, vol. 55, no. 3, pp. 1–44, 2022
2022
-
[76]
Fairness definitions explained,
S. Verma and J. Rubin, “Fairness definitions explained,” inProceedings of the International Workshop on Software Fairness, pp. 1–7, May 2018
2018
-
[77]
Measuring algorithmic fairness,
D. Hellman, “Measuring algorithmic fairness,”Virginia Law Review, vol. 106, no. 4, pp. 811–866, 2020
2020
-
[78]
Bias preservation in machine learning: The legality of fairness metrics under EU non-discrimination law,
S. Wachter, B. Mittelstadt, and C. Russell, “Bias preservation in machine learning: The legality of fairness metrics under EU non-discrimination law,”West Virginia Law Review, vol. 123, no. 3, pp. 735–790, 2021
2021
-
[79]
Inclusive speaker verification with adaptive thresholding,
N. Jain and H. Wang, “Inclusive speaker verification with adaptive thresholding,”arXiv preprint arXiv:2111.05501, 2021
Pith/arXiv arXiv 2021
-
[80]
On fairness and calibration,
G. Pleiss, M. Raghavan, F. Wu, J. Kleinberg, and K. Q. Weinberger, “On fairness and calibration,” inAdvances in Neural Information Processing Systems, vol. 30, 2017
2017
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.