Pith. sign in

REVIEW 3 major objections 5 minor 40 references

Current audio deepfake detectors fail on Brazilian Portuguese political speech, and how a fake is made matters far more than the speaker's gender or region.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

State-of-the-art audio deepfake detectors severely degrade on Brazilian Portuguese political speech, and the main source of performance gaps is the synthesis method, not demographic traits.

T0 review reviewed 2026-08-03 challenge →

load-bearing objection A useful new Portuguese political-speech deepfake benchmark with a credible core result, but the headline bias ranking needs statistical rework before it is taken at face value. the 3 major comments →

arxiv 2607.28770 v1 pith:2L6QNEAZ submitted 2026-07-30 eess.AS

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

classification eess.AS
keywords audio deepfake detectionParlaSpoof-BRpolitical speechelectoral integrityBrazilian Portuguesebias analysispartial manipulationanti-spoofing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds ParlaSpoof-BR, a benchmark of 134,400 audio files drawn from real recordings of Brazilian parliamentary sessions plus synthetic speech from ten voice-cloning systems and four partial-manipulation conditions, balanced by gender and region. It then asks whether state-of-the-art detectors can tell real political speech from fake speech, and which factors drive their errors. The paper finds that they cannot, at least not reliably: the best tested detector, DF-Arena-1B, reaches only a 32.30% equal-error rate, and even it misclassifies about half of genuine recordings at its default threshold. It also finds that the synthesis model used and the fraction of audio manipulated matter much more than speaker demographics: a 25% infill edit evades the detector more than 70% of the time, while gender and region account for under 4 percentage points of gap. The intended takeaway is that forensic tools for electoral integrity must be evaluated on domain-specific political speech and must defend against partial edits, not just fully synthetic clips.

Core claim

The paper shows that existing detectors fail on Brazilian Portuguese political speech: AASIST and AASIST-L are near chance (EER ~51-54%), while the strongest system, DF-Arena-1B, reaches only 32.30% EER and misclassifies about half of genuine clips. Partial manipulation is the most effective attack: a 25% infill edit drops DF-Arena-1B's recall to 29.2%, so an adversary succeeds over 70% of the time. The bias analysis ranks synthesis-model choice (68.5-pp recall gap) and manipulation extent (44.3-pp gap) far above gender (0.7-pp EER) and region (3.7-pp EER), concluding that methodological factors dominate demographic ones. Deployment risks include a 94.8% false-positive rate on OGG-compressed

What carries the argument

The load-bearing object is ParlaSpoof-BR, a public benchmark of 2,000 genuine utterances by 40 speakers balanced for gender and Brazil's five regions, expanded to 134,400 files with ten synthetic generators and audio infilling. Its design varies three dimensions independently—synthesis model, manipulation extent, and speaker demographics—which is what allows gap sizes to be compared across factors. The comparison device is the gap between best and worst levels of a factor (recall points for methodological factors, EER points for demographic factors). The central attack protocol is masked infilling: regenerating 25-75% of an utterance from surrounding audio, with word-level alignments and cro

Load-bearing premise

The paper's headline ranking depends on treating its recall-based gaps for synthesis model and infilling extent as directly comparable to its equal-error-rate-based gaps for gender and region; if those metrics are not interchangeable, the conclusion that methodological factors dominate demographic disparities is not established by the numbers as presented.

What would settle it

Recompute the bias ranking using one metric for every factor—for instance, report recall for gender and region at the same operating point used for synthesis model. If the gender or region recall gap approaches the 68.5-point synthesis-model gap, the dominance conclusion reverses; if it stays near 1-4 points on the same metric, the conclusion survives. This is a single calculation a reader can perform from the released dataset.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Detectors trained mainly on generic spoof-detection data should not be treated as deployment-ready for Brazilian Portuguese political speech; even the strongest tested model commits 32.3% errors at its equal-error operating point.
  • Electoral forensics must treat a single altered word as an open threat: a 25% infill edit evades the best tested detector with over 70% probability, and less manipulation produces better evasion than full synthesis.
  • If the methodological-bias ranking holds, improving detector diversity across synthesis models and manipulation degrees is a more direct route to fair and accurate detection than demographic reweighting alone.
  • Real-world pipelines need codec- and enhancement-aware handling: OGG roundtrips and speech enhancement turn genuine parliamentary recordings into false positives at rates (94.8% and 98.9%) that would swamp any alert queue.
  • The benchmark's regional stratification shows why language-level coverage is not enough: within Brazilian Portuguese, the North versus South accent gap (3.7 pp) is five times the gender gap (0.7 pp), so fairness audits for Portuguese need a regional axis.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Read strictly, the finding that synthesis model and manipulation extent dominate suggests the most effective fairness intervention is wider generator diversity in training, not demographic reweighting—an implication the paper does not spell out.
  • The partial-manipulation result implies that for consequential political audio, a detector's 'bona fide' verdict should not be taken at face value; content-level verification (e.g., confirming the original full recording exists) becomes necessary—a workflow consequence beyond the paper's stated scope.
  • The paper's dominance ranking mixes recall-based gaps for methods with EER-based gaps for demographics, so the quantitative comparison (68.5 pp vs 0.7 pp) is suggestive rather than metric-identical; re-running the ranking on a single metric would test the claim more directly.
  • Because the benchmark draws on a single parliamentary recording pipeline, a natural extension is to test whether the factor ranking replicates on televised debates, radio interviews, or messaging-app voice notes, which have different acoustics and codec histories.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces ParlaSpoof-BR, a Brazilian Portuguese political-speech audio deepfake benchmark built from 2,000 bona fide utterances of 40 speakers of the Brazilian Chamber of Deputies, balanced by gender and region, plus 123,200 spoofed files generated by five TTS systems, five VC systems, OmniVoice infilling at 25/50/75% and LLM-guided edits, and robustness variants (enhancement, lossy codecs, babble noise). It evaluates three pretrained detectors (AASIST, AASIST-L, DF-Arena-1B) on a 30,000-file core set. Main findings: AASIST variants are near chance (EER 50.98% and 53.70%), DF-Arena-1B is best but still weak (EER 32.30%, AUC 0.715); recall varies strongly by generator (31.2–99.7%); 25% partial manipulation yields 29.2% recall; enhancement/lossy compression/babble conditions cause large errors. The paper's third contribution is a bias-factor ranking (Table V) concluding that methodological factors dominate demographic disparities.

Significance. If the bias ranking is repaired, this is a valuable contribution. ParlaSpoof-BR is the first public Portuguese political-speech deepfake benchmark with regional and gender stratification, and it provides concrete evidence that pretrained detectors degrade sharply on spontaneous parliamentary Portuguese. The partial-manipulation protocol and the explicit separation of core and robustness sets are strengths, and the public release of the dataset supports reproducibility and follow-up work. However, the headline bias conclusion is not established by the numbers as presented: Table V mixes non-commensurable metrics and lacks uncertainty quantification. The underlying data could support a corrected analysis, so the paper is a strong candidate after a major revision.

major comments (3)
  1. [IV.B, Table V] Table V ranks bias factors by mixing two non-commensurable metrics: recall gaps for methodological factors (Synthesis Model 68.5, Partial Manipulation 44.3, Speaker Similarity 17.9, UTMOS 13.2, SNR 11.1) and EER gaps for demographic factors (Region 3.7, Gender 0.7). Recall is a threshold-dependent true-positive rate; EER is a threshold-independent summary of the ROC curve. Moreover, the max–min range is an order statistic that increases with the number of factor levels: Synthesis Model has 10 generators, Partial Manipulation has 4 levels, Gender has 2, Region has 5. With 40 speakers and no confidence intervals or significance tests, the 0.7 and 3.7 pp demographic gaps may be sampling noise. Because this table is the basis for contribution (3) and the abstract's 'methodological factors dominating' claim, the conclusion is not supported as presented. Please recompute all gaps on a common m
  2. [IV.B, Partial Manipulation] The claim that 'less manipulation produces better evasion' compares the 25%-infilling recall (29.2%) with recall for Qwen3-TTS (31.2%) and VoxCPM2 (41.4%) — the two weakest TTS systems. The same detector's recall on full OmniVoice synthesis is 87.5%, and the mean recall across full TTS is 70.5% (Table IV). Thus 25% partial manipulation is not more evasive than full synthesis on average or for the same generator; the headline 'evade with over 70% probability' conflates attack type with generator choice. To support contribution (2), compare partial manipulation with full synthesis of the same underlying generator and with the mean of the full-synthesis family, and rephrase the conclusion accordingly.
  3. [IV.B, Region and Gender] The demographic analysis treats utterances as independent, but each of the 40 speakers contributes 50 utterances, and each region has only 8 speakers. The 3.7 pp EER gap between 'Norte' and 'Sul' and the even smaller 0.7 pp gender gap could be driven by a few individual speakers rather than by regional or gender effects. The text interprets the regional pattern as 'regional acoustic patterns (prosodic rhythm, vowel reduction) correlate with detection cues' without a speaker-level mixed-effects model or per-speaker aggregation. Please report per-speaker error distributions and a speaker-level significance test before drawing conclusions about demographic disparities.
minor comments (5)
  1. [Throughout] There are several LaTeX artifacts such as 'V oice' instead of 'Voice' (e.g., 'V oice Conversion' in Section III.B, 'V oice References' in Table I). Please proofread before final submission.
  2. [Table V] The table would be clearer if it explicitly listed, for each factor, the metric used (recall vs. EER) and the number of factor levels, so readers can see the non-commensurability at a glance.
  3. [IV.A] The text says 'On average, the detector performs better on VC than on TTS' based on the mean rows of Table IV. Since each generator contributes the same number of samples, the unweighted mean is appropriate, but saying 'unweighted average across generators' would make the definition explicit.
  4. [III.A] The paper uses recordings of real parliamentarians and creates synthetic impersonations. A short data-statement paragraph addressing consent, intended use, and potential misuse would strengthen the paper's ethical presentation.
  5. [Abstract] The abstract says 'state-of-the-art audio deepfake detectors' but only three detectors are evaluated. Suggest 'three state-of-the-art pretrained detectors' or similar to avoid overgeneralization.

Circularity Check

0 steps flagged

No significant circularity: the benchmark evaluation is self-contained; only a minor overlapping-authorship citation appears, and it is not load-bearing.

full rationale

The central claims (generalization failure, partial-manipulation evasion, and the bias ranking) are obtained by running three pretrained public detectors on a newly constructed dataset; no parameter is fitted to ParlaSpoof-BR and then reported as a prediction. The detector weights are external artifacts (ASVspoof-trained AASIST/AASIST-L; Speech DF Arena's DF-Arena-1B), and the new data are separate from their training. The only self-reference is the citation of BRSpeech-DF [9] as the prior Portuguese resource; two authors of the present paper are among BRSpeech-DF's authors, but that citation is used only to position the new dataset and does not supply any load-bearing equation or theorem. The paper itself flags a genuine non-identifiability (Section IV.A: "the present experiment cannot distinguish language mismatch from other generator-specific factors"), and the Table V comparison of recall gaps with EER gaps is a validity concern, not circularity: it does not make the result equivalent to its inputs by construction. Under the stated rules, no fitted-input-called-prediction, self-definitional, or imported-uniqueness pattern is present, so the circularity score is low.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

No fitted parameters exist in this empirical benchmark paper. The central claims rest on standard evaluation metrics and several domain assumptions about the reliability of pretrained detectors, archive metadata, WhisperX alignments, and ECAPA similarity. No new theoretical entities are introduced.

axioms (5)
  • domain assumption Pretrained detector weights (AASIST, AASIST-L, DF-Arena-1B) as released are valid representatives of state-of-the-art audio deepfake detection.
    The entire cross-domain generalization claim rests on evaluating these off-the-shelf checkpoints; Section III.E.
  • domain assumption Chamber of Deputies archive metadata (speaker identity, regional affiliation, segment boundaries) is accurate.
    Speaker and region labels for the 40 speakers come directly from archive metadata without independent verification; Section III.A.
  • domain assumption WhisperX word-level timestamps are sufficiently accurate for constructing semantically targeted infilling targets and for measuring WER/CER in quality assessment.
    WhisperX drives both the infilling frame boundaries and the intelligibility metrics; Sections III.A and III.D.
  • domain assumption ECAPA-TDNN cosine similarity is a valid proxy for perceived target-speaker similarity.
    Used to rank speaker-similarity bias in Table V; Section III.D.
  • standard math Standard definitions of EER, AUC, and macro-F1 are used throughout.
    These are standard evaluation metrics; no non-standard definitions introduced.

reviewed 2026-08-03 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil." pith.science (2026). https://pith.science/paper/2L6QNEAZ

@misc{pith2026260728770,
  author       = {Pith},
  title        = {Pith review of: Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2L6QNEAZ}},
  note         = {Machine review of arXiv:2607.28770}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent advances in generative artificial intelligence have made it easier to fabricate statements and amplify political disinformation during elections. We introduce ParlaSpoof-BR, an audio deepfake dataset derived from recordings of the Brazilian Chamber of Deputies and expanded with synthetic utterances from diverse text-to-speech and voice conversion models. Using ParlaSpoof-BR, we benchmark state-of-the-art audio deepfake detectors, examine their ability to generalize to Brazilian Portuguese political speech, and investigate potential biases in their predictions. Our analysis reveals that current systems struggle to provide consistent decisions across the diversity represented in the dataset, with methodological factors (synthesis model choice, manipulation extent) dominating over demographic disparities. ParlaSpoof-BR provides a domain-specific benchmark for studying audio deepfake detection in a socially consequential and underrepresented setting, supporting the development of more robust detection systems for electoral integrity in Brazil.

Figures

Figures reproduced from arXiv: 2607.28770 by Alef Iury Ferreira, Anderson da Silva Soares, Beatriz Almeida Fel\'icio, Daniel Casanova, Frederico Santos de Oliveira, Lucas Rafael Stefanel Gris, Raul C\'esar Reis Mata.

Figure 1
Figure 1. Figure 1: Confusion matrices for all detectors. AASIST and AASIST-L classify nearly all samples as spoof regardless of true label. DF-Arena-1B shows better [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

40 extracted references · 9 linked inside Pith

  1. [1]

    Deepfakes and disinformation: Ex- ploring the impact of synthetic political video on deception, uncer- tainty, and trust in news,

    C. Vaccari and A. Chadwick, “Deepfakes and disinformation: Ex- ploring the impact of synthetic political video on deception, uncer- tainty, and trust in news,”Social media+ society, vol. 6, no. 1, p. 2056305120903408, 2020

  2. [2]

    Deepfakes and democracy (theory): How synthetic audio- visual media for disinformation and hate speech threaten core democratic functions,

    M. Pawelec, “Deepfakes and democracy (theory): How synthetic audio- visual media for disinformation and hate speech threaten core democratic functions,”Digital Society, vol. 1, no. 2, p. 19, 2022

  3. [3]

    Do deepfake videos undermine our epistemic trust? a thematic analysis of tweets that discuss deepfakes in the Russian invasion of Ukraine,

    J. Twomey, D. Ching, M. P. Aylett, M. Quayle, C. Linehan, and G. Mur- phy, “Do deepfake videos undermine our epistemic trust? a thematic analysis of tweets that discuss deepfakes in the Russian invasion of Ukraine,”PLOS ONE, vol. 18, no. 10, p. e0291668, 2023

  4. [4]

    Analyzing misinformation claims during the 2022 Brazilian general election on WhatsApp, Twitter, and Kwai,

    S. A. Hale, A. Belisario, A. N. Mostafa, N. Marchal, C. C. Bento, F. M. Simon, F. Cant ´u, C. Scannavino, and P. N. Howard, “Analyzing misinformation claims during the 2022 Brazilian general election on WhatsApp, Twitter, and Kwai,”International Journal of Public Opinion Research, vol. 36, no. 2, 2024

  5. [5]

    ASVspoof2019 vs. ASVspoof5: Assessment and comparison,

    A. Weizman, Y . Ben-Shimol, and I. Lapidot, “ASVspoof2019 vs. ASVspoof5: Assessment and comparison,” inProc. Interspeech, 2025

  6. [6]

    Speech is silver, silence is golden: What do ASVspoof- trained models really learn?

    N. M. M ¨uller, F. Dieckmann, P. Czempin, R. U. Canals, K. B ¨ottinger, and J. Williams, “Speech is silver, silence is golden: What do ASVspoof- trained models really learn?” inProc. ASVspoof Workshop, 2021

  7. [7]

    Is synthetic voice detection research going into the right direction?

    S. Borz `‘i, O. Giudice, F. Stanco, and D. Allegra, “Is synthetic voice detection research going into the right direction?” inProc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2022

  8. [8]

    FairSSD: Understanding bias in synthetic speech detectors,

    A. K. Singh Yadav, K. Bhagtani, D. Salvi, P. Bestagini, and E. J. Delp, “FairSSD: Understanding bias in synthetic speech detectors,” inProc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2024

  9. [9]

    BRSpeech-DF: A deep fake synthetic speech dataset for Portuguese zero-shot TTS,

    A. C. Ferro Filho, R. Virgilli, L. A. Souza, F. S. de Oliveira, M. H. L. Ferreira, D. Tunnermann, G. R. Oliveira, A. S. Soares, and A. R. Galv˜ao Filho, “BRSpeech-DF: A deep fake synthetic speech dataset for Portuguese zero-shot TTS,” inProc. Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025, pp. 35 122–35 127

  10. [10]

    Brazil’s electoral deepfake law tested as ai-generated content targeted local elections,

    B. Farrugia, “Brazil’s electoral deepfake law tested as ai-generated content targeted local elections,” Digital Forensic Research Lab, Nov. 2024, accessed: 2026-07-29. [Online]. Available: https://dfrlab.org/ 2024/11/26/brazil-election-ai-deepfakes/

  11. [11]

    Does audio deepfake detection generalize?

    N. M. M ¨uller, P. Czempin, F. Dieckmann, A. Froghyar, and K. B¨ottinger, “Does audio deepfake detection generalize?” inProc. Interspeech, 2022, pp. 2783–2787

  12. [12]

    Speech DF arena: A leaderboard for speech DeepFake detection models,

    S. Dowerah, A. Kulkarni, A. Kulkarni, H. M. Tran, J. Kalda, A. Fe- dorchenko, B. Fauve, D. Lolive, T. Alum ¨ae, and M. Magimai-Doss, “Speech DF arena: A leaderboard for speech DeepFake detection models,”IEEE Open Journal of Signal Processing, 2026

  13. [13]

    Is audio spoof detection robust to laundering attacks?

    H. Ali, S. Subramani, S. Sudhir, R. Varahamurthy, and H. Malik, “Is audio spoof detection robust to laundering attacks?” inProc. ACM Work- shop on Information Hiding and Multimedia Security (IH&MMSec), 2024

  14. [14]

    Measuring the robustness of audio deep- fake detection under real-world corruption,

    X. Li, P.-Y . Chen, and W. Wei, “Measuring the robustness of audio deep- fake detection under real-world corruption,” inProc. ACM Conference on Data and Application Security and Privacy (CODASPY), 2026

  15. [15]

    Gender fairness in audio deepfake detection: Performance and disparity analysis,

    A. Fursule, S. Kshirsagar, and A. R. Avila, “Gender fairness in audio deepfake detection: Performance and disparity analysis,”arXiv preprint arXiv:2603.09007, 2026

  16. [16]

    Towards trustworthy audio deepfake detection: A systematic framework for diagnosing and mitigating gender bias,

    ——, “Towards trustworthy audio deepfake detection: A systematic framework for diagnosing and mitigating gender bias,”arXiv preprint arXiv:2605.09087, 2026

  17. [17]

    Towards quantifying and reducing language mismatch effects in cross- lingual speech anti-spoofing,

    T. Liu, I. Kukanov, Z. Pan, Q. Wang, H. B. Sailor, and K. A. Lee, “Towards quantifying and reducing language mismatch effects in cross- lingual speech anti-spoofing,” inProc. IEEE Spoken Language Technol- ogy Workshop (SLT), 2024

  18. [18]

    Revealing cross-lingual bias in synthetic speech detection under controlled conditions,

    V . Moreno, J. Lima, F. Sim ˜oes, R. Violato, M. Uliani Neto, F. Runstein, and P. Costa, “Revealing cross-lingual bias in synthetic speech detection under controlled conditions,” inProc. Symposium on Security and Privacy in Speech Communication (SPSC), 2025

  19. [19]

    How do neural spoofing countermeasures detect partially spoofed audio?

    T. Liu, L. Zhang, R. K. Das, Y . Ma, R. Tao, and H. Li, “How do neural spoofing countermeasures detect partially spoofed audio?” in Proc. Interspeech, 2024

  20. [20]

    Audio deepfake detectors vs. real fraud: The fall of benchmarks,

    J. Gajewska, A. Martinek, and E. Bartuzi-Trokielewicz, “Audio deepfake detectors vs. real fraud: The fall of benchmarks,” inProc. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Work- shops, 2026

  21. [21]

    WhisperX: Time-accurate speech transcription of long-form audio,

    M. Bain, J. Huh, T. Han, and A. Zisserman, “WhisperX: Time-accurate speech transcription of long-form audio,” inProc. Interspeech, 2023, pp. 4489–4493

  22. [22]

    OmniV oice: Masked generative speech modeling with bidirec- tional attention,

    k2-fsa, “OmniV oice: Masked generative speech modeling with bidirec- tional attention,” https://github.com/k2-fsa/OmniV oice, 2025

  23. [23]

    XTTS: A massively multilingual zero-shot text-to-speech model,

    E. Casanova, K. Davis, E. G ¨olge, G. G ¨oknar, I. Gulea, L. Hart, A. Alja- fari, J. Meyer, R. Morais, S. Olayemi, and J. Weber, “XTTS: A massively multilingual zero-shot text-to-speech model,” inProc. Interspeech, 2024, pp. 4978–4982

  24. [24]

    Chatterbox: Sota open-source tts,

    Resemble AI, “Chatterbox: Sota open-source tts,” 2025. [Online]. Available: https://github.com/resemble-ai/chatterbox

  25. [25]

    V oxCPM2 technical report,

    Y . Zhou, G. Zeng, X. Liu, X. Li, R. Yu, J. Gui, J. Wu, Z. Wang, X. Shen, R. Ye, Z. Zhang, J. Zhou, B. Bai, W. Sun, M. Deng, Q. Shi, Z. Wu, and Z. Liu, “V oxCPM2 technical report,” 2026. [Online]. Available: https://arxiv.org/abs/2606.06928

  26. [26]

    Qwen3-TTS technical report,

    H. Hu, X. Zhu, T. He, D. Guo, B. Zhang, X. Wang, Z. Guo, Z. Jiang, H. Hao, Z. Guo, X. Zhang, P. Zhang, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-TTS technical report,” 2026. [Online]. Available: https://arxiv.org/abs/2601.15621

  27. [27]

    Seed-VC: Zero-shot voice conversion and singing voice con- version,

    S. Liu, “Seed-VC: Zero-shot voice conversion and singing voice con- version,” https://github.com/Plachtaa/seed-vc, 2024

  28. [28]

    V oice conversion with just nearest neighbors,

    M. Baas, B. van Niekerk, and H. Kamper, “V oice conversion with just nearest neighbors,” inProc. Interspeech, 2023, pp. 2053–2057

  29. [29]

    OpenV oice: Versatile instant voice cloning,

    Z. Qin, W. Zhao, X. Yu, and X. Sun, “OpenV oice: Versatile instant voice cloning,”arXiv preprint arXiv:2312.01479, 2023

  30. [30]

    X-VC: Zero-shot streaming voice conversion in codec space,

    Q. Zheng, Y . Zhao, T. Wang, W. Chen, K. Xu, Y . Li, Q. Chen, X. Qiu, K. Yu, and X. Chen, “X-VC: Zero-shot streaming voice conversion in codec space,” 2026. [Online]. Available: https: //arxiv.org/abs/2604.12456

  31. [31]

    EZ-VC: Easy zero-shot any-to-any voice conversion,

    A. Joglekar, D. Singh, R. R. Bhatia, and S. Umesh, “EZ-VC: Easy zero-shot any-to-any voice conversion,” 2025. [Online]. Available: https://arxiv.org/abs/2505.16691

  32. [32]

    Claude sonnet model family,

    Anthropic, “Claude sonnet model family,” https://www.anthropic.com/ claude, 2024

  33. [33]

    Real time speech enhancement in the waveform domain,

    A. Defossez, G. Synnaeve, and Y . Adi, “Real time speech enhancement in the waveform domain,”arXiv preprint arXiv:2006.12847, 2020

  34. [34]

    Metricgan+: An improved version of metricgan for speech enhancement,

    S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y . Tsao, “Metricgan+: An improved version of metricgan for speech enhancement,”arXiv preprint arXiv:2104.03538, 2021

  35. [35]

    Silero vad: pre-trained enterprise-grade voice activity detec- tor (vad), number detector and language classifier,

    S. Team, “Silero vad: pre-trained enterprise-grade voice activity detec- tor (vad), number detector and language classifier,” https://github.com/ snakers4/silero-vad, 2024

  36. [36]

    UTMOS: Utokyo-sarulab system for voicemos challenge 2022,

    T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: Utokyo-sarulab system for voicemos challenge 2022,” inProceedings of Interspeech 2022, 2022, pp. 4521–4525

  37. [37]

    ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,

    B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” inProceedings of Interspeech 2020, 2020, pp. 3830–3834

  38. [38]

    AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,

    J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.- J. Yu, and N. Evans, “AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” inProc. IEEE ICASSP, 2022, pp. 6367–6371

  39. [39]

    End-to-end anti-spoofing with rawnet2,

    H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with rawnet2,” inICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 6369–6373

  40. [40]

    Harder or different? understanding generalization of audio deepfake detection,

    N. M. M ¨uller, N. Evans, H. Tak, P. Sperl, and K. B ¨ottinger, “Harder or different? understanding generalization of audio deepfake detection,” in Proc. Interspeech, 2024

This paper was first reviewed by deepseek-v4-flash on August 3, 2026.