REVIEW 3 major objections 5 minor 40 references
Current audio deepfake detectors fail on Brazilian Portuguese political speech, and how a fake is made matters far more than the speaker's gender or region.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
State-of-the-art audio deepfake detectors severely degrade on Brazilian Portuguese political speech, and the main source of performance gaps is the synthesis method, not demographic traits.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection A useful new Portuguese political-speech deepfake benchmark with a credible core result, but the headline bias ranking needs statistical rework before it is taken at face value. the 3 major comments →
Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper shows that existing detectors fail on Brazilian Portuguese political speech: AASIST and AASIST-L are near chance (EER ~51-54%), while the strongest system, DF-Arena-1B, reaches only 32.30% EER and misclassifies about half of genuine clips. Partial manipulation is the most effective attack: a 25% infill edit drops DF-Arena-1B's recall to 29.2%, so an adversary succeeds over 70% of the time. The bias analysis ranks synthesis-model choice (68.5-pp recall gap) and manipulation extent (44.3-pp gap) far above gender (0.7-pp EER) and region (3.7-pp EER), concluding that methodological factors dominate demographic ones. Deployment risks include a 94.8% false-positive rate on OGG-compressed
What carries the argument
The load-bearing object is ParlaSpoof-BR, a public benchmark of 2,000 genuine utterances by 40 speakers balanced for gender and Brazil's five regions, expanded to 134,400 files with ten synthetic generators and audio infilling. Its design varies three dimensions independently—synthesis model, manipulation extent, and speaker demographics—which is what allows gap sizes to be compared across factors. The comparison device is the gap between best and worst levels of a factor (recall points for methodological factors, EER points for demographic factors). The central attack protocol is masked infilling: regenerating 25-75% of an utterance from surrounding audio, with word-level alignments and cro
Load-bearing premise
The paper's headline ranking depends on treating its recall-based gaps for synthesis model and infilling extent as directly comparable to its equal-error-rate-based gaps for gender and region; if those metrics are not interchangeable, the conclusion that methodological factors dominate demographic disparities is not established by the numbers as presented.
What would settle it
Recompute the bias ranking using one metric for every factor—for instance, report recall for gender and region at the same operating point used for synthesis model. If the gender or region recall gap approaches the 68.5-point synthesis-model gap, the dominance conclusion reverses; if it stays near 1-4 points on the same metric, the conclusion survives. This is a single calculation a reader can perform from the released dataset.
If this is right
- Detectors trained mainly on generic spoof-detection data should not be treated as deployment-ready for Brazilian Portuguese political speech; even the strongest tested model commits 32.3% errors at its equal-error operating point.
- Electoral forensics must treat a single altered word as an open threat: a 25% infill edit evades the best tested detector with over 70% probability, and less manipulation produces better evasion than full synthesis.
- If the methodological-bias ranking holds, improving detector diversity across synthesis models and manipulation degrees is a more direct route to fair and accurate detection than demographic reweighting alone.
- Real-world pipelines need codec- and enhancement-aware handling: OGG roundtrips and speech enhancement turn genuine parliamentary recordings into false positives at rates (94.8% and 98.9%) that would swamp any alert queue.
- The benchmark's regional stratification shows why language-level coverage is not enough: within Brazilian Portuguese, the North versus South accent gap (3.7 pp) is five times the gender gap (0.7 pp), so fairness audits for Portuguese need a regional axis.
Where Pith is reading between the lines
- Read strictly, the finding that synthesis model and manipulation extent dominate suggests the most effective fairness intervention is wider generator diversity in training, not demographic reweighting—an implication the paper does not spell out.
- The partial-manipulation result implies that for consequential political audio, a detector's 'bona fide' verdict should not be taken at face value; content-level verification (e.g., confirming the original full recording exists) becomes necessary—a workflow consequence beyond the paper's stated scope.
- The paper's dominance ranking mixes recall-based gaps for methods with EER-based gaps for demographics, so the quantitative comparison (68.5 pp vs 0.7 pp) is suggestive rather than metric-identical; re-running the ranking on a single metric would test the claim more directly.
- Because the benchmark draws on a single parliamentary recording pipeline, a natural extension is to test whether the factor ranking replicates on televised debates, radio interviews, or messaging-app voice notes, which have different acoustics and codec histories.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ParlaSpoof-BR, a Brazilian Portuguese political-speech audio deepfake benchmark built from 2,000 bona fide utterances of 40 speakers of the Brazilian Chamber of Deputies, balanced by gender and region, plus 123,200 spoofed files generated by five TTS systems, five VC systems, OmniVoice infilling at 25/50/75% and LLM-guided edits, and robustness variants (enhancement, lossy codecs, babble noise). It evaluates three pretrained detectors (AASIST, AASIST-L, DF-Arena-1B) on a 30,000-file core set. Main findings: AASIST variants are near chance (EER 50.98% and 53.70%), DF-Arena-1B is best but still weak (EER 32.30%, AUC 0.715); recall varies strongly by generator (31.2–99.7%); 25% partial manipulation yields 29.2% recall; enhancement/lossy compression/babble conditions cause large errors. The paper's third contribution is a bias-factor ranking (Table V) concluding that methodological factors dominate demographic disparities.
Significance. If the bias ranking is repaired, this is a valuable contribution. ParlaSpoof-BR is the first public Portuguese political-speech deepfake benchmark with regional and gender stratification, and it provides concrete evidence that pretrained detectors degrade sharply on spontaneous parliamentary Portuguese. The partial-manipulation protocol and the explicit separation of core and robustness sets are strengths, and the public release of the dataset supports reproducibility and follow-up work. However, the headline bias conclusion is not established by the numbers as presented: Table V mixes non-commensurable metrics and lacks uncertainty quantification. The underlying data could support a corrected analysis, so the paper is a strong candidate after a major revision.
major comments (3)
- [IV.B, Table V] Table V ranks bias factors by mixing two non-commensurable metrics: recall gaps for methodological factors (Synthesis Model 68.5, Partial Manipulation 44.3, Speaker Similarity 17.9, UTMOS 13.2, SNR 11.1) and EER gaps for demographic factors (Region 3.7, Gender 0.7). Recall is a threshold-dependent true-positive rate; EER is a threshold-independent summary of the ROC curve. Moreover, the max–min range is an order statistic that increases with the number of factor levels: Synthesis Model has 10 generators, Partial Manipulation has 4 levels, Gender has 2, Region has 5. With 40 speakers and no confidence intervals or significance tests, the 0.7 and 3.7 pp demographic gaps may be sampling noise. Because this table is the basis for contribution (3) and the abstract's 'methodological factors dominating' claim, the conclusion is not supported as presented. Please recompute all gaps on a common m
- [IV.B, Partial Manipulation] The claim that 'less manipulation produces better evasion' compares the 25%-infilling recall (29.2%) with recall for Qwen3-TTS (31.2%) and VoxCPM2 (41.4%) — the two weakest TTS systems. The same detector's recall on full OmniVoice synthesis is 87.5%, and the mean recall across full TTS is 70.5% (Table IV). Thus 25% partial manipulation is not more evasive than full synthesis on average or for the same generator; the headline 'evade with over 70% probability' conflates attack type with generator choice. To support contribution (2), compare partial manipulation with full synthesis of the same underlying generator and with the mean of the full-synthesis family, and rephrase the conclusion accordingly.
- [IV.B, Region and Gender] The demographic analysis treats utterances as independent, but each of the 40 speakers contributes 50 utterances, and each region has only 8 speakers. The 3.7 pp EER gap between 'Norte' and 'Sul' and the even smaller 0.7 pp gender gap could be driven by a few individual speakers rather than by regional or gender effects. The text interprets the regional pattern as 'regional acoustic patterns (prosodic rhythm, vowel reduction) correlate with detection cues' without a speaker-level mixed-effects model or per-speaker aggregation. Please report per-speaker error distributions and a speaker-level significance test before drawing conclusions about demographic disparities.
minor comments (5)
- [Throughout] There are several LaTeX artifacts such as 'V oice' instead of 'Voice' (e.g., 'V oice Conversion' in Section III.B, 'V oice References' in Table I). Please proofread before final submission.
- [Table V] The table would be clearer if it explicitly listed, for each factor, the metric used (recall vs. EER) and the number of factor levels, so readers can see the non-commensurability at a glance.
- [IV.A] The text says 'On average, the detector performs better on VC than on TTS' based on the mean rows of Table IV. Since each generator contributes the same number of samples, the unweighted mean is appropriate, but saying 'unweighted average across generators' would make the definition explicit.
- [III.A] The paper uses recordings of real parliamentarians and creates synthetic impersonations. A short data-statement paragraph addressing consent, intended use, and potential misuse would strengthen the paper's ethical presentation.
- [Abstract] The abstract says 'state-of-the-art audio deepfake detectors' but only three detectors are evaluated. Suggest 'three state-of-the-art pretrained detectors' or similar to avoid overgeneralization.
Circularity Check
No significant circularity: the benchmark evaluation is self-contained; only a minor overlapping-authorship citation appears, and it is not load-bearing.
full rationale
The central claims (generalization failure, partial-manipulation evasion, and the bias ranking) are obtained by running three pretrained public detectors on a newly constructed dataset; no parameter is fitted to ParlaSpoof-BR and then reported as a prediction. The detector weights are external artifacts (ASVspoof-trained AASIST/AASIST-L; Speech DF Arena's DF-Arena-1B), and the new data are separate from their training. The only self-reference is the citation of BRSpeech-DF [9] as the prior Portuguese resource; two authors of the present paper are among BRSpeech-DF's authors, but that citation is used only to position the new dataset and does not supply any load-bearing equation or theorem. The paper itself flags a genuine non-identifiability (Section IV.A: "the present experiment cannot distinguish language mismatch from other generator-specific factors"), and the Table V comparison of recall gaps with EER gaps is a validity concern, not circularity: it does not make the result equivalent to its inputs by construction. Under the stated rules, no fitted-input-called-prediction, self-definitional, or imported-uniqueness pattern is present, so the circularity score is low.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Pretrained detector weights (AASIST, AASIST-L, DF-Arena-1B) as released are valid representatives of state-of-the-art audio deepfake detection.
- domain assumption Chamber of Deputies archive metadata (speaker identity, regional affiliation, segment boundaries) is accurate.
- domain assumption WhisperX word-level timestamps are sufficiently accurate for constructing semantically targeted infilling targets and for measuring WER/CER in quality assessment.
- domain assumption ECAPA-TDNN cosine similarity is a valid proxy for perceived target-speaker similarity.
- standard math Standard definitions of EER, AUC, and macro-F1 are used throughout.
Cite this review
Pith. "Pith review of Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil." pith.science (2026). https://pith.science/paper/2L6QNEAZ
@misc{pith2026260728770,
author = {Pith},
title = {Pith review of: Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil},
year = {2026},
howpublished = {\url{https://pith.science/paper/2L6QNEAZ}},
note = {Machine review of arXiv:2607.28770}
}
read the original abstract
Recent advances in generative artificial intelligence have made it easier to fabricate statements and amplify political disinformation during elections. We introduce ParlaSpoof-BR, an audio deepfake dataset derived from recordings of the Brazilian Chamber of Deputies and expanded with synthetic utterances from diverse text-to-speech and voice conversion models. Using ParlaSpoof-BR, we benchmark state-of-the-art audio deepfake detectors, examine their ability to generalize to Brazilian Portuguese political speech, and investigate potential biases in their predictions. Our analysis reveals that current systems struggle to provide consistent decisions across the diversity represented in the dataset, with methodological factors (synthesis model choice, manipulation extent) dominating over demographic disparities. ParlaSpoof-BR provides a domain-specific benchmark for studying audio deepfake detection in a socially consequential and underrepresented setting, supporting the development of more robust detection systems for electoral integrity in Brazil.
Figures
Reference graph
Works this paper leans on
-
[1]
Deepfakes and disinformation: Ex- ploring the impact of synthetic political video on deception, uncer- tainty, and trust in news,
C. Vaccari and A. Chadwick, “Deepfakes and disinformation: Ex- ploring the impact of synthetic political video on deception, uncer- tainty, and trust in news,”Social media+ society, vol. 6, no. 1, p. 2056305120903408, 2020
2020
-
[2]
Deepfakes and democracy (theory): How synthetic audio- visual media for disinformation and hate speech threaten core democratic functions,
M. Pawelec, “Deepfakes and democracy (theory): How synthetic audio- visual media for disinformation and hate speech threaten core democratic functions,”Digital Society, vol. 1, no. 2, p. 19, 2022
2022
-
[3]
Do deepfake videos undermine our epistemic trust? a thematic analysis of tweets that discuss deepfakes in the Russian invasion of Ukraine,
J. Twomey, D. Ching, M. P. Aylett, M. Quayle, C. Linehan, and G. Mur- phy, “Do deepfake videos undermine our epistemic trust? a thematic analysis of tweets that discuss deepfakes in the Russian invasion of Ukraine,”PLOS ONE, vol. 18, no. 10, p. e0291668, 2023
2023
-
[4]
Analyzing misinformation claims during the 2022 Brazilian general election on WhatsApp, Twitter, and Kwai,
S. A. Hale, A. Belisario, A. N. Mostafa, N. Marchal, C. C. Bento, F. M. Simon, F. Cant ´u, C. Scannavino, and P. N. Howard, “Analyzing misinformation claims during the 2022 Brazilian general election on WhatsApp, Twitter, and Kwai,”International Journal of Public Opinion Research, vol. 36, no. 2, 2024
2022
-
[5]
ASVspoof2019 vs. ASVspoof5: Assessment and comparison,
A. Weizman, Y . Ben-Shimol, and I. Lapidot, “ASVspoof2019 vs. ASVspoof5: Assessment and comparison,” inProc. Interspeech, 2025
2025
-
[6]
Speech is silver, silence is golden: What do ASVspoof- trained models really learn?
N. M. M ¨uller, F. Dieckmann, P. Czempin, R. U. Canals, K. B ¨ottinger, and J. Williams, “Speech is silver, silence is golden: What do ASVspoof- trained models really learn?” inProc. ASVspoof Workshop, 2021
2021
-
[7]
Is synthetic voice detection research going into the right direction?
S. Borz `‘i, O. Giudice, F. Stanco, and D. Allegra, “Is synthetic voice detection research going into the right direction?” inProc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2022
2022
-
[8]
FairSSD: Understanding bias in synthetic speech detectors,
A. K. Singh Yadav, K. Bhagtani, D. Salvi, P. Bestagini, and E. J. Delp, “FairSSD: Understanding bias in synthetic speech detectors,” inProc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2024
2024
-
[9]
BRSpeech-DF: A deep fake synthetic speech dataset for Portuguese zero-shot TTS,
A. C. Ferro Filho, R. Virgilli, L. A. Souza, F. S. de Oliveira, M. H. L. Ferreira, D. Tunnermann, G. R. Oliveira, A. S. Soares, and A. R. Galv˜ao Filho, “BRSpeech-DF: A deep fake synthetic speech dataset for Portuguese zero-shot TTS,” inProc. Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025, pp. 35 122–35 127
2025
-
[10]
Brazil’s electoral deepfake law tested as ai-generated content targeted local elections,
B. Farrugia, “Brazil’s electoral deepfake law tested as ai-generated content targeted local elections,” Digital Forensic Research Lab, Nov. 2024, accessed: 2026-07-29. [Online]. Available: https://dfrlab.org/ 2024/11/26/brazil-election-ai-deepfakes/
2024
-
[11]
Does audio deepfake detection generalize?
N. M. M ¨uller, P. Czempin, F. Dieckmann, A. Froghyar, and K. B¨ottinger, “Does audio deepfake detection generalize?” inProc. Interspeech, 2022, pp. 2783–2787
2022
-
[12]
Speech DF arena: A leaderboard for speech DeepFake detection models,
S. Dowerah, A. Kulkarni, A. Kulkarni, H. M. Tran, J. Kalda, A. Fe- dorchenko, B. Fauve, D. Lolive, T. Alum ¨ae, and M. Magimai-Doss, “Speech DF arena: A leaderboard for speech DeepFake detection models,”IEEE Open Journal of Signal Processing, 2026
2026
-
[13]
Is audio spoof detection robust to laundering attacks?
H. Ali, S. Subramani, S. Sudhir, R. Varahamurthy, and H. Malik, “Is audio spoof detection robust to laundering attacks?” inProc. ACM Work- shop on Information Hiding and Multimedia Security (IH&MMSec), 2024
2024
-
[14]
Measuring the robustness of audio deep- fake detection under real-world corruption,
X. Li, P.-Y . Chen, and W. Wei, “Measuring the robustness of audio deep- fake detection under real-world corruption,” inProc. ACM Conference on Data and Application Security and Privacy (CODASPY), 2026
2026
-
[15]
Gender fairness in audio deepfake detection: Performance and disparity analysis,
A. Fursule, S. Kshirsagar, and A. R. Avila, “Gender fairness in audio deepfake detection: Performance and disparity analysis,”arXiv preprint arXiv:2603.09007, 2026
Pith/arXiv arXiv 2026
-
[16]
——, “Towards trustworthy audio deepfake detection: A systematic framework for diagnosing and mitigating gender bias,”arXiv preprint arXiv:2605.09087, 2026
Pith/arXiv arXiv 2026
-
[17]
Towards quantifying and reducing language mismatch effects in cross- lingual speech anti-spoofing,
T. Liu, I. Kukanov, Z. Pan, Q. Wang, H. B. Sailor, and K. A. Lee, “Towards quantifying and reducing language mismatch effects in cross- lingual speech anti-spoofing,” inProc. IEEE Spoken Language Technol- ogy Workshop (SLT), 2024
2024
-
[18]
Revealing cross-lingual bias in synthetic speech detection under controlled conditions,
V . Moreno, J. Lima, F. Sim ˜oes, R. Violato, M. Uliani Neto, F. Runstein, and P. Costa, “Revealing cross-lingual bias in synthetic speech detection under controlled conditions,” inProc. Symposium on Security and Privacy in Speech Communication (SPSC), 2025
2025
-
[19]
How do neural spoofing countermeasures detect partially spoofed audio?
T. Liu, L. Zhang, R. K. Das, Y . Ma, R. Tao, and H. Li, “How do neural spoofing countermeasures detect partially spoofed audio?” in Proc. Interspeech, 2024
2024
-
[20]
Audio deepfake detectors vs. real fraud: The fall of benchmarks,
J. Gajewska, A. Martinek, and E. Bartuzi-Trokielewicz, “Audio deepfake detectors vs. real fraud: The fall of benchmarks,” inProc. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Work- shops, 2026
2026
-
[21]
WhisperX: Time-accurate speech transcription of long-form audio,
M. Bain, J. Huh, T. Han, and A. Zisserman, “WhisperX: Time-accurate speech transcription of long-form audio,” inProc. Interspeech, 2023, pp. 4489–4493
2023
-
[22]
OmniV oice: Masked generative speech modeling with bidirec- tional attention,
k2-fsa, “OmniV oice: Masked generative speech modeling with bidirec- tional attention,” https://github.com/k2-fsa/OmniV oice, 2025
2025
-
[23]
XTTS: A massively multilingual zero-shot text-to-speech model,
E. Casanova, K. Davis, E. G ¨olge, G. G ¨oknar, I. Gulea, L. Hart, A. Alja- fari, J. Meyer, R. Morais, S. Olayemi, and J. Weber, “XTTS: A massively multilingual zero-shot text-to-speech model,” inProc. Interspeech, 2024, pp. 4978–4982
2024
-
[24]
Chatterbox: Sota open-source tts,
Resemble AI, “Chatterbox: Sota open-source tts,” 2025. [Online]. Available: https://github.com/resemble-ai/chatterbox
2025
-
[25]
Y . Zhou, G. Zeng, X. Liu, X. Li, R. Yu, J. Gui, J. Wu, Z. Wang, X. Shen, R. Ye, Z. Zhang, J. Zhou, B. Bai, W. Sun, M. Deng, Q. Shi, Z. Wu, and Z. Liu, “V oxCPM2 technical report,” 2026. [Online]. Available: https://arxiv.org/abs/2606.06928
Pith/arXiv arXiv 2026
-
[26]
H. Hu, X. Zhu, T. He, D. Guo, B. Zhang, X. Wang, Z. Guo, Z. Jiang, H. Hao, Z. Guo, X. Zhang, P. Zhang, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-TTS technical report,” 2026. [Online]. Available: https://arxiv.org/abs/2601.15621
Pith/arXiv arXiv 2026
-
[27]
Seed-VC: Zero-shot voice conversion and singing voice con- version,
S. Liu, “Seed-VC: Zero-shot voice conversion and singing voice con- version,” https://github.com/Plachtaa/seed-vc, 2024
2024
-
[28]
V oice conversion with just nearest neighbors,
M. Baas, B. van Niekerk, and H. Kamper, “V oice conversion with just nearest neighbors,” inProc. Interspeech, 2023, pp. 2053–2057
2023
-
[29]
OpenV oice: Versatile instant voice cloning,
Z. Qin, W. Zhao, X. Yu, and X. Sun, “OpenV oice: Versatile instant voice cloning,”arXiv preprint arXiv:2312.01479, 2023
Pith/arXiv arXiv 2023
-
[30]
X-VC: Zero-shot streaming voice conversion in codec space,
Q. Zheng, Y . Zhao, T. Wang, W. Chen, K. Xu, Y . Li, Q. Chen, X. Qiu, K. Yu, and X. Chen, “X-VC: Zero-shot streaming voice conversion in codec space,” 2026. [Online]. Available: https: //arxiv.org/abs/2604.12456
Pith/arXiv arXiv 2026
-
[31]
EZ-VC: Easy zero-shot any-to-any voice conversion,
A. Joglekar, D. Singh, R. R. Bhatia, and S. Umesh, “EZ-VC: Easy zero-shot any-to-any voice conversion,” 2025. [Online]. Available: https://arxiv.org/abs/2505.16691
Pith/arXiv arXiv 2025
-
[32]
Claude sonnet model family,
Anthropic, “Claude sonnet model family,” https://www.anthropic.com/ claude, 2024
2024
-
[33]
Real time speech enhancement in the waveform domain,
A. Defossez, G. Synnaeve, and Y . Adi, “Real time speech enhancement in the waveform domain,”arXiv preprint arXiv:2006.12847, 2020
Pith/arXiv arXiv 2006
-
[34]
Metricgan+: An improved version of metricgan for speech enhancement,
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y . Tsao, “Metricgan+: An improved version of metricgan for speech enhancement,”arXiv preprint arXiv:2104.03538, 2021
Pith/arXiv arXiv 2021
-
[35]
Silero vad: pre-trained enterprise-grade voice activity detec- tor (vad), number detector and language classifier,
S. Team, “Silero vad: pre-trained enterprise-grade voice activity detec- tor (vad), number detector and language classifier,” https://github.com/ snakers4/silero-vad, 2024
2024
-
[36]
UTMOS: Utokyo-sarulab system for voicemos challenge 2022,
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: Utokyo-sarulab system for voicemos challenge 2022,” inProceedings of Interspeech 2022, 2022, pp. 4521–4525
2022
-
[37]
ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” inProceedings of Interspeech 2020, 2020, pp. 3830–3834
2020
-
[38]
AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,
J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.- J. Yu, and N. Evans, “AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” inProc. IEEE ICASSP, 2022, pp. 6367–6371
2022
-
[39]
End-to-end anti-spoofing with rawnet2,
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with rawnet2,” inICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 6369–6373
2021
-
[40]
Harder or different? understanding generalization of audio deepfake detection,
N. M. M ¨uller, N. Evans, H. Tak, P. Sperl, and K. B ¨ottinger, “Harder or different? understanding generalization of audio deepfake detection,” in Proc. Interspeech, 2024
2024
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.