REVIEW 2 major objections 2 minor 2 cited by
The RADAR Challenge 2026 benchmark shows that audio deepfake detectors still produce high equal error rates when tested on multilingual utterances after compression, resampling, noise, and reverberation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-30 22:51 UTC pith:FI2SAV6A
load-bearing objection A standard challenge announcement paper that sets up a multilingual audio deepfake benchmark with media transformations and reports participation numbers, but stays descriptive without new methods or detailed results. the 2 major comments →
RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that binary real/fake classifiers evaluated by equal error rate continue to exhibit elevated error rates on the multilingual, media-transformed test set, indicating that robustness under compression, resampling, additive noise, and reverberation across languages remains an open problem.
What carries the argument
Equal error rate (EER) computed on the multilingual media-transformed evaluation utterances, which serves as the sole performance metric for ranking submissions.
Load-bearing premise
The particular choices of media transformations and language distribution in the test set accurately mirror the conditions that real-world audio deepfakes will encounter.
What would settle it
A follow-up experiment in which the top-performing systems from the challenge are re-evaluated on a fresh collection of real and synthetic audio that has passed through the same transformation pipeline but was collected independently of the challenge data.
If this is right
- Any production detector must maintain low EER after the audio has undergone lossy compression and resampling.
- Performance must generalize from English to at least five additional languages without retraining on target-language fakes.
- Noise and reverberation must be treated as first-class distortions rather than optional augmentations.
- Future systems will need to report EER on held-out multilingual transformed data rather than clean English test sets alone.
Where Pith is reading between the lines
- Deployment pipelines that ingest user-uploaded audio will require additional preprocessing stages or model updates to close the observed performance gap.
- The challenge protocol could be extended by adding a third phase that measures latency and memory use on edge devices under the same transformations.
- Cross-lingual transfer methods developed for this task may also improve robustness in related audio forensics problems such as speaker verification under distortion.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript announces the RADAR Challenge 2026, an APSIPA Grand Challenge on robust audio deepfake recognition under media transformations (compression, resampling, noise, reverberation). It describes a two-phase structure—an English-labeled development phase and a multilingual evaluation phase with >100,000 utterances across English, Singapore English, Mandarin Chinese, Taiwanese Mandarin, Japanese, and Vietnamese—along with the EER-based binary classification protocol, participation (33 development submissions, 22 evaluation submissions), and the observation that the reported results indicate remaining challenges in multilingual and transformed conditions.
Significance. If the dataset construction and protocol are reproducible, the challenge supplies a large-scale, multilingual benchmark that can serve as a community testbed for systems intended to operate in realistic media pipelines, potentially accelerating progress on robustness beyond clean English conditions.
major comments (2)
- [Results / overall results paragraph] The central claim that the reported results 'highlight the remaining challenges' is load-bearing for the paper's contribution, yet the manuscript supplies no quantitative EER values, baseline comparisons, or per-language/per-transformation breakdowns to support it; without these numbers the assertion remains unsupported by evidence.
- [Data set construction] Dataset construction section: no information is given on how the media transformations were implemented (specific codecs, SNR ranges, reverberation parameters) or on any validation that the transformed utterances preserve the original real/fake labels, which directly affects whether the EER metric can be interpreted as measuring robustness.
minor comments (2)
- [Abstract] The abstract states the evaluation contains 'more than 100,000 utterances' but does not break down the counts by language or by real/fake class; adding these counts would improve clarity.
- [Participation summary] The paper mentions '33 teams submitted to the development phase and 22 teams submitted to the final evaluation phase' but does not indicate whether the same teams participated in both phases or whether any overlap analysis was performed.
Simulated Author's Rebuttal
We appreciate the referee's constructive comments on our manuscript describing the RADAR Challenge 2026. We address each major comment below and will revise the paper to strengthen the presentation of results and dataset details.
read point-by-point responses
-
Referee: [Results / overall results paragraph] The central claim that the reported results 'highlight the remaining challenges' is load-bearing for the paper's contribution, yet the manuscript supplies no quantitative EER values, baseline comparisons, or per-language/per-transformation breakdowns to support it; without these numbers the assertion remains unsupported by evidence.
Authors: We agree that the manuscript would benefit from explicit quantitative support for this claim. In the revised version, we will add the EER results from the top-performing teams in both phases, include a baseline system for comparison, and provide breakdowns by language and transformation type where space permits. This will allow readers to better assess the remaining challenges. revision: yes
-
Referee: [Data set construction] Dataset construction section: no information is given on how the media transformations were implemented (specific codecs, SNR ranges, reverberation parameters) or on any validation that the transformed utterances preserve the original real/fake labels, which directly affects whether the EER metric can be interpreted as measuring robustness.
Authors: We acknowledge that additional details are needed. The revised manuscript will provide the specific implementation parameters for the media transformations and describe the procedures used to validate that the transformations do not alter the real/fake labels. revision: yes
Circularity Check
No significant circularity; purely descriptive challenge paper
full rationale
The paper is a standard challenge announcement describing task setup, multilingual dataset construction, evaluation protocol using EER, and participant outcomes (33 dev, 22 eval submissions). It advances no derivations, equations, predictions, models, or theorems. The central claim is observational: reported EERs indicate remaining difficulties under the defined conditions. No self-citations are load-bearing for any derivation, and no steps reduce by construction to inputs or prior author work. This is self-contained against external benchmarks as a simulation tool.
Axiom & Free-Parameter Ledger
read the original abstract
RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an English development phase with labeled data for analysis and paper writing, and a multilingual evaluation phase containing more than 100,000 utterances in English, Singapore English, Mandarin Chinese, Taiwanese Mandarin, Japanese, and Vietnamese. Systems are evaluated using equal error rate (EER) for binary real/fake classification. This paper describes the challenge task, the construction of the data set, the evaluation protocol, and the overall results. During the challenge, 33 teams submitted to the development phase and 22 teams submitted to the final evaluation phase. The reported results highlight the remaining challenges of robust audio deepfake detection under multilingual and media-transformed conditions.
Figures
Forward citations
Cited by 2 Pith papers
-
Large Audio Language Models for Spoofing-Aware Speaker Verification
Adapted LALMs can reach competitive spoofing-aware speaker verification (89.3% accuracy, 0.19 min a-DCF on an ASVspoof5 subset), though zero-shot performance is near chance.
-
Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
Fine-tuning Chatterbox and CosyVoice 3 on 50 Singlish speakers measurably raises accent similarity, and the gain persists on held-out speakers.
Reference graph
Works this paper leans on
-
[1]
ASVspoof 2019: A large-scale pub- lic database of synthesized, converted and replayed speech,
X. Wang et al., “ASVspoof 2019: A large-scale pub- lic database of synthesized, converted and replayed speech,”Computer Speech & Language, vol. 64, p. 101 114, Nov. 2020,ISSN: 08852308
work page 2019
-
[2]
Insights into deep non-linear filters for improved multi-channel speech enhancement,
X. Liu et al., “ASVspoof 2021: Towards spoofed and deepfake speech detection in the wild,”IEEE/ACM Transactions on Audio, Speech, and Language Process- ing, vol. 31, pp. 2507–2522, 2023.DOI: 10.1109/TASLP. 2023.3285283
-
[3]
X. Wang et al., “ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and ad- versarial attack detection using crowdsourced speech,” Computer Speech & Language, vol. 95, p. 101 825, 2026,ISSN: 0885-2308.DOI: https://doi.org/10.1016/j. csl.2025.101825
work page doi:10.1016/j 2026
-
[4]
Safe: Synthetic audio forensics evalua- tion challenge,
T. Kirill et al., “Safe: Synthetic audio forensics evalua- tion challenge,” inProc. ACM IH&MMSEC Workshop, 2025, pp. 174–180
work page 2025
-
[5]
Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,
J. Yi et al., “ADD 2022: The first audio deep synthesis detection challenge,” inProc. ICASSP, 2022, pp. 9216– 9220.DOI: 10.1109/ICASSP43922.2022.9746939
-
[6]
ADD 2023: The Second Audio Deepfake Detection Challenge,
J. Yi et al., “ADD 2023: The Second Audio Deepfake Detection Challenge,” inProc. IJCAI DADA Workshop, May 2023
work page 2023
-
[7]
Perturbed public voices (p 2v): A dataset for robust audio deepfake detection,
C. Gao, M. Postiglione, I. Gortner, S. Kraus, and V . Sub- rahmanian, “Perturbed public voices (p 2v): A dataset for robust audio deepfake detection,”arXiv preprint arXiv:2508.10949, 2025
-
[8]
Room impulse responses help attackers to evade deep fake detection,
H.-T. Luong, D.-T. Truong, K. A. Lee, and E. S. Chng, “Room impulse responses help attackers to evade deep fake detection,” inProc. SLT 2024, IEEE, 2024, pp. 623–629
work page 2024
-
[9]
DeePen: Penetration Testing for Audio Deepfake Detection
N. M ¨uller et al., “Deepen: Penetration testing for audio deepfake detection,”arXiv preprint arXiv:2502.20427, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[10]
Investigating the impact of speech enhancement on audio deep- fake detection in noisy environments,
S. Kshirsagar, A. R. Avila, et al., “Investigating the impact of speech enhancement on audio deep- fake detection in noisy environments,”arXiv preprint arXiv:2603.14767, 2026
-
[11]
Mlaad: The multi-language audio anti-spoofing dataset,
N. M. M ¨uller et al., “Mlaad: The multi-language audio anti-spoofing dataset,” inProc. IJCNN 2024, IEEE, 2024, pp. 1–7
work page 2024
-
[12]
Sea-spoof: Bridging the gap in multilingual audio deepfake detection for south-east asian,
J. Wu, N. Hou, Z. Pan, Q. Zhang, S. H. Bhupendra, and S. Mondal, “Sea-spoof: Bridging the gap in multilingual audio deepfake detection for south-east asian,”arXiv preprint arXiv:2509.19865, 2025
-
[13]
Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation,
H.-T. Luong, H. Li, L. Zhang, K. A. Lee, and E. S. Chng, “Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation,” inProc. ICASSP, 2025.DOI: 10 . 1109 / ICASSP49660 . 2025 . 10888070
work page 2025
-
[14]
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,
H. Zen et al., “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” inInterspeech 2019, 2019, pp. 2638–2642
work page 2019
-
[15]
JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech,
J. Lim, J. Ye, S. Chun, S. Kim, and J. Cho, “JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech,” inInterspeech 2022, 2022, pp. 2338–2342
work page 2022
-
[16]
YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G¨olge, and M. A. Ponti, “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,” inICML, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., ser. Proceed- ings of Machine Learning Research, vol. 162, PMLR, 17–23 Jul 2022, pp. 2709–2720
work page 2022
-
[17]
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,
E. Casanova et al., “XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model,” inInterspeech 2024, 2024, pp. 4978–4982
work page 2024
-
[18]
Z. Du et al., “Cosyvoice: A scalable multilingual zero- shot text-to-speech synthesizer based on supervised se- mantic tokens,”arXiv preprint arXiv:2407.05407, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[19]
Common voice: A massively- multilingual speech corpus,
R. Ardila et al., “Common voice: A massively- multilingual speech corpus,” inProceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 2020, pp. 4211–4215
work page 2020
-
[20]
arXiv preprint arXiv:2111.09344 , year=
D. Galvez et al., “The people’s speech: A large-scale di- verse english speech recognition dataset for commercial usage,”arXiv preprint arXiv:2111.09344, 2021
-
[21]
Building the Singapore English National Speech Corpus,
J. X. Koh et al., “Building the Singapore English National Speech Corpus,” inInterspeech 2019, pp. 321– 325
work page 2019
-
[22]
imagicdatatech.com/index.php/home/dataopensource/ data info/id/101, Accessed: 2019-05, 2019
Magic Data Technology Co., Ltd.,Openslr68: Magic- data mandarin chinese read speech corpus, http://www. imagicdatatech.com/index.php/home/dataopensource/ data info/id/101, Accessed: 2019-05, 2019
work page 2019
-
[23]
Formosa speech recognition challenge 2020 and taiwanese across taiwan corpus,
Y .-F. Liao et al., “Formosa speech recognition challenge 2020 and taiwanese across taiwan corpus,” inProc. O- COCOSDA 2020, IEEE, 2020, pp. 65–70
work page 2020
-
[24]
Cpjd corpus: Crowd- sourced parallel speech corpus of japanese dialects,
S. Takamichi and H. Saruwatari, “Cpjd corpus: Crowd- sourced parallel speech corpus of japanese dialects,” in Proc. LREC 2018, 2018
work page 2018
-
[25]
D. C. Tran,FPT Open Speech Dataset (FOSD) - Viet- namese, version V4, Mendeley Data, 2020.DOI: 10 . 17632/k9sxg2twv4.4
work page 2020
-
[26]
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Z. Du et al., “CosyV oice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training,”arXiv preprint arXiv:2505.17589, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[27]
H. Hu et al., “Qwen3-tts technical report,”arXiv preprint arXiv:2601.15621, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[28]
Fish audio s2 technical report,
S. Liao et al., “Fish audio s2 technical report,”arXiv preprint arXiv:2603.08823, 2026
-
[29]
Statistics of natural reverberation enable perceptual separation of sound and space,
J. Traer and J. H. McDermott, “Statistics of natural reverberation enable perceptual separation of sound and space,”PNAS, vol. 113, no. 48, E7856–E7865, 2016
work page 2016
-
[30]
MUSAN: A Music, Speech, and Noise Corpus
D. Snyder, G. Chen, and D. Povey, “Musan: A music, speech, and noise corpus,”arXiv preprint arXiv:1510.08484, 2015
work page internal anchor Pith review Pith/arXiv arXiv 2015
-
[31]
Fma: A dataset for music analysis,
M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson, “Fma: A dataset for music analysis,” in18th International Society for Music Information Retrieval Conference, 2017
work page 2017
-
[32]
A binaural room im- pulse response database for the evaluation of dereverber- ation algorithms,
M. Jeub, M. Schafer, and P. Vary, “A binaural room im- pulse response database for the evaluation of dereverber- ation algorithms,” in2009 16th international conference on digital signal processing, IEEE, 2009, pp. 1–5
work page 2009
-
[33]
Image method for ef- ficiently simulating small-room acoustics,
J. B. Allen and D. A. Berkley, “Image method for ef- ficiently simulating small-room acoustics,”The Journal of the Acoustical Society of America, vol. 65, no. 4, pp. 943–950, 1979
work page 1979
-
[34]
Hierarchical and multimodal learning for hetero- geneous sound classification,
P. Anastasopoulou, F. A. Dal R ´ı, X. Serra, and F. Font, “Hierarchical and multimodal learning for hetero- geneous sound classification,” inProc. DCASE 2025, 2025
work page 2025
-
[35]
H. Tak, M. Todisco, X. Wang, J.-w. Jung, J. Yamagishi, and N. Evans, “Automatic Speaker Verification Spoofing and Deepfake Detection Using Wav2vec 2.0 and Data Augmentation,” inProc. Odyssey 2022, 2022, pp. 112– 119
work page 2022
-
[36]
Robust localization of partially fake speech: Metrics and out-of-domain evaluation,
H.-T. Luong, I. Rimon, H. Permuter, K. A. Lee, and E. S. Chng, “Robust localization of partially fake speech: Metrics and out-of-domain evaluation,” inProc. APSIPA ASC 2025, IEEE, 2025, pp. 2205–2210
work page 2025
-
[37]
Li, X., Chen, P.-Y ., and Wei, W
X. Li, P.-Y . Chen, and W. Wei, “Measuring the ro- bustness of audio deepfake detectors,”arXiv preprint arXiv:2503.17577, 2025
work page internal anchor Pith review arXiv 2025
-
[38]
Replay attacks against audio deepfake detection,
N. M ¨uller et al., “Replay attacks against audio deepfake detection,”Interspeech 2025, 2025
work page 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.