REVIEW 2 major objections 5 minor 33 references
Speech-trained deepfake detectors fail on synthetic sound effects, and even after fine-tuning they overfit to known generators rather than learning general acoustic fakes.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 12:33 UTC pith:MVPJ77KP
load-bearing objection Solid multi-generator sound-effect deepfake dataset that cleanly documents forgetting and generator-overfitting; the OOD claim is a bit over-sold but the resource itself is useful. the 2 major comments →
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Current deepfake detectors, even after joint-domain fine-tuning on speech and SynSFX, overfit to the synthesis artifacts of the seven seen text-to-audio models; they do not learn transferable acoustic signatures that separate natural from synthetic sound effects, so equal-error rates remain high (27–37 percent) on an unseen commercial generator while speech separability is preserved.
What carries the argument
The Shared Prompt Subset—1,890 identical text prompts rendered by every generator—plus a diagnostic split into Seen-Real/Unseen-Fake versus Unseen-Real/Seen-Fake conditions. These controls let the authors separate content effects from generator-specific artifacts and show that the dominant failure mode is the unseen fake generator.
Load-bearing premise
That one proprietary commercial text-to-audio API together with UrbanSound8K samples is a fair and representative out-of-domain test of whether detectors have learned general acoustic anomalies rather than mere generator fingerprints.
What would settle it
Train the same joint-domain models on SynSFX, then evaluate on several additional open or closed text-to-audio generators never seen during training; if equal-error rates fall to the low-single-digit range obtained on the original seven models, the overfitting claim is weakened.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SynSFX, a 178-hour corpus of 43,374 clips (26,452 synthetic from seven text-to-audio models A1–A7, 16,922 real from AudioCaps/Clotho/ESC-50/TACoS/WavCaps) for isolated non-speech audio deepfake detection, including a shared-prompt subset of 1,890 prompts. Zero-shot evaluation shows speech-centric detectors (AASIST, RawNet2) near chance and EAT-AASIST at 23.71% EER. Fine-tuning on SynSFX alone yields strong in-domain EERs (3.23%/2.36%) but catastrophic forgetting on speech (33.61%/24.04% EER). Joint SynSFX+ASVspoof training restores speech performance while preserving low SynSFX EERs, yet Table III shows high EERs (27–37%) on unseen proprietary generator + UrbanSound8K. t-SNE and shared-prompt score variance analyses are used to argue that detectors overfit to seen-generator artifacts rather than learning universal acoustic anomalies.
Significance. If the central claims hold, SynSFX fills a genuine gap relative to EnvSDD and CompSpoof-style mixed speech–background sets by providing isolated multi-generator sound effects with transparent provenance and a controlled shared-prompt subset. The joint-training results and the explicit demonstration of speech-domain catastrophic forgetting are useful empirical findings for multi-domain audio forensics. The shared-prompt design is a concrete methodological contribution that future work can exploit for prompt-matched generator comparisons. Strengths include clear split protocols, consistent reporting of EER/AUC/F1, and public data availability. The main limitation is that the strongest generalization claim rests on a single closed-source unseen generator of modest scale.
major comments (2)
- §IV-E and §IV-H / Table III: The central claim that detectors overfit to the seven seen TTA artifacts (rather than learning universal anomalies) is supported primarily by the high EERs under Seen Real + Unseen Fake (27–34%) and Unseen Real + Unseen Fake (32–37%). That claim is only as strong as the assumption that the single proprietary commercial API (A8, 1,113 clips total, 1:1 balance) is a clean, representative generator-shift probe. Because A8 is closed-source, its pipeline and residual artifacts are unknown; it could be an outlier relative to A1–A7. Without at least one additional independent open unseen generator (or a multi-generator OOD suite), the attribution of collapse specifically to “generator-specific artifacts” remains under-supported relative to other uncontrolled distribution shifts.
- §IV-E (Out-of-Domain Test Set): The authentic half of the OOD set is drawn exclusively from UrbanSound8K, whose short-clip urban taxonomy, recording conditions, and class distribution differ from the caption-driven real sources used in training (AudioCaps, Clotho, WavCaps, etc.). Although the paper’s own decomposition shows that real-source shift alone is mild (Unseen Real + Seen Fake EERs of 2–5%), a matched real-source control (or multiple real OOD sources) would more cleanly isolate generator shift from real-domain mismatch and strengthen the causal interpretation offered in §IV-H.
minor comments (5)
- Table I: A6 and A7 exclusive-prompt counts (1888/1884) are slightly below the 1890 shared figure; a one-sentence note on the missing prompts would remove any appearance of imbalance.
- §III-C and Table I: Sample-rate heterogeneity (16 kHz vs 44.1 kHz) is acknowledged and all audio is later resampled to 16 kHz; a brief statement that native-rate artifacts are intentionally retained until the common preprocessing step would clarify the design choice.
- Figure 1 caption: The four class labels are clear, but adding the exact number of points per class (already given as n=300) and the perplexity/seed used for t-SNE would aid reproducibility.
- §IV-D: Fixed crop lengths (64 600 / 64 000 samples) and the CosineAnnealingWarmRestarts schedule are stated; reporting the final validation EER or early-stopping criterion would complete the training protocol description.
- References: Several 2025–2026 arXiv entries (EnvSDD, CompSpoof, ESDD2, MMAudio, TangoFlux) are appropriately cited; ensure final DOIs or camera-ready versions are updated if available at publication time.
Circularity Check
Empirical dataset-and-benchmark paper with measured EERs and diagnostics; no derivation chain that reduces to its own inputs by construction.
full rationale
SynSFX is a corpus-construction and transfer-evaluation study. Its central claims (speech-centric detectors collapse on isolated sound effects; SynSFX-only fine-tuning produces catastrophic forgetting on ASVspoof; joint-domain training restores speech performance while preserving low in-domain EER; zero-shot EER remains high on a held-out commercial generator) are direct numerical outcomes of training and testing protocols on explicitly partitioned data (Tables II–III, Sections IV-B–H). The shared-prompt subset is used only for post-hoc variance/range statistics of detector scores, not to define or force any prediction. t-SNE visualizations are likewise diagnostic. Citations to CompSpoofV2/ESDD2 (partial author overlap) and to the EAT-AASIST baseline supply background and initialization weights; they do not underwrite the SynSFX measurements or the overfitting conclusion. No equation equates a claimed result to a fitted parameter, no uniqueness theorem is imported to forbid alternatives, and no ansatz is smuggled via self-citation. The work is therefore self-contained against its own held-out splits and external speech corpora; circularity score is zero.
Axiom & Free-Parameter Ledger
free parameters (2)
- learning rate / weight decay / CosineAnnealing schedule
- input crop length (64600 / 64000 samples)
axioms (4)
- domain assumption Equal Error Rate and related binary metrics on held-out clips are valid proxies for deepfake detection quality in non-speech audio.
- domain assumption The seven selected open TTA models plus one proprietary commercial API adequately sample the space of modern text-to-audio generators for studying cross-generator generalization.
- domain assumption Real clips drawn from AudioCaps, Clotho, ESC-50, TACoS, WavCaps and UrbanSound8K are sufficiently authentic and distributionally distinct from the synthetic set.
- ad hoc to paper Resampling all audio to 16 kHz mono with peak normalization does not erase the generator-specific artifacts that detectors should learn.
read the original abstract
While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models.
Figures
Reference graph
Works this paper leans on
-
[1]
Audio deepfakes: A survey,
Z. Khanjani, G. Watson, and V . P. Janeja, “Audio deepfakes: A survey,”Frontiers in Big Data, vol. V olume 5 - 2022,
2022
-
[2]
Available: https://www.frontiersin.org/journals/big- data/articles/10.3389/fdata.2022.1001063
[Online]. Available: https://www.frontiersin.org/journals/big- data/articles/10.3389/fdata.2022.1001063
-
[3]
Audiobox: Unified audio generation with natural language prompts,
A. Vyas, B. Shi, M. Le, A. Tjandra, Y .-C. Wu, B. Guo, J. Zhang, X. Zhang, R. Adkins, W. Ngan, J. Wang, I. Cruz, B. Akula, A. Akinyemi, B. Ellis, R. Moritz, Y . Yungster, A. Rakotoarison, L. Tan, C. Summers, C. Wood, J. Lane, M. Williamson, and W.-N. Hsu, “Audiobox: Unified audio generation with natural language prompts,”
-
[4]
Available: https://arxiv.org/abs/2312.15821
[Online]. Available: https://arxiv.org/abs/2312.15821
-
[5]
Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, and K. A. Lee, “Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, p. 2507–2522, 2023. [Online]. Available: http://dx.doi.org/10.1109/TASLP.2023.3285283
-
[6]
Environmental sound recognition: A survey,
S. Chachada and C.-C. J. Kuo, “Environmental sound recognition: A survey,” in2013 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, 2013, pp. 1–9
2013
-
[7]
Audio surveillance: a systematic review,
M. Crocco, M. Cristani, A. Trucco, and V . Murino, “Audio surveillance: a systematic review,” 2014. [Online]. Available: https: //arxiv.org/abs/1409.7787
Pith/arXiv arXiv 2014
-
[8]
ASVspoof 2019: Future horizons in spoofed and fake audio detection,
M. Todisco, X. Wang, V . Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. W. D. Evans, T. H. Kinnunen, and K. A. Lee, “ASVspoof 2019: Future horizons in spoofed and fake audio detection,” inProceedings of INTERSPEECH 2019, 2019, pp. 1008–1012. [Online]. Available: https://dblp.org/rec/conf/interspeech/ Todisco0VSDNYEK19.html
2019
-
[9]
ASVspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation plan,
H. Delgado, N. W. D. Evans, T. Kinnunen, K. A. Lee, X. Liu, A. Nautsch, J. Patino, M. Sahidullah, M. Todisco, X. Wang, and J. Yamagishi, “ASVspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation plan,” ASVspoof Consortium, Evaluation Plan / Technical Report, Jul. 2021
2021
-
[10]
Fakeavceleb: A novel audio-video multimodal deepfake dataset,
H. Khalid, S. Tariq, M. Kim, and S. S. Woo, “Fakeavceleb: A novel audio-video multimodal deepfake dataset,” 2022. [Online]. Available: https://arxiv.org/abs/2108.05080
Pith/arXiv arXiv 2022
-
[11]
To-rawnet: Improving RawNet with TCN and orthogonal regularization for fake audio detection,
C. Wang, J. Yi, J. Tao, C. Zhang, S. Zhang, R. Fu, and X. Chen, “To-rawnet: Improving RawNet with TCN and orthogonal regularization for fake audio detection,” 2023. [Online]. Available: https://arxiv.org/abs/2305.13701
Pith/arXiv arXiv 2023
-
[12]
Twice attention networks for synthetic speech detection,
D. Yao, X. Yuan, X. Hu, and G. Guo, “Twice attention networks for synthetic speech detection,”Neurocomputing, vol. 559, p. 126799, 2023. [Online]. Available: https://doi.org/10.1016/j.neucom.2023.126799
-
[13]
Wavefake: A data set to facilitate audio deepfake detection,
J. Frank and L. Sch ¨onherr, “Wavefake: A data set to facilitate audio deepfake detection,” inProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), 2021. [Online]. Available: https://datasets-benchmarks-proceedings.neurips.cc/ paper/2021/file/c74d97b01eae257e44aa9d5bade97baf-Pap...
2021
-
[14]
Multi- lingual deepfake speech dataset for robust and generalizable detection,
C. O. Mawalim, Y . Wang, A. Adila, S. Okada, and M. Unoki, “Multi- lingual deepfake speech dataset for robust and generalizable detection,” IEEE Access, vol. 14, pp. 57 144–57 161, 2026
2026
-
[15]
Simple and controllable music generation,
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y . Adi, and A. D´efossez, “Simple and controllable music generation,” inProceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY , USA: Curran Associates Inc., 2023
2023
-
[16]
Audioldm: Text-to-audio generation with latent diffusion models,
H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. P. Mandic, W. Wang, and M. D. Plumbley, “Audioldm: Text-to-audio generation with latent diffusion models,” 2023. [Online]. Available: https: //arxiv.org/abs/2301.12503
Pith/arXiv arXiv 2023
-
[17]
Audioldm 2: Learning holistic audio generation with self-supervised pretraining,
H. Liu, Y . Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y . Wang, W. Wang, Y . Wang, and M. D. Plumbley, “Audioldm 2: Learning holistic audio generation with self-supervised pretraining,” 2023. [Online]. Available: https://arxiv.org/abs/2308.05734
Pith/arXiv arXiv 2023
-
[18]
Stable audio: Fast timing-conditioned latent audio diffusion,
Stability AI, “Stable audio: Fast timing-conditioned latent audio diffusion,” Stability AI, Technical Report (web publication), Sep
-
[19]
Available: https://stability.ai/research/stable-audio-fast- timing-conditioned-latent-audio-diffusion
[Online]. Available: https://stability.ai/research/stable-audio-fast- timing-conditioned-latent-audio-diffusion
-
[20]
Diffsound: Discrete diffusion model for text-to-sound generation,
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y . Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 1720–1733, 2023. [Online]. Available: https://dl.acm.org/doi/10.1109/TASLP.2023.3268730
-
[21]
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,
R. Huang, J. Huang, D. Yang, Y . Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,” inProceedings of the 40th International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol. 202, 2023, pp. 13 916–13 932. [Online]. Available: https...
2023
-
[22]
Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis,
H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y . Mitsufuji, “Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis,” 2025. [Online]. Available: https://arxiv.org/abs/2412.15322
Pith/arXiv arXiv 2025
-
[23]
C.-Y . Hung, N. Majumder, Z. Kong, A. Mehrish, A. A. Bagherzadeh, C. Li, R. Valle, B. Catanzaro, and S. Poria, “Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,” 2025. [Online]. Available: https://arxiv.org/abs/2412.21037
Pith/arXiv arXiv 2025
-
[24]
Envsdd: Benchmarking environmental sound deepfake detection,
H. Yin, Y . Xiao, R. K. Das, J. Bai, H. Liu, W. Wang, and M. D. Plumbley, “Envsdd: Benchmarking environmental sound deepfake detection,” 2025. [Online]. Available: https://arxiv.org/abs/2505.19203
arXiv 2025
-
[25]
X. Zhang, Y . Wang, L. Li, L. Jin, and M. Li, “Compspoof: A dataset and joint learning framework for component-level audio anti-spoofing countermeasures,” 2026. [Online]. Available: https: //arxiv.org/abs/2509.15804
arXiv 2026
-
[26]
Overview of esdd2: Environment-aware speech and sound deepfake detection challenge,
X. Zhang, H. Yin, Y . Xiao, L. Zhang, T. Dang, R. K. Das, and M. Li, “Overview of esdd2: Environment-aware speech and sound deepfake detection challenge,” 2026. [Online]. Available: https://arxiv.org/abs/2606.10791
Pith/arXiv arXiv 2026
-
[27]
Efficient audio transformer and aasist for environment sound deepfake detection in the esdd 2026 challenge,
J. Cao, C. Fan, J. Xue, Y . Xie, R. Fu, Z. Wen, J. Yi, Y . Ren, Z. Lv, and J. Tao, “Efficient audio transformer and aasist for environment sound deepfake detection in the esdd 2026 challenge,” inICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 21 781–21 783
2026
-
[28]
Clotho: an audio captioning dataset,
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: an audio captioning dataset,” in2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8,
2020
-
[29]
Barcelona, Spain: IEEE, May 2020, pp. 736–740. [Online]. Available: https://doi.org/10.1109/ICASSP40776.2020.9052990
-
[30]
TACOS: Temporally-aligned audio captions for language-audio pretraining,
P. Primus, F. Schmid, and G. Widmer, “TACOS: Temporally-aligned audio captions for language-audio pretraining,” inIEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025, Tahoe City, CA, USA, October 12-15, 2025. Tahoe City, CA, USA: IEEE, October 2025, pp. 1–5. [Online]. Available: https://doi.org/10.1109/W ASPAA66052.2025.11230997
doi:10.1109/w 2025
-
[31]
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y . Zou, and W. Wang, “WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,”IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 32, pp. 3339–3354, 2024. [Online]. Available: https://doi.org/10.1109/TASLP.2024.3419446
-
[32]
AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,
J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. W. D. Evans, “AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” in2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 6367–6371. [Online]. Available: https://dblp.org/rec/conf/icassp/JungH...
2022
-
[33]
A dataset and taxonomy for urban sound research,
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” inProceedings of the 22nd ACM International Conference on Multimedia (MM ’14), 2014, pp. 1041–1044. [Online]. Available: https://dl.acm.org/doi/10.1145/2647868.2655045
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.