Pith. sign in

REVIEW 2 major objections 5 minor 33 references

Speech-trained deepfake detectors fail on synthetic sound effects, and even after fine-tuning they overfit to known generators rather than learning general acoustic fakes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 12:33 UTC pith:MVPJ77KP

load-bearing objection Solid multi-generator sound-effect deepfake dataset that cleanly documents forgetting and generator-overfitting; the OOD claim is a bit over-sold but the resource itself is useful. the 2 major comments →

arxiv 2607.04848 v1 pith:MVPJ77KP submitted 2026-07-06 cs.SD cs.AI

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

classification cs.SD cs.AI
keywords audio deepfake detectionsound effectsspoofingnon-speech audiotext-to-audioSynSFX datasetcross-generator generalization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Speech deepfake detectors have improved rapidly, but the same systems largely fail when asked to spot synthetic environmental sound effects. The paper releases SynSFX, a 178-hour corpus of more than 43,000 clips that pairs real environmental audio with fakes from seven modern text-to-audio models, plus a shared-prompt subset that holds the text description fixed across generators. Zero-shot evaluation shows pure speech models collapse to near-random performance; even a more general model still yields a 23.7 percent equal-error rate. Fine-tuning on SynSFX alone restores strong in-domain accuracy yet erases speech-detection skill; joint training on speech and sound-effect data recovers both domains for seen generators. The decisive finding is that performance still collapses on an unseen commercial generator, and t-SNE embeddings of real versus unseen-fake sound effects remain entangled. The authors conclude that current detectors latch onto generator-specific fingerprints rather than universal acoustic anomalies, and that the shared-prompt design is intended to help future work isolate those fingerprints from content bias.

Core claim

Current deepfake detectors, even after joint-domain fine-tuning on speech and SynSFX, overfit to the synthesis artifacts of the seven seen text-to-audio models; they do not learn transferable acoustic signatures that separate natural from synthetic sound effects, so equal-error rates remain high (27–37 percent) on an unseen commercial generator while speech separability is preserved.

What carries the argument

The Shared Prompt Subset—1,890 identical text prompts rendered by every generator—plus a diagnostic split into Seen-Real/Unseen-Fake versus Unseen-Real/Seen-Fake conditions. These controls let the authors separate content effects from generator-specific artifacts and show that the dominant failure mode is the unseen fake generator.

Load-bearing premise

That one proprietary commercial text-to-audio API together with UrbanSound8K samples is a fair and representative out-of-domain test of whether detectors have learned general acoustic anomalies rather than mere generator fingerprints.

What would settle it

Train the same joint-domain models on SynSFX, then evaluate on several additional open or closed text-to-audio generators never seen during training; if equal-error rates fall to the low-single-digit range obtained on the original seven models, the overfitting claim is weakened.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces SynSFX, a 178-hour corpus of 43,374 clips (26,452 synthetic from seven text-to-audio models A1–A7, 16,922 real from AudioCaps/Clotho/ESC-50/TACoS/WavCaps) for isolated non-speech audio deepfake detection, including a shared-prompt subset of 1,890 prompts. Zero-shot evaluation shows speech-centric detectors (AASIST, RawNet2) near chance and EAT-AASIST at 23.71% EER. Fine-tuning on SynSFX alone yields strong in-domain EERs (3.23%/2.36%) but catastrophic forgetting on speech (33.61%/24.04% EER). Joint SynSFX+ASVspoof training restores speech performance while preserving low SynSFX EERs, yet Table III shows high EERs (27–37%) on unseen proprietary generator + UrbanSound8K. t-SNE and shared-prompt score variance analyses are used to argue that detectors overfit to seen-generator artifacts rather than learning universal acoustic anomalies.

Significance. If the central claims hold, SynSFX fills a genuine gap relative to EnvSDD and CompSpoof-style mixed speech–background sets by providing isolated multi-generator sound effects with transparent provenance and a controlled shared-prompt subset. The joint-training results and the explicit demonstration of speech-domain catastrophic forgetting are useful empirical findings for multi-domain audio forensics. The shared-prompt design is a concrete methodological contribution that future work can exploit for prompt-matched generator comparisons. Strengths include clear split protocols, consistent reporting of EER/AUC/F1, and public data availability. The main limitation is that the strongest generalization claim rests on a single closed-source unseen generator of modest scale.

major comments (2)
  1. §IV-E and §IV-H / Table III: The central claim that detectors overfit to the seven seen TTA artifacts (rather than learning universal anomalies) is supported primarily by the high EERs under Seen Real + Unseen Fake (27–34%) and Unseen Real + Unseen Fake (32–37%). That claim is only as strong as the assumption that the single proprietary commercial API (A8, 1,113 clips total, 1:1 balance) is a clean, representative generator-shift probe. Because A8 is closed-source, its pipeline and residual artifacts are unknown; it could be an outlier relative to A1–A7. Without at least one additional independent open unseen generator (or a multi-generator OOD suite), the attribution of collapse specifically to “generator-specific artifacts” remains under-supported relative to other uncontrolled distribution shifts.
  2. §IV-E (Out-of-Domain Test Set): The authentic half of the OOD set is drawn exclusively from UrbanSound8K, whose short-clip urban taxonomy, recording conditions, and class distribution differ from the caption-driven real sources used in training (AudioCaps, Clotho, WavCaps, etc.). Although the paper’s own decomposition shows that real-source shift alone is mild (Unseen Real + Seen Fake EERs of 2–5%), a matched real-source control (or multiple real OOD sources) would more cleanly isolate generator shift from real-domain mismatch and strengthen the causal interpretation offered in §IV-H.
minor comments (5)
  1. Table I: A6 and A7 exclusive-prompt counts (1888/1884) are slightly below the 1890 shared figure; a one-sentence note on the missing prompts would remove any appearance of imbalance.
  2. §III-C and Table I: Sample-rate heterogeneity (16 kHz vs 44.1 kHz) is acknowledged and all audio is later resampled to 16 kHz; a brief statement that native-rate artifacts are intentionally retained until the common preprocessing step would clarify the design choice.
  3. Figure 1 caption: The four class labels are clear, but adding the exact number of points per class (already given as n=300) and the perplexity/seed used for t-SNE would aid reproducibility.
  4. §IV-D: Fixed crop lengths (64 600 / 64 000 samples) and the CosineAnnealingWarmRestarts schedule are stated; reporting the final validation EER or early-stopping criterion would complete the training protocol description.
  5. References: Several 2025–2026 arXiv entries (EnvSDD, CompSpoof, ESDD2, MMAudio, TangoFlux) are appropriately cited; ensure final DOIs or camera-ready versions are updated if available at publication time.

Circularity Check

0 steps flagged

Empirical dataset-and-benchmark paper with measured EERs and diagnostics; no derivation chain that reduces to its own inputs by construction.

full rationale

SynSFX is a corpus-construction and transfer-evaluation study. Its central claims (speech-centric detectors collapse on isolated sound effects; SynSFX-only fine-tuning produces catastrophic forgetting on ASVspoof; joint-domain training restores speech performance while preserving low in-domain EER; zero-shot EER remains high on a held-out commercial generator) are direct numerical outcomes of training and testing protocols on explicitly partitioned data (Tables II–III, Sections IV-B–H). The shared-prompt subset is used only for post-hoc variance/range statistics of detector scores, not to define or force any prediction. t-SNE visualizations are likewise diagnostic. Citations to CompSpoofV2/ESDD2 (partial author overlap) and to the EAT-AASIST baseline supply background and initialization weights; they do not underwrite the SynSFX measurements or the overfitting conclusion. No equation equates a claimed result to a fitted parameter, no uniqueness theorem is imported to forbid alternatives, and no ansatz is smuggled via self-citation. The work is therefore self-contained against its own held-out splits and external speech corpora; circularity score is zero.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The paper is an empirical resource-and-benchmark contribution. Its claims rest on standard deep-learning training assumptions, the representativeness of the chosen generators and real sources, and the validity of EER/AUC as detection metrics. No free parameters are fitted to produce a theoretical prediction; hyper-parameters are ordinary training knobs. No new physical or mathematical entities are postulated.

free parameters (2)
  • learning rate / weight decay / CosineAnnealing schedule
    Standard AdamW settings (1e-4, 1e-5, T0=10) chosen for fine-tuning; affect absolute EER numbers but not the qualitative forgetting/generalization claims.
  • input crop length (64600 / 64000 samples)
    Fixed lengths for AASIST and EAT-AASIST; arbitrary but conventional for the architectures.
axioms (4)
  • domain assumption Equal Error Rate and related binary metrics on held-out clips are valid proxies for deepfake detection quality in non-speech audio.
    Used throughout Tables II–III and Section IV without further justification; standard in ASVspoof-style literature.
  • domain assumption The seven selected open TTA models plus one proprietary commercial API adequately sample the space of modern text-to-audio generators for studying cross-generator generalization.
    Stated in Sections III-A and IV-E; load-bearing for the 'illusion of generalization' claim.
  • domain assumption Real clips drawn from AudioCaps, Clotho, ESC-50, TACoS, WavCaps and UrbanSound8K are sufficiently authentic and distributionally distinct from the synthetic set.
    Section III-A and IV-E; required for labeling and OOD real evaluation.
  • ad hoc to paper Resampling all audio to 16 kHz mono with peak normalization does not erase the generator-specific artifacts that detectors should learn.
    Section IV-D preprocessing step; could systematically remove high-frequency cues present in 44.1 kHz generators.

pith-pipeline@v1.1.0-grok45 · 16465 in / 2828 out tokens · 32142 ms · 2026-07-11T12:33:38.574886+00:00 · methodology

0 comments
read the original abstract

While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models.

Figures

Figures reproduced from arXiv: 2607.04848 by Carsten Maple, Linxi Li, Liwei Jin, Qianwei Guo, Yechen Wang, Yuncong Yu.

Figure 1
Figure 1. Figure 1: t-SNE visualization of the penultimate-layer embeddings for the joint-trained (a) AASIST and (b) EAT-AASIST models. Both architectures maintain [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 3 canonical work pages

  1. [1]

    Audio deepfakes: A survey,

    Z. Khanjani, G. Watson, and V . P. Janeja, “Audio deepfakes: A survey,”Frontiers in Big Data, vol. V olume 5 - 2022,

  2. [2]

    Available: https://www.frontiersin.org/journals/big- data/articles/10.3389/fdata.2022.1001063

    [Online]. Available: https://www.frontiersin.org/journals/big- data/articles/10.3389/fdata.2022.1001063

  3. [3]

    Audiobox: Unified audio generation with natural language prompts,

    A. Vyas, B. Shi, M. Le, A. Tjandra, Y .-C. Wu, B. Guo, J. Zhang, X. Zhang, R. Adkins, W. Ngan, J. Wang, I. Cruz, B. Akula, A. Akinyemi, B. Ellis, R. Moritz, Y . Yungster, A. Rakotoarison, L. Tan, C. Summers, C. Wood, J. Lane, M. Williamson, and W.-N. Hsu, “Audiobox: Unified audio generation with natural language prompts,”

  4. [4]

    Available: https://arxiv.org/abs/2312.15821

    [Online]. Available: https://arxiv.org/abs/2312.15821

  5. [5]

    Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,

    X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, and K. A. Lee, “Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, p. 2507–2522, 2023. [Online]. Available: http://dx.doi.org/10.1109/TASLP.2023.3285283

  6. [6]

    Environmental sound recognition: A survey,

    S. Chachada and C.-C. J. Kuo, “Environmental sound recognition: A survey,” in2013 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, 2013, pp. 1–9

  7. [7]

    Audio surveillance: a systematic review,

    M. Crocco, M. Cristani, A. Trucco, and V . Murino, “Audio surveillance: a systematic review,” 2014. [Online]. Available: https: //arxiv.org/abs/1409.7787

  8. [8]

    ASVspoof 2019: Future horizons in spoofed and fake audio detection,

    M. Todisco, X. Wang, V . Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. W. D. Evans, T. H. Kinnunen, and K. A. Lee, “ASVspoof 2019: Future horizons in spoofed and fake audio detection,” inProceedings of INTERSPEECH 2019, 2019, pp. 1008–1012. [Online]. Available: https://dblp.org/rec/conf/interspeech/ Todisco0VSDNYEK19.html

  9. [9]

    ASVspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation plan,

    H. Delgado, N. W. D. Evans, T. Kinnunen, K. A. Lee, X. Liu, A. Nautsch, J. Patino, M. Sahidullah, M. Todisco, X. Wang, and J. Yamagishi, “ASVspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation plan,” ASVspoof Consortium, Evaluation Plan / Technical Report, Jul. 2021

  10. [10]

    Fakeavceleb: A novel audio-video multimodal deepfake dataset,

    H. Khalid, S. Tariq, M. Kim, and S. S. Woo, “Fakeavceleb: A novel audio-video multimodal deepfake dataset,” 2022. [Online]. Available: https://arxiv.org/abs/2108.05080

  11. [11]

    To-rawnet: Improving RawNet with TCN and orthogonal regularization for fake audio detection,

    C. Wang, J. Yi, J. Tao, C. Zhang, S. Zhang, R. Fu, and X. Chen, “To-rawnet: Improving RawNet with TCN and orthogonal regularization for fake audio detection,” 2023. [Online]. Available: https://arxiv.org/abs/2305.13701

  12. [12]

    Twice attention networks for synthetic speech detection,

    D. Yao, X. Yuan, X. Hu, and G. Guo, “Twice attention networks for synthetic speech detection,”Neurocomputing, vol. 559, p. 126799, 2023. [Online]. Available: https://doi.org/10.1016/j.neucom.2023.126799

  13. [13]

    Wavefake: A data set to facilitate audio deepfake detection,

    J. Frank and L. Sch ¨onherr, “Wavefake: A data set to facilitate audio deepfake detection,” inProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), 2021. [Online]. Available: https://datasets-benchmarks-proceedings.neurips.cc/ paper/2021/file/c74d97b01eae257e44aa9d5bade97baf-Pap...

  14. [14]

    Multi- lingual deepfake speech dataset for robust and generalizable detection,

    C. O. Mawalim, Y . Wang, A. Adila, S. Okada, and M. Unoki, “Multi- lingual deepfake speech dataset for robust and generalizable detection,” IEEE Access, vol. 14, pp. 57 144–57 161, 2026

  15. [15]

    Simple and controllable music generation,

    J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y . Adi, and A. D´efossez, “Simple and controllable music generation,” inProceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY , USA: Curran Associates Inc., 2023

  16. [16]

    Audioldm: Text-to-audio generation with latent diffusion models,

    H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. P. Mandic, W. Wang, and M. D. Plumbley, “Audioldm: Text-to-audio generation with latent diffusion models,” 2023. [Online]. Available: https: //arxiv.org/abs/2301.12503

  17. [17]

    Audioldm 2: Learning holistic audio generation with self-supervised pretraining,

    H. Liu, Y . Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y . Wang, W. Wang, Y . Wang, and M. D. Plumbley, “Audioldm 2: Learning holistic audio generation with self-supervised pretraining,” 2023. [Online]. Available: https://arxiv.org/abs/2308.05734

  18. [18]

    Stable audio: Fast timing-conditioned latent audio diffusion,

    Stability AI, “Stable audio: Fast timing-conditioned latent audio diffusion,” Stability AI, Technical Report (web publication), Sep

  19. [19]

    Available: https://stability.ai/research/stable-audio-fast- timing-conditioned-latent-audio-diffusion

    [Online]. Available: https://stability.ai/research/stable-audio-fast- timing-conditioned-latent-audio-diffusion

  20. [20]

    Diffsound: Discrete diffusion model for text-to-sound generation,

    D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y . Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 1720–1733, 2023. [Online]. Available: https://dl.acm.org/doi/10.1109/TASLP.2023.3268730

  21. [21]

    Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,

    R. Huang, J. Huang, D. Yang, Y . Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,” inProceedings of the 40th International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol. 202, 2023, pp. 13 916–13 932. [Online]. Available: https...

  22. [22]

    Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis,

    H. K. Cheng, M. Ishii, A. Hayakawa, T. Shibuya, A. Schwing, and Y . Mitsufuji, “Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis,” 2025. [Online]. Available: https://arxiv.org/abs/2412.15322

  23. [23]

    Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,

    C.-Y . Hung, N. Majumder, Z. Kong, A. Mehrish, A. A. Bagherzadeh, C. Li, R. Valle, B. Catanzaro, and S. Poria, “Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,” 2025. [Online]. Available: https://arxiv.org/abs/2412.21037

  24. [24]

    Envsdd: Benchmarking environmental sound deepfake detection,

    H. Yin, Y . Xiao, R. K. Das, J. Bai, H. Liu, W. Wang, and M. D. Plumbley, “Envsdd: Benchmarking environmental sound deepfake detection,” 2025. [Online]. Available: https://arxiv.org/abs/2505.19203

  25. [25]

    Compspoof: A dataset and joint learning framework for component-level audio anti-spoofing countermeasures,

    X. Zhang, Y . Wang, L. Li, L. Jin, and M. Li, “Compspoof: A dataset and joint learning framework for component-level audio anti-spoofing countermeasures,” 2026. [Online]. Available: https: //arxiv.org/abs/2509.15804

  26. [26]

    Overview of esdd2: Environment-aware speech and sound deepfake detection challenge,

    X. Zhang, H. Yin, Y . Xiao, L. Zhang, T. Dang, R. K. Das, and M. Li, “Overview of esdd2: Environment-aware speech and sound deepfake detection challenge,” 2026. [Online]. Available: https://arxiv.org/abs/2606.10791

  27. [27]

    Efficient audio transformer and aasist for environment sound deepfake detection in the esdd 2026 challenge,

    J. Cao, C. Fan, J. Xue, Y . Xie, R. Fu, Z. Wen, J. Yi, Y . Ren, Z. Lv, and J. Tao, “Efficient audio transformer and aasist for environment sound deepfake detection in the esdd 2026 challenge,” inICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 21 781–21 783

  28. [28]

    Clotho: an audio captioning dataset,

    K. Drossos, S. Lipping, and T. Virtanen, “Clotho: an audio captioning dataset,” in2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8,

  29. [29]

    Barcelona, Spain: IEEE, May 2020, pp. 736–740. [Online]. Available: https://doi.org/10.1109/ICASSP40776.2020.9052990

  30. [30]

    TACOS: Temporally-aligned audio captions for language-audio pretraining,

    P. Primus, F. Schmid, and G. Widmer, “TACOS: Temporally-aligned audio captions for language-audio pretraining,” inIEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2025, Tahoe City, CA, USA, October 12-15, 2025. Tahoe City, CA, USA: IEEE, October 2025, pp. 1–5. [Online]. Available: https://doi.org/10.1109/W ASPAA66052.2025.11230997

  31. [31]

    WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

    X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y . Zou, and W. Wang, “WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,”IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 32, pp. 3339–3354, 2024. [Online]. Available: https://doi.org/10.1109/TASLP.2024.3419446

  32. [32]

    AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,

    J.-w. Jung, H.-S. Heo, H. Tak, H.-j. Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. W. D. Evans, “AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” in2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 6367–6371. [Online]. Available: https://dblp.org/rec/conf/icassp/JungH...

  33. [33]

    A dataset and taxonomy for urban sound research,

    J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” inProceedings of the 22nd ACM International Conference on Multimedia (MM ’14), 2014, pp. 1041–1044. [Online]. Available: https://dl.acm.org/doi/10.1145/2647868.2655045