{"id":"453a400b-98c5-4880-9180-736b1049cb22","arxiv_id":"2507.21463","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.","lead":"This paper introduces SpeechFake, a dataset of over 3 million fake audio clips in 46 languages, generated with 40 speech synthesis tools. It shows that models trained on this data detect unseen deepfakes far better than those trained on standard benchmarks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"External test sets are not methodologically unseen: WaveFake's generation methods all appear in SpeechFake training, so the claimed 50-80% 'unseen' EER gains may reflect seen-artifact matching.","rationale":"The reader's verdict identified representativeness of the generation toolset after filtering as the key assumption. My pass finds a more immediate and falsifiable issue: the external benchmarks used to claim 'unseen' generalization are not unseen in the dimension the claim requires. WaveFake in particular is constructed from exactly the vocoders included in SpeechFake's BD-NV training split (Table 8). A model trained on BD has therefore already learned artifacts of those generators; its low EER on WF is a measure of within-distribution performance, not cross-method generalization. This is a concrete correctness risk to the strongest claim, not a disagreement with consensus. The paper's own cross-generator results confirm that genuine method novelty degrades accuracy, so the external gains likely stem from method overlap. The suggested test is to tag external test utterances by the generating tool and to retrain on a version of BD that excludes overlapping methods, then compare EERs. If the gain persists on truly unseen tools, the claim survives; otherwise it should be weakened to 'generalization to unseen speakers/recording conditions under seen generators.' This does not diminish the dataset's potential utility, but it changes what Table 2 demonstrates. The verdict remains CONDITIONAL because the dataset's value and baseline results are still plausible, but the generalization claim needs re-evaluation under method-disjoint conditions.","tokens_in":16668,"tokens_out":7355,"duration_ms":81938,"concrete_test":"Compare the generation method list of each external test set (WaveFake, In-the-Wild, CD-ADD) against the 40 tools in Table 8. For WaveFake, verify whether its six vocoders (MelGAN, Multi-band MelGAN, Parallel WaveGAN, HiFi-GAN, WaveGlow, StyleMelGAN) are present in the BD training set; if so, retrain on BD excluding those vocoders (or evaluate on a hold-out of BD methods not in WaveFake) and recompute Table 2 EER for WF. If the EER improvement over ASV19 shrinks or disappears, the 'unseen' generalization claim is not supported. Repeat for CD-ADD and ITW by tagging fake utterances with known generation tools.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2 is presented as evidence that models trained on the Bilingual Dataset 'generalize substantially better to unseen deepfake test sets' (Section 4.2, abstract). This interpretation presupposes that the external test sets contain deepfake generation methods absent from the training data. That presupposition is violated at least for WaveFake. WaveFake consists of audio generated by six neural vocoders: MelGAN, Multi-band MelGAN, Parallel WaveGAN, HiFi-GAN, WaveGlow, and StyleMelGAN. All six appear in SpeechFake's BD-NV training portion (Table 8: MelGAN, WaveGlow, Parallel WaveGAN, HiFi-GAN, Fullband-MelGAN, StyleMelGAN). Hence the WF column in Table 2 measures recognition of the same artifact-generating architectures, not generalization to an unseen method. The same issue may affect CD-ADD and In-the-Wild; the paper never reports which external test-set methods overlap with its 40 tools. The authors' own cross-generator experiments (Section 4.3, Table 3) show that truly unseen categories cause large EER degradations, so the external improvements are consistent with seen-method matching. Without a method-disjoint evaluation, the strongest claim overstates what the data demonstrate.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SpeechFake, a large-scale multilingual speech deepfake dataset containing over 3 million fake utterances (more than 3,000 hours) generated by 40 speech synthesis tools, split into a Bilingual Dataset (English/Chinese) and a Multilingual Dataset (46 languages). The authors provide a detailed description of data collection, statistics, and metadata, and present baseline experiments with AASIST and W2V+AASIST. The central claim is that models trained on SpeechFake generalize substantially better to unseen deepfake test sets than models trained on ASVspoof2019, supported by large EER reductions on WaveFake, In-the-Wild, and CD-ADD, along with analyses of cross-generator, cross-lingual, and cross-speaker robustness.","tokens_in":16965,"tokens_out":4595,"duration_ms":53691,"significance":"If the dataset is released and the generalization claims hold, SpeechFake would be a valuable community resource due to its scale, diversity of generation methods (including recent TTS/VC/NV tools), multilingual coverage, and rich metadata. The paper includes several controlled experiments (cross-generator, cross-lingual, cross-speaker) and provides training details in the appendix. The external generalization result is important for deepfake detection research. However, the central 'unseen generalization' claim currently rests on external test sets whose method overlap with the training data is not analyzed, and the headline numbers lack error bars; these issues need to be addressed before the claims are fully supported.","major_comments":[{"comment":"The claim that models trained on SpeechFake generalize to 'unseen' test sets is not fully supported for WaveFake, because the WaveFake test set contains six neural vocoders (MelGAN, Multi-band MelGAN, Parallel WaveGAN, HiFi-GAN, WaveGlow, StyleMelGAN; see Frank and Schönherr, 2021), and five of those architectures (MelGAN, Parallel WaveGAN, HiFi-GAN, WaveGlow, StyleMelGAN; Table 8) also appear in the SpeechFake BD training data. The WaveFake column therefore largely measures recognition of seen artifact-generating architectures, not generalization to a truly unseen method. Please provide a systematic overlap analysis for all external test sets (WF, ITW, CDADD, ASV24, FOR, CFAD) and either remove overlapping generators from training, evaluate on per-method subsets, or explicitly qualify the 'unseen' wording.","section":"Section 4.2 / Table 2 / Abstract"},{"comment":"All headline EER results in Tables 2 and 3 come from a single training run (Table 6 reports '1 run' for the main experiments). The cross-speaker experiments in Section 4.5 do report standard deviations across three runs, so the infrastructure to report means and variances exists. Given that some external-test EER differences (e.g., the W2V+AASIST row in Table 2) are small, single-run numbers are insufficient to establish the claimed ranking. Please report multiple-seed averages with standard deviations for the main comparisons, or provide a stability analysis.","section":"Section 4.1 / Table 6 and Tables 2-3"},{"comment":"The main BD train/dev/test partition is not documented as speaker-disjoint for either real or fake data. The dedicated cross-speaker trials in Section 4.5 (where the authors explicitly construct seen/unseen speaker settings) suggest that the default partition does not guarantee speaker separation. If speakers overlap between training and the BD test sets, internal EERs such as BD 3.48 and BD-EN 3.98 in Table 2 may be optimistic, and cross-lingual or cross-generator comparisons could be confounded by speaker identity. Please clarify whether speakers are disjoint in BD and MD splits; if they are not, re-run the affected experiments with a speaker-disjoint split or state this limitation explicitly.","section":"Section 4.2 / Table 7 and Section 4.5"},{"comment":"The quality-filtering step ('a selective human review is conducted to discard generated speech with noticeable distortions, excessive noise, or unnatural artifacts') is described only qualitatively, without specifying the number of human reviewers, the sampling procedure beyond 'approximately 1% of samples from each method', or any inter-annotator agreement. The paper also does not test whether this filtering systematically biases the dataset toward easy-to-detect utterances, which would weaken the usefulness of the dataset for training robust detectors. Please provide more detail on the human review protocol and, ideally, an analysis of how filtering affects detection difficulty.","section":"Section 3.1 (Data Post-processing)"}],"minor_comments":[{"comment":"The explanation for NV underperformance is unclear: the text says NV models 'often rely on older methods that generate lower-quality deepfakes, making detection more challenging for models trained on NV data.' Lower-quality synthesized speech should generally make fake samples more distinguishable; please clarify the intended mechanism (e.g., that the model learns low-quality artifacts that do not transfer to high-quality TTS).","section":"Section 4.3"},{"comment":"The phrase '50%-80% better performance' is ambiguous; Table 2 reports EER reductions, so please define the relative improvement metric (e.g., relative EER reduction) and state it explicitly.","section":"Section 4.2"},{"comment":"The 'Generator Types' column uses inconsistent separators (e.g., 'TTS, VC' vs 'TTS, VC, AT'); please standardize the formatting and verify the counts for ASVspoof2015 and other rows.","section":"Table 1"},{"comment":"The abstract uses 'voice id' while the metadata section uses 'Speaker/Voice ID'; please unify the terminology across the paper.","section":"Section 3.1 / Appendix A.2"},{"comment":"The caption states 'Eight latest generation methods were selected' but the figure is not visible in the provided text; please ensure the figure is included and legible, and specify which methods are shown.","section":"Appendix B.2 / Figure 5"},{"comment":"Some references have informal author names (e.g., 'gil Lee et al., 2023' for BigVGAN) and incomplete entries; please correct the bibliography.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the typical scope of a speech/audio processing journal and the dataset itself is a potentially significant contribution. The main concern is that the 'unseen generalization' claim is overstated for WaveFake due to training/test method overlap, and the lack of speaker-disjoint splits and error bars further weakens the headline results. These are fixable with additional analysis and reruns, so major revision rather than rejection seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing here is the dataset itself: three million-plus samples, 40 generators, 46 languages, with metadata that lets you slice by method, language, and speaker. That is a real resource for the audio deepfake detection community, and the baseline experiments are a solid starting point. The cross-generator and cross-lingual analyses are worth reading; they confirm that method mismatch hurts, and that multilingual pretraining helps.\n\nThe soft spot is the strongest claim. The paper says models trained on SpeechFake generalize to 'unseen' test sets, with 50-80% EER gains on WaveFake, In-the-Wild, and CD-ADD. But WaveFake's six vocoders are all present in SpeechFake's own BD-NV training portion (Table 8). So the WaveFake result is not evidence of generalization to an unseen method; it is recognition of the same artifact-generating architectures. The paper never reports which external methods overlap with its 40 tools, and its own cross-generator table shows truly unseen categories cause large degradations. That suggests the external improvements are largely seen-method matching. The claim should be reworded or the evaluation should use method-disjoint splits.\n\nOther issues are minor but real: the main baselines are single-run without error bars; the BD test sets are not speaker-disjoint from training; and the human-review filtering may remove hard-to-detect samples, which could inflate measured performance. The cross-speaker conclusion is based on one TTS system and a small set, so it is narrower than the conclusion section implies.\n\nIf these are fixed—an explicit method-overlap table, speaker-disjoint splits, error bars, and a toned-down generalization claim—this becomes a very solid dataset paper. As it stands, it is a useful contribution with an overreaching headline result.\n\nI'd send it to review; the dataset deserves the community's eyes, and the methodological overclaim is fixable. For a reading group, it is a good example of how dataset papers can overstate 'unseen' generalization. I'd cite it if I work in this area.","headline":"SpeechFake is a genuinely useful dataset resource, but the paper oversells its external generalization results because the 'unseen' test sets are not method-disjoint from training.","tokens_in":17442,"tokens_out":2342,"would_cite":true,"duration_ms":24509,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SpeechFake, a 3-million-clip multilingual dataset built from 40 speech-generation tools, makes trained detectors generalize to unseen deepfake methods far better than ASVspoof2019 training does, with 50–80% lower error rates on several…","keywords":["speech deepfake detection","deepfake dataset","text-to-speech","voice conversion","neural vocoder","multilingual speech","generalization","equal error rate"],"falsifier":"Take a commercial TTS API that is not among the 40 tools used in SpeechFake, generate a held-out test set with no quality filtering, and compare equal error rates of models trained on SpeechFake-BD and on ASVspoof2019; if the SpeechFake-trained model's relative EER reduction on this fresh generator is much smaller than the 50–80% range reported on WaveFake, In-the-Wild, and CD-ADD, the generalization claim is not general.","tokens_in":16491,"feed_emoji":"🎙️","tokens_out":15372,"duration_ms":140290,"temperature":0.7,"pith_summary":"This paper introduces SpeechFake, a public dataset of more than three million fake speech clips and over three thousand hours of audio, produced with 40 speech-synthesis tools and covering 46 languages, and argues that training on this diversity makes deepfake detectors generalize to generators they have never seen. The core evidence is a side-by-side comparison in which AASIST and W2V+AASIST models trained on SpeechFake's Bilingual Dataset consistently outperform the same models trained on the standard ASVspoof2019 benchmark, with relative equal-error-rate (EER) reductions of 50–80% on the unseen WaveFake, In-the-Wild, and CD-ADD test sets. Supporting experiments separate the axes that drive the gain: modern text-to-speech data transfers best to unseen commercial TTS APIs, multilingual pretraining nearly closes the language gap, and speaker identity contributes little. If the results hold, the dataset provides a training resource that is better matched to the current generation of synthesis tools than existing benchmarks.","feed_headline":"SpeechFake training cuts unseen deepfake errors by 50–80%","feed_subtitle":"A 3-million-clip, 40-tool, 46-language dataset trains detectors that beat the ASVspoof benchmark on new generators.","key_machinery":"The load-bearing object is the dataset itself, and specifically its controlled diversity. SpeechFake is partitioned into a Bilingual Dataset (English and Chinese) and a Multilingual Dataset (46 languages), and every fake clip carries metadata for generation method, voice/speaker ID, language, and text transcription, so the training distribution can be matched to or deliberately separated from the test distribution. A three-way generator taxonomy (TTS, VC, NV, defined by input modality at inference) lets the authors construct train/test splits that isolate whether a generation method was seen during training. The evaluation machinery is paired training of two standard detectors—AASIST, a spectro-temporal graph attention network, and W2V+AASIST, the same network with a Wav2Vec 2.0 XLS-R frontend—measured by equal error rate across a common set of external benchmarks.","core_discovery":"SpeechFake is a large-scale, public, multilingual benchmark built to reflect current speech generation rather than the older and narrower method mix of earlier datasets. Its Bilingual Dataset contains 2,003,016 fake utterances in English and Chinese, generated by 30 open-source tools and 10 commercial APIs spanning text-to-speech, voice conversion, and neural vocoders; its Multilingual Dataset adds 1,335,492 utterances across 46 languages using six multilingual generators. The paper's central discovery is that detection models trained on this diversity generalize to unseen generators: on the WaveFake, In-the-Wild, and CD-ADD test sets, SpeechFake-trained models beat ASVspoof2019-trained models by 50–80% relative EER, and the W2V+AASIST variant stays competitive on ASVspoof2019's own evaluation set. The authors then use cross-generator, cross-lingual, and cross-speaker splits to show that generation method and language diversity are the factors that matter, while speaker identity has only a small effect.","pith_inferences":["Because the ten commercial APIs appear only in the unseen test set, SpeechFake defines a deployment-like protocol: a detector is trained entirely on open-source tools and then faces proprietary synthesis, which is closer to real-world conditions than a random split within one benchmark.","If the generalization gains are driven by the diversity and recency of the generator set rather than dataset size alone, then periodically refreshing the listed tools will matter as much as adding more hours; the paper's own limitations section acknowledges that the method coverage is incomplete and time-sensitive.","The cross-lingual result suggests that collecting fake audio in every target language may be unnecessary if a multilingual self-supervised encoder is used, which could lower the cost of extending detection to low-resource languages.","A natural next experiment the paper does not run is a direct comparison against other large recent datasets such as ASVspoof5 under equal training budgets, which would separate the contribution of dataset composition from total scale."],"forward_implications":["Training on SpeechFake's Bilingual Dataset gives 50–80% relative EER improvements over ASVspoof2019 training on the unseen WaveFake, In-the-Wild, and CD-ADD test sets, so large, method-diverse training data is a direct route to better cross-generator detection.","Models trained only on TTS data transfer best to the held-out commercial API set, indicating that exposure to modern text-to-speech output is the most useful ingredient for finding commercially generated deepfakes.","Language mismatch hurts detection even when generation methods are seen in training, but a multilingual self-supervised frontend (XLS-R) almost closes the gap, pointing to a practical recipe for multilingual detection.","Cross-speaker experiments on one multi-speaker TTS system show that detectors learn deepfake-specific artifacts rather than memorizing speaker identity; only completely unseen fake speakers cause a small EER increase.","The rich per-sample metadata lets future studies control for generation method, language, and speaker independently instead of treating them as confounded variables."],"supporting_citations":[{"why":"Provides the ASVspoof2019-LA train/eval sets that serve as the standard benchmark baseline and the main comparison condition.","marker":"Nautsch et al. (2021)"},{"why":"WaveFake is one of the three unseen test sets where SpeechFake-trained models achieve 50–80% relative EER improvement.","marker":"Frank and Schönherr (2021)"},{"why":"In-the-Wild is one of the three unseen test sets in the main generalization comparison.","marker":"Müller et al. (2022)"},{"why":"CD-ADD is the third unseen test set in the main generalization comparison.","marker":"Li et al. (2024c)"},{"why":"Defines the AASIST detector used for all baseline and analysis experiments.","marker":"Jung et al. (2022)"},{"why":"Defines the W2V+AASIST model and the training configuration the experiments follow.","marker":"Tak et al. (2022)"},{"why":"Supplies the XLS-R multilingual self-supervised representation that drives the strong cross-lingual results.","marker":"Babu et al. (2021)"},{"why":"ASVspoof5 is an additional unseen test set and a scale/composition comparison point in the dataset survey.","marker":"Wang et al. (2024)"},{"why":"Common Voice is the source of real speech for the 46-language Multilingual Dataset.","marker":"Ardila et al. (2019)"},{"why":"TorToiSe is the TTS system used in the cross-speaker generalization experiments.","marker":"Betker (2023)"}],"fun_headline_variants":["SpeechFake cuts unseen deepfake errors by up to 80%","3M deepfakes, 40 tools: dataset boosts detection","46-language deepfake corpus slashes new generator misses","SpeechFake: training on diverse fakes generalizes better"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 40 synthesis tools, after voice-activity filtering and review of roughly 1% of samples, produce fake audio whose artifacts are representative of what detectors will meet in practice, and if important synthesis channels are missing or hard cases are filtered out, the measured generalization gains may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["SpeechFake cuts unseen deepfake errors by up to 80%","3M deepfakes, 40 tools: dataset boosts detection","46-language deepfake corpus slashes new generator misses","SpeechFake: training on diverse fakes generalizes better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1344,"prompt_tokens":1002,"completion_tokens":342,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":270}},"tokens_in":618,"tokens_out":342,"duration_ms":4625,"temperature":1.0,"reasoning_tokens":270,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:43:37.398269+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a commercial TTS API that is not among the 40 tools used in SpeechFake, generate a held-out test set with no quality filtering, and compare equal error rates of models trained on SpeechFake-BD and on ASVspoof2019; if the SpeechFake-trained model's relative EER reduction on this fresh generator is much smaller than the 50–80% range reported on WaveFake, In-the-Wild, and CD-ADD, the generalization claim is not general.","supporting_citations":[],"review_version":1}