{"id":"22a8cabc-0e07-4254-ae52-f15dfed6790e","arxiv_id":"2412.14890","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Speech enhancement quality scales with speaker and noise diversity in training data, not with text or language diversity.","lead":"This paper builds training sets for speech enhancement with a text-to-speech engine, changing only one attribute at a time: the words, the language, the speaker, or the background noise. The authors find that adding more speakers and more noise types helps the models most, while adding more sentences or languages helps little, which could make future data collection cheaper and more targeted.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The language-independence result may stem from English-prompted synthetic non-English speech, and the paper's own Figure 2 already shows SGMSE generalizes poorly from English to Chinese; the synthetic proxy for language scaling needs validation.","rationale":"The reader's weakest_assumption identifies exactly this risk: XTTS-generated speech may not preserve real-data scaling behavior, and the use of English speaker prompts for non-English languages could attenuate the true language effect. Our concern narrows this to the language attribute specifically, where the threat is most concrete and the evidence is weakest. Section III-A states that all speaker prompts come from LibriSpeech, so non-English utterances are English-voice clones; Section IV-A validates synthetic data only on LibriMix, an English test set, leaving language transfer unverified; and Section IV-C/Figure 2 already contains a counterexample to 'largely language-independent' for SGMSE. Because the headline recommendation—spend scaling budgets on speaker and noise rather than language—depends on this flat language curve, the concern is load-bearing. The proposed test with native speaker prompts for each language is a direct, feasible check: if the scaling curves diverge between English-prompted and native-prompted synthetic data, the central claim needs to be weakened to 'acoustic attributes are more important than semantic attributes among the synthetic conditions tested here.' Since the current verdict is already CONDITIONAL, our analysis supports keeping that verdict rather than moving it; no change is needed.","tokens_in":8125,"tokens_out":6077,"duration_ms":54951,"concrete_test":"Re-run the language-scaling study L_p with two matched synthetic conditions: (A) speaker prompts from LibriSpeech English speakers exactly as in the paper, and (B) native speaker prompts from CommonVoice for each target language, holding transcriptions, noise, duration, and speaker count per language fixed; train BSRNN and SGMSE in both conditions and compare multilingual test metrics and scaling slopes across p. If condition (B) yields better non-English test scores or a non-flat language-scaling curve where condition (A) was flat, the paper's language-independence conclusion fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central ranking ('acoustic >> semantic') depends on language and text scaling curves measured on XTTS-synthesized data. The generation protocol (Section III-A) uses speaker prompts extracted from LibriSpeech for every utterance, including Chinese, Czech, German, and other languages. XTTS therefore clones English voices to speak non-English text. If cross-lingual synthesis yields accented or artifact-prone audio, or flattens genuine phonotactic differences, the trained SE models will appear language-insensitive even when real multilingual data would matter. The only synthetic-versus-real validation (Section IV-A, Table II) compares average quality on the English LibriMix test set, not per-attribute scaling slopes, so it cannot detect attenuation of the language effect. Moreover, the paper's own Figure 2 reports that an SGMSE model trained only on English generalizes poorly to Chinese, which is hard to reconcile with the abstract's 'largely language-independent' claim and shows that language can matter at least for the generative model. The abstract's 'much more important' conclusion therefore outruns the available evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a generation-training-evaluation framework that uses a zero-shot multilingual TTS system (XTTS) to synthesize speech enhancement training corpora in which text, language, speaker, and noise attributes are varied independently. The authors train two representative SE models (BSRNN and SGMSE) on these controlled synthetic datasets and evaluate them on LibriMix plus multilingual out-of-domain test sets. Their main empirical finding is that acoustic attributes (speaker and noise) matter much more than semantic attributes (text and language) for current SE models, and they conclude that data scaling budgets should prioritize speaker and noise diversity. The paper also reports that purely synthetic TTS-based training data performs comparably to real speech data.","tokens_in":8381,"tokens_out":3330,"duration_ms":32131,"significance":"If the central ranking of attribute importance holds, the paper provides actionable guidance for SE data collection and augmentation: spend scaling resources on speaker and noise diversity rather than on text or language coverage. The study is valuable for its controlled experimental design, the use of two model families (discriminative and generative), evaluation on external multilingual and out-of-domain noise sets, and the plan to open-source the generation code. The result that purely synthetic TTS data can train competitive SE models is itself a useful contribution. However, the headline claim that models are 'largely text- and language-independent' is only as strong as the synthetic proxy used to measure language scaling, and the paper's own cross-lingual transfer results complicate the claim. The lack of repeated runs also limits the certainty of comparisons between flat and rising scaling curves.","major_comments":[{"comment":"The language manipulation uses speaker prompts extracted from LibriSpeech for every utterance, including non-English languages. This means Chinese, Czech, German, and other languages are spoken by cloned English voices, which can introduce non-native accents or TTS artifacts that flatten genuine phonological and phonetic differences across languages. The only synthetic-to-real validation (Section IV-A, Table II) compares average LibriMix quality between models trained on real versus synthetic data; it does not validate that language scaling slopes measured on synthetic data transfer to real multilingual data. To support the abstract's claim that language is much less important than acoustic attributes, the authors should either validate the language scaling with native-prompt or real multilingual training data, or substantially soften the claim.","section":"Section III-A, Table I, Section IV-A"},{"comment":"Figure 2 shows that SGMSE trained only on English generalizes poorly to Chinese, and the text states that 'the performance of the generative model varies based on the training language.' This is difficult to reconcile with the abstract's statement that models are 'largely language-independent.' At minimum, the paper should distinguish between (a) adding language diversity during training and (b) zero-shot cross-lingual transfer, and should explain why poor English-to-Chinese transfer does not contradict the claimed unimportance of the language attribute. As written, the paper's own evidence suggests that language can matter for generative models.","section":"Section IV-C, Figure 2"},{"comment":"Every data point in the scaling curves comes from a single training run with no error bars, repeated seeds, or significance tests. The central conclusion relies on distinguishing flat curves (text, language) from increasing curves (speaker, noise), and on small differences in some conditions. Without variance estimates, it is impossible to assess whether the observed flatness is meaningful or within run-to-run noise. The authors should provide at least three seeds for key comparisons or report confidence intervals.","section":"Figures 1-4"},{"comment":"The speaker diversity experiments keep the total number of utterances m fixed while varying the number of speakers s. Consequently, increasing s reduces the number of utterances per speaker, confounding speaker diversity with the degree of repeated exposure to each speaker's voice. The observed saturation beyond 100 speakers may reflect per-speaker data scarcity rather than a true limit of speaker diversity. This confound should be controlled or explicitly discussed.","section":"Section III-A, Figure 1(e,f)"}],"minor_comments":[{"comment":"The caption contains a typo: 'Evaluaion' should be 'Evaluation.'","section":"Figure 2 caption"},{"comment":"The abbreviations in the table, such as 'W,T' and '#NT', are not defined in the caption; please define them in a footnote or in the table caption.","section":"Table I"},{"comment":"The sentence 'we may conclude that it is safe to scale the SE training data by introducing new languages' is too strong given the cross-lingual transfer results in Figure 2; a more cautious formulation would reflect the observed discrepancy between discriminative and generative models.","section":"Section IV-C"},{"comment":"The contribution list says the analysis reveals that models are 'largely text- and language-independent,' but this phrasing is already an interpretation before the experimental results are presented; consider rephrasing to state the finding after the experiments.","section":"Section I"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-designed empirical study with a clear potential impact on SE data curation. The main risk is the unvalidated synthetic proxy for language scaling, which is load-bearing for the headline claim. I believe the concern is addressable with additional experiments or a more nuanced claim, so the appropriate decision is major revision rather than rejection. The lack of repeated runs is also a significant issue for a scaling analysis, but it is fixable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is worth a look, but the headline needs tempering. The paper's real contribution is the framework: using zero-shot TTS to synthesize SE training sets where text, language, speaker, and noise are varied one at a time. That is genuinely new in this literature, which has mostly studied data size and model complexity. The finding that speaker and noise-type diversity drive performance more than text recycling is plausible and practically useful; the noise-duration-vs-noise-type contrast is clean and insightful.\n\nThe soft spots are where the strong claims live. First, every curve in Figures 1-4 is a single training run, no seeds, no error bars. The differences between conditions are often small, so without a sense of variance the ranking is fragile. Second, the language manipulation uses English speaker prompts from LibriSpeech for every language. XTTS clones English voices to speak Chinese, Czech, etc. That will likely flatten phonotactic and prosodic differences and may systematically attenuate the language effect. The synthetic-vs-real validation on LibriMix cannot catch that. And the paper's own Figure 2 shows SGMSE trained on English transfers poorly to Chinese — the authors acknowledge this contradiction but still stand behind a largely language-independent conclusion. That is an internal tension that needs resolution, not just a footnote.\n\nI would send this to peer review. The framework is reusable, the experiments are thoughtfully designed, and the practical guidance on data budgets is valuable even if it only applies to the attribute ranges tested. But the authors should be asked to add multiple seeds, to re-run the language condition with native-speaker prompts for at least a few languages, and to spend more effort validating that synthetic scaling slopes match real-data slopes. The central claim should be softened to something like 'acoustic attributes, especially noise type, dominate in the ranges tested on these two models.'\n\nFor a reading group, it is a good discussion piece about synthetic data and scaling. I would not cite the headline conclusion uncritically, but I would cite the framework.","headline":"A useful controlled-study framework for SE data scaling, but the 'language-independence' claim is partly an artifact of English-voice TTS prompts and single runs.","tokens_in":8812,"tokens_out":1915,"would_cite":true,"duration_ms":15855,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that when scaling training data for speech enhancement, acoustic attributes such as speaker and noise diversity matter far more than semantic attributes such as language and text, and demonstrates this with a…","keywords":["speech enhancement","dataset scaling","synthetic speech","zero-shot text-to-speech","speaker diversity","noise diversity","language variability","text variability"],"falsifier":"Run the same language-scaling experiment with real speech from native speakers recorded in each language while holding speaker identity and noise distribution fixed, and test on held-out languages: if PESQ, STOI, SDR, or DNSMOS drop substantially as the number of languages increases from one to ten, the language-independence claim would be contradicted. A complementary check is to compare the performance of models trained on a real multilingual corpus against models trained on the synthetic corpus under identical evaluation; a large gap would indicate that the synthetic proxy is not faithful.","tokens_in":7870,"feed_emoji":"🎙️","tokens_out":7139,"duration_ms":51267,"temperature":0.7,"pith_summary":"This paper argues that when scaling up training data for speech enhancement, not all diversity is equal: acoustic variation in speakers and noise matters far more than semantic variation in text and language. To test this, the authors build a generation-training-evaluation pipeline using a multi-lingual zero-shot text-to-speech model, synthesizing training corpora in which only one attribute (text, language, speaker, or noise) changes at a time. They train both a discriminative (BSRNN) and a generative (SGMSE) enhancement model on each controlled corpus and evaluate on real multilingual test sets. Their results suggest that models trained on a single sentence or a single language perform nearly as well as models trained on rich text and language diversity, whereas reducing speaker or noise diversity hurts performance. If correct, this gives data collectors a concrete priority: spend scaling budgets on speaker and noise coverage, not on language or textual breadth.","feed_headline":"Acoustic diversity beats language in speech enhancement scaling","feed_subtitle":"Speech enhancement models gain the most from more speakers and noise types, not more text or languages.","key_machinery":"The load-bearing mechanism is a generation-training-evaluation pipeline built on a pre-trained multi-lingual zero-shot text-to-speech model. The TTS model takes a short speaker prompt and a text transcription and produces speech in a chosen language with a controlled speaker identity, so the authors can synthesize corpora in which only the target attribute varies. They generate datasets varying the number of unique transcriptions, the number of languages, the number of speakers (in single-prompt and multi-prompt modes), and the noise duration or noise-type count, then train BSRNN and SGMSE on fixed noisy-clean mixtures simulated from paired clean speech and environmental noise. The comparison of model performance across these controlled corpora is what attributes the observed scaling effects to the manipulated attribute rather than to confounded differences in content.","core_discovery":"The central claim is that current speech enhancement models are largely text- and language-independent while being sensitive to speaker and noise diversity. Using purely synthetic speech generated by a zero-shot text-to-speech model, the authors manipulate one dataset attribute at a time while holding total duration and word counts roughly constant. Across both a discriminative model (band-split RNN) and a generative model (diffusion-based), they find that collapsing textual diversity to a single sentence costs little in PESQ, STOI, SDR, and DNSMOS on in-domain and out-of-domain multilingual evaluations, and that adding languages up to ten languages does not improve or degrade performance much. In contrast, increasing the number of speakers and the variety of speaker prompts improves enhancement, and increasing noise type diversity helps generalization to unseen noise, especially for the discriminative model. The finding is framed as guidance for efficient dataset scaling: spend resources on acoustic attribute diversity first.","pith_inferences":["Editorial inference: the language-independence result may be specific to speech enhancement, which does not need to understand content; tasks like ASR or translation would likely show a much larger language and text effect, so the ranking of attributes should not be transferred across tasks.","Editorial inference: because the multilingual synthetic speech was produced from English speaker prompts, the language comparison may understate acoustic differences between languages; a test using native speaker prompts per language would clarify whether the TTS proxy masked a real language effect.","Editorial inference: a natural next experiment is to apply the same controlled-generation framework to real human speech where possible, or to measure how far the synthetic-to-real gap grows as TTS quality degrades, since the entire argument depends on TTS preserving real-data scaling behavior.","Editorial inference: the saturation of speaker gains beyond 100 speakers in this setup may reflect the fixed total data size; with a larger utterance budget, speaker gains might continue further, so the '100 speakers' number should not be read as a universal ceiling."],"forward_implications":["Data scaling budgets for speech enhancement should prioritize adding speakers and noise types over adding text or language coverage.","Training on synthetic speech is a viable proxy for real speech when studying scaling laws, at least for the models and test conditions examined.","A model trained almost entirely on a single-sentence, single-language corpus can generalize to multilingual, varied-content test conditions, so small-domain synthetic corpora may be sufficient for many enhancement deployments.","Improving noise-type diversity in training data is a more effective route to generalization on unseen noise than simply increasing noise duration.","For generative models, the effect of speaker and noise diversity is visible but weaker than for discriminative models, implying separate scaling strategies for the two model families."],"supporting_citations":[{"why":"Supplies the multi-lingual zero-shot TTS model used for all synthetic speech generation.","marker":"[10]"},{"why":"Establishes the scalability context and motivates the need to study dataset variability in speech enhancement.","marker":"[4]"},{"why":"Defines the band-split RNN architecture used as the discriminative model.","marker":"[19]"},{"why":"Defines the diffusion-based architecture used as the generative model.","marker":"[20]"},{"why":"Provides the fixed noisy-clean mixture simulation pipeline used to pair synthetic speech with noise.","marker":"[24]"},{"why":"Supplies the real speech corpus used as the baseline and as the source of speaker prompts and transcriptions.","marker":"[25]"},{"why":"Supplies one environmental noise corpus used for training and evaluation.","marker":"[26]"},{"why":"Supplies another environmental noise corpus used for training and evaluation.","marker":"[27]"},{"why":"Supplies multilingual speech used for evaluation and for non-English transcriptions.","marker":"[28]"}],"fun_headline_variants":["Speech enhancement scales with speakers and noise, not language","Acoustic diversity beats text for scaling speech enhancement","For speech enhancement, add speakers and noise, not words","Synthetic data shows acoustic attributes matter most","Speaker and noise variety trump language in speech enhancement"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole ranking of attributes rests on whether speech produced by a synthetic voice behaves like real speech for scaling experiments; if TTS artifacts or non-native voice prompts distort the true language or speaker effects, the ranking could change.","fun_headline_variants_meta":{"raw":{"variants":["Speech enhancement scales with speakers and noise, not language","Acoustic diversity beats text for scaling speech enhancement","For speech enhancement, add speakers and noise, not words","Synthetic data shows acoustic attributes matter most","Speaker and noise variety trump language in speech enhancement"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00016,"raw_usage":{"total_tokens":1211,"prompt_tokens":902,"completion_tokens":309,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":236}},"tokens_in":518,"tokens_out":309,"duration_ms":3170,"temperature":1.0,"reasoning_tokens":236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:48:57.098119+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same language-scaling experiment with real speech from native speakers recorded in each language while holding speaker identity and noise distribution fixed, and test on held-out languages: if PESQ, STOI, SDR, or DNSMOS drop substantially as the number of languages increases from one to ten, the language-independence claim would be contradicted. A complementary check is to compare the performance of models trained on a real multilingual corpus against models trained on the synthetic corpus under identical evaluation; a large gap would indicate that the synthetic proxy is not faithful.","supporting_citations":[{"cited_title":"XTTS: A massively multilingual zero-shot text-to-speech model,","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-lingual zero-shot TTS model used for all synthetic speech generation."},{"cited_title":"Beyond performance plateaus: A comprehensive study on scalability in speech enhancement,","cited_arxiv_id":null,"evidence_quote":"Establishes the scalability context and motivates the need to study dataset variability in speech enhancement."},{"cited_title":"Music source separation with band-split RNN,","cited_arxiv_id":null,"evidence_quote":"Defines the band-split RNN architecture used as the discriminative model."},{"cited_title":"WHAM!: Extending speech separation to noisy environments,","cited_arxiv_id":null,"evidence_quote":"Supplies one environmental noise corpus used for training and evaluation."},{"cited_title":"TUT database for acoustic scene classification and sound event detection,","cited_arxiv_id":null,"evidence_quote":"Supplies another environmental noise corpus used for training and evaluation."},{"cited_title":"Common voice: A massively-multilingual speech corpus,","cited_arxiv_id":null,"evidence_quote":"Supplies multilingual speech used for evaluation and for non-English transcriptions."}],"review_version":1}