{"id":"271e9718-4760-415f-a827-94de3e8faf21","arxiv_id":"2412.12512","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A TSE dataset using clean LibriTTS targets, noisy VoxCeleb2 interference, synthetic speaker augmentation and curriculum learning reports iSDR gains of 1.39 dB and 0.78 dB on Libri2Talker and Libri2Vox test sets.","lead":"Libri2Vox is a new training dataset for target speaker extraction that pairs clean LibriTTS target speech with noisy VoxCeleb2 interference speech and adds synthetic speakers plus curriculum learning. The authors report signal-to-distortion ratio gains of up to 1.39 dB over baseline training and 0.78 dB from curriculum learning with synthetic speakers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Libri2Talker gains are not isolated from data volume, and the synthetic/CL gains are only shown on the Libri2Vox test set.","rationale":"The reader's weakest assumption was that VoxCeleb2-based training generalizes to Libri2Talker despite the negative Libri2Vox-only iSDR. I agree that domain mismatch is real, but the more load-bearing issue is that the experimental design cannot attribute the reported Libri2Talker gain to the proposed dataset properties. The 1.39 dB improvement is measured by adding Libri2Vox to Libri2Talker, which simultaneously increases training data size, speaker count, and noise realism; without a matched clean-data control, the gain could simply come from having more data and more speakers. This matters because the paper's core claim is that realistic noisy interference from VoxCeleb2 helps. The second headline number is even more fragile: it is computed on the Libri2Vox test set, which is drawn from the same target/interference pools as the training data, and the synthetic-only increment over the real-only CL control is 0.19 dB, not 0.78 dB. I am not claiming the paper is wrong; I am claiming the current evidence does not rule out a data-volume or in-domain-matching explanation. If the proposed clean-data control reproduces the gain, the concern is resolved and the dataset contribution is strengthened. Because the paper already ships no data or code and reports no error bars, these confounds keep it at conditional rather than accept. I do not see grounds to reject: the experiments are systematic, multi-architecture, and the direction of the results is internally consistent; the weaknesses are addressable with additional controls and cross-test evaluation.","tokens_in":18431,"tokens_out":10666,"duration_ms":91826,"concrete_test":"Run one experiment with two arms: (a) train SpeakerBeam on Libri2Talker plus a same-size, same-speaker-count clean interference set from LibriTTS/LibriSpeech with the same SNR distribution and noise augmentation; (b) evaluate the Table IV Conformer (3-stage CL, Real+SALT) on Libri2Talker. If arm (a) matches the 7.65 dB joint-training result within seed noise, and arm (b) does not beat the 9.58 dB baseline, then the claimed benefits are confounded or domain-specific.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Two headline results support the paper's central claim, and neither isolates the proposed mechanism.\n\nFirst, the abstract's 1.39 dB improvement for SpeakerBeam on Libri2Talker (Table III: 7.65 vs 6.26 dB, both with noise augmentation) comes from training on Libri2Talker plus Libri2Vox jointly. This condition adds about 150k mixtures and roughly 5,900 additional interference speakers while also introducing VoxCeleb2's real-world noise. There is no control that adds an equal volume of clean mixtures with a matched speaker count. The reported gain can therefore be explained by data scale or speaker diversity alone, not by realistic noisy interference. The negative Libri2Vox-only iSDR on Libri2Talker (-1.28 for Conformer) confirms a substantial domain gap and shows why joint training is needed, but it does not identify which component of the joint training is responsible.\n\nSecond, the 0.78 dB improvement for Conformer with curriculum learning and synthetic speakers is reported only on the Libri2Vox test set (Table IV: 16.20 vs 15.42 dB). The abstract presents it as a continuation of the Libri2Talker result, yet no synthetic/CL model is evaluated on Libri2Talker. Given the domain mismatch in Table III, there is no basis to assume this gain transfers. Moreover, the comparison against 'w/ 3-stage CL (Real only)' shows a synthetic-only increment of 0.19 dB (16.20 vs 16.01); the headline 0.78 dB bundles curriculum learning with synthetic augmentation. Section VI-A states three seeds were averaged, but no variance or significance is reported, so differences of 0.02-0.19 dB may be within seed noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Libri2Vox, a target speaker extraction (TSE) training dataset that mixes clean LibriTTS target utterances with overlapping speech from VoxCeleb2 as natural noisy interference, plus a synthetic extension generated via SynVox2 or SALT. It also proposes a three-stage curriculum learning scheme based on cosine similarity of ECAPA-TDNN speaker embeddings, with synthetic interference speakers introduced at the final stage. Experiments with Conformer, BLSTM, SpeakerBeam, and VoiceFilter compare training on Libri2Talker, Libri2Vox, and their union, with and without DNS noise augmentation, and evaluate curriculum-learning and synthetic-data configurations on the Libri2Vox test set. The main reported findings are that joint training on Libri2Talker plus Libri2Vox improves iSDR on Libri2Talker (e.g., +1.39 dB for SpeakerBeam), and that three-stage curriculum learning with synthetic speakers improves iSDR on Libri2Vox (e.g., +0.78 dB for Conformer over the no-CL real-data baseline).","tokens_in":18750,"tokens_out":7713,"duration_ms":63265,"significance":"If the gains were cleanly attributable to the proposed dataset and training strategy, Libri2Vox would be a useful community resource: it is large (149,691 training mixtures, 250 hours), offers far more interference speakers than existing TSE datasets, and includes real-world acoustic variability. The paper also runs four architectures and three seeds, and includes a useful negative control (synthetic-only training from scratch gives 7.17 dB). However, the experimental design does not currently isolate the causal contribution of the dataset's realistic noisy interference from data volume or speaker count, and the synthetic/curriculum gains are demonstrated only on the matched Libri2Vox test set. With additional controls and clearer reporting, the contribution could be significant; in its current form the central causal claims are not fully supported.","major_comments":[{"comment":"The headline gain for SpeakerBeam on Libri2Talker (7.65 dB vs. 6.26 dB, i.e., 1.39 dB) compares joint training on Libri2Talker+Libri2Vox with training on Libri2Talker alone, both with noise augmentation. These conditions differ simultaneously in training-set size (Libri2Vox adds roughly 149,691 mixtures and 250 hours), interference-speaker count (about 5,900 additional speakers), and acoustic condition (real-world VoxCeleb2 noise). There is no control that adds an equal amount of clean interference data with a matched speaker count, so the gain cannot be attributed to 'more realistic acoustic conditions' rather than to data volume or speaker diversity. The negative iSDR values for Libri2Vox-only training on Libri2Talker (e.g., -1.28 dB for Conformer) also indicate a large domain gap; the paper needs a matched control or an explicit caveat to support its cross-dataset attribution. The abstract should also state that the 1.39 dB figure is measured against the Libri2Talker-with-noise-augmentation baseline.","section":"§VII-A, Table III"},{"comment":"The 'additional 0.78 dB improvement' (16.20 vs. 15.42 dB) is reported only on the Libri2Vox test set; no curriculum-learning or synthetic-speaker model is evaluated on Libri2Talker. Because the abstract presents this result immediately after the Libri2Talker result, readers may infer cross-dataset gains that have not been measured. Moreover, this bundled figure includes both curriculum learning and synthetic augmentation; the comparison against 'w/ 3-stage CL (Real only)' at 16.01 dB shows a synthetic-only increment of 0.19 dB. The paper should state the test set explicitly in the abstract and report Libri2Talker results for the curriculum/synthetic configurations, or at least qualify the transfer claim.","section":"§VII-B, Table IV and abstract"},{"comment":"The experimental setup states 'No additional data augmentation ... was applied during training,' yet §VI-B1 describes a DNS Challenge noise-augmentation schedule applied with 50% probability. This is a direct contradiction; the text must be reconciled because all Table III rows with noise augmentation depend on this setting.","section":"§VI-A vs. §VI-B"},{"comment":"The paper reports only point averages across three independent seeded runs. No standard deviations, confidence intervals, or per-seed results are given, so the reader cannot assess whether differences such as 0.19 dB (16.01 vs. 16.20 in Table IV) or the ratio-ablation values in Fig. 5 are significant. Please provide variance information for at least the headline comparisons.","section":"§VI-A and Tables III–V"}],"minor_comments":[{"comment":"The text says 'For the training set, LibriTTS provides 1,151 speakers, with 8.97 hours of data,' but the training set is later reported as 250 hours; 8.97 hours appears to be the validation duration. Please correct this inconsistency.","section":"§III-B, Table II"},{"comment":"The abstract and body use 'SDR' while the tables report 'iSDR'; please define iSDR (improvement in SDR, presumably) and use a single consistent term throughout.","section":"Abstract and Tables III–V"},{"comment":"No dataset download URL or release plan is given. Since Libri2Vox is a central contribution, please include availability information.","section":"Dataset availability"},{"comment":"The sentence 'To demonstrate the benefits of CL in utilizing synthetic data, starting with 50% real and 50% synthetic data at stage 1 only (e.g., w/o CL (Real+SynVox2)) does not show substantial improvements compared to using synthetic and real data in Stage 3 (e.g., w/ 3-stage CL (Real + SynVox2)), while curriculum learning further enhances the performance of the latter setup' is grammatically tangled and should be split into two sentences.","section":"§VII-B"},{"comment":"The table is titled 'different numbers of synthetic speakers' but does not list the actual number of synthetic speakers used in each configuration; please add those counts.","section":"Table V"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Libri2Vox is a sensible new training resource for target speaker extraction: clean LibriTTS targets mixed with real noisy VoxCeleb2 interference, plus optional synthetic speakers from SynVox2/SALT and a three-stage curriculum. The main new results are that adding Libri2Vox to Libri2Talker training helps, and that curriculum plus synthetic speakers helps on the Libri2Vox test set. The experiments are careful in places: they include a real-data-only control for the synthetic stage, report a strong negative transfer result (trained on Libri2Vox alone does badly on Libri2Talker, -1.28 dB iSDR), and ablate the synthetic ratio and the number of synthetic speakers. Those negative results are refreshingly honest.\n\nThe soft spots are real but not fatal. The abstract's headline 1.39 dB gain for SpeakerBeam on Libri2Talker comes from joint training on Libri2Talker + Libri2Vox versus Libri2Talker alone. That comparison conflates data volume, speaker count, and real-world noise; there is no control that adds an equal number of clean mixtures with a comparable speaker set. The gain could be mostly about scale or diversity rather than realistic interference. The 0.78 dB Conformer gain similarly bundles curriculum learning with synthetic speakers (0.59 dB from CL alone, 0.19 dB from synthetic) and is only shown on the Libri2Vox test set, not Libri2Talker. To the paper's credit, the real-only control is in Table IV, so the components are partially disentangled, but the abstract overstates the synthetic contribution.\n\nAlso, the paper says three seeds were averaged but gives no error bars. Differences like 0.19 dB may be within seed noise. That is a fixable reporting issue, but it matters for a dataset paper where the quantitative claims are the point. The dataset itself is not yet released, which limits immediate use.\n\nOverall, this is a credible, useful contribution for the TSE subfield. It deserves serious review. My recommendation: send it to peer review, ask for error bars, a cleaner ablation of the data-source effect, and a direct statement of release plans. The core idea is sound.","headline":"A solid, honest dataset contribution whose headline gains are real but partly confounded by data scale and underreported without error bars.","tokens_in":19312,"tokens_out":3199,"would_cite":false,"duration_ms":27098,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that target speaker extraction improves when clean LibriTTS target speech is trained against naturally noisy VoxCeleb2 interference and synthetic speakers under a speaker-similarity curriculum, with reported gains of 1.39…","keywords":["target speaker extraction","Libri2Vox","curriculum learning","synthetic speakers","VoxCeleb2","LibriTTS","speaker similarity","signal-to-distortion ratio"],"falsifier":"Train the same TSE model on Libri2Vox alone and on Libri2Talker alone, then evaluate both on a held-out set of real overlapping-speaker recordings with ground truth; if the Libri2Vox model does not beat the clean-trained model in iSDR there, the realistic-noise transfer claim fails. A second check is to keep the curriculum fixed but replace synthetic speakers with repeated real speakers, which should remove the reported gains if the synthetic-speaker mechanism is what carries them.","tokens_in":18226,"feed_emoji":"🎙️","tokens_out":8594,"duration_ms":68082,"temperature":0.7,"pith_summary":"The paper introduces Libri2Vox, a training dataset for target speaker extraction that pairs clean target speech from LibriTTS with naturally noisy interference speech from VoxCeleb2, yielding over seven thousand speakers in conditions closer to real recordings. It claims that this combination, together with synthetic interference speakers generated by two speech-generation methods and a curriculum that orders training from easy to hard speaker pairs, makes extraction models more robust. The reported headline numbers are a 1.39 dB SDR improvement for SpeakerBeam on the Libri2Talker test set over baseline training and an additional 0.78 dB for the Conformer from similarity-based curriculum learning over random sampling. The intended contribution is a data recipe that improves TSE generalization to unseen speakers and real-world noise, rather than a new network architecture.","feed_headline":"Noisy real-world training data lift speaker extraction by 1.39 dB","feed_subtitle":"Clean speech plus real-world and synthetic interference improves speaker extraction across four models","key_machinery":"The load-bearing object is Libri2Vox, a dataset construction in which each clean LibriTTS target utterance is mixed with a randomly chosen male or female VoxCeleb2 interference utterance at an SNR drawn uniformly from -5 to 5 dB. Two synthetic variants expand the interference side: SynVox2 anonymizes VoxCeleb2 speakers through an orthogonal Householder neural network and reconstructs via HiFi-GAN, while SALT interpolates k-nearest-neighbour WavLM representations of reference speakers and also reconstructs with HiFi-GAN. The training procedure that unlocks these data is a three-stage curriculum: low-similarity real speaker pairs first, then high-similarity real pairs, then real pairs mixed with synthetic interference speakers. This ordering matters because training from scratch on synthetic data alone collapses to 7.17 dB iSDR, whereas the curriculum reaches 16.20 dB with Conformer.","core_discovery":"On its own terms, the paper's central discovery is that target speaker extraction models benefit when the interference side of their training mixtures carries real-world recording noise and a much wider speaker population than existing clean-mix datasets provide. Libri2Vox embodies this by fixing target speech as clean LibriTTS utterances and adding VoxCeleb2 interference at SNRs between -5 and 5 dB, so the noise in the interference is authentic rather than synthesized. The paper then shows that synthetic interference speakers from SynVox2 and SALT add further gains, but only when they enter late in a three-stage curriculum ordered by speaker similarity, and that joint training on Libri2Vox and Libri2Talker beats either dataset alone across Conformer, BLSTM, SpeakerBeam, and VoiceFilter. The strongest reported results are a 1.39 dB SDR gain for SpeakerBeam on Libri2Talker and a further 0.78 dB for Conformer when similarity-based curriculum learning replaces random sampling.","pith_inferences":["The paper's own Table III shows that Libri2Vox-only training gives negative iSDR on Libri2Talker, for example -1.28 dB for Conformer; an inference is that Libri2Vox is best treated as a complementary training signal for mixed clean and noisy deployment rather than a standalone replacement for clean paired data.","Because VoxCeleb2 contains multilingual, in-the-wild recordings, the paper's recipe could plausibly extend to target speaker extraction in languages and reverberant settings not represented in Libri2Talker, though the paper does not test that directly.","A natural ablation the paper does not run is to replace the generative synthetic speakers with pitch-shifted or noise-corrupted real speakers in the same curriculum; if the gains persist, the active ingredient is diversity rather than the generative model's naturalness."],"forward_implications":["Joint training on Libri2Vox and Libri2Talker improves iSDR on both test sets over single-dataset training for every architecture the paper tests.","Synthetic interference speakers add reliable gains only when introduced in the final curriculum stage; the gains surpass simply training longer on real data.","The synthetic-to-real ratio inside a mini-batch is not monotonic: 0.2 and 0.5 gave the best Conformer results at 16.20 dB, while 90% or 100% synthetic data degraded performance toward 15.82 dB and 13.61 dB.","Sequentially presenting two different synthetic speaker sets, SALT then SynVox2 or the reverse, outperforms combining both at once, reaching 13.44 dB iSDR for BLSTM.","Optional DNS-challenge noise augmentation during joint training yields small, consistent robustness gains, for example 0.41 dB for BLSTM on the Libri2Vox test set."],"supporting_citations":[{"why":"Supplies the naturally noisy, real-world interference speech that is the dataset's key departure from clean artificial mixing.","marker":"[6]"},{"why":"Supplies the clean target speech that Libri2Vox pairs with VoxCeleb2 interference.","marker":"[7]"},{"why":"Defines the Libri2Talker test set used as the main cross-dataset evaluation target.","marker":"[5]"},{"why":"Provides the Conformer-based TSE model and the initial curriculum-learning scheme this paper extends.","marker":"[22]"},{"why":"Prior result showing synthetic speakers combined with curriculum learning improves TSE, which this paper builds on with VoxCeleb2.","marker":"[24]"},{"why":"Supplies the SALT method that produces one family of synthetic interference speakers.","marker":"[25]"},{"why":"Supplies the SynVox2 privacy-friendly VoxCeleb2 generation method for the other synthetic-speaker family.","marker":"[40]"},{"why":"Provides the ECAPA-TDNN speaker encoder whose embeddings guide all four TSE models.","marker":"[30]"},{"why":"Supplies the DNS Challenge noise used for optional noise augmentation in training.","marker":"[53]"}],"fun_headline_variants":["Real-world noise boosts speaker extraction by 1.39 dB","Libri2Vox: real-world noise lifts speaker extraction 1.39 dB","Speaker extraction gains 1.39 dB from real-world noise data","Diverse speakers and synthetic data improve speaker extraction","Curriculum learning adds 0.78 dB on top of real-world data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that VoxCeleb2's naturally noisy utterances, when used as interference, are close enough to real deployment conditions to improve generalization, which is strained by the paper's own finding that Libri2Vox-only training yields negative iSDR on Libri2Talker and must be rescued by joint training.","fun_headline_variants_meta":{"raw":{"variants":["Real-world noise boosts speaker extraction by 1.39 dB","Libri2Vox: real-world noise lifts speaker extraction 1.39 dB","Speaker extraction gains 1.39 dB from real-world noise data","Diverse speakers and synthetic data improve speaker extraction","Curriculum learning adds 0.78 dB on top of real-world data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1320,"prompt_tokens":1016,"completion_tokens":304,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":213}},"tokens_in":632,"tokens_out":304,"duration_ms":2973,"temperature":1.0,"reasoning_tokens":213,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:00:12.505781+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same TSE model on Libri2Vox alone and on Libri2Talker alone, then evaluate both on a held-out set of real overlapping-speaker recordings with ground truth; if the Libri2Vox model does not beat the clean-trained model in iSDR there, the realistic-noise transfer claim fails. A second check is to keep the curriculum fixed but replace synthetic speakers with repeated real speakers, which should remove the reported gains if the synthetic-speaker mechanism is what carries them.","supporting_citations":[{"cited_title":"V oxCeleb2: Deep speaker recognition,","cited_arxiv_id":null,"evidence_quote":"Supplies the naturally noisy, real-world interference speech that is the dataset's key departure from clean artificial mixing."},{"cited_title":"LibriTTS: A corpus derived from librispeech for text-to- speech,","cited_arxiv_id":null,"evidence_quote":"Supplies the clean target speech that Libri2Vox pairs with VoxCeleb2 interference."},{"cited_title":"Target speaker verification with se- lective auditory attention for single and multi-talker speech,","cited_arxiv_id":null,"evidence_quote":"Defines the Libri2Talker test set used as the main cross-dataset evaluation target."},{"cited_title":"Target speaker extraction with curriculum learning,","cited_arxiv_id":null,"evidence_quote":"Provides the Conformer-based TSE model and the initial curriculum-learning scheme this paper extends."},{"cited_title":"Improving curriculum learning for target speaker extraction with synthetic speakers","cited_arxiv_id":"2410.00811","evidence_quote":"Prior result showing synthetic speakers combined with curriculum learning improves TSE, which this paper builds on with VoxCeleb2."},{"cited_title":"SALT: Distinguish- able speaker anonymization through latent space transformation,","cited_arxiv_id":null,"evidence_quote":"Supplies the SALT method that produces one family of synthetic interference speakers."},{"cited_title":"Synvox2: Towards a privacy-friendly V oxCeleb2 dataset,","cited_arxiv_id":null,"evidence_quote":"Supplies the SynVox2 privacy-friendly VoxCeleb2 generation method for the other synthetic-speaker family."},{"cited_title":"ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,","cited_arxiv_id":null,"evidence_quote":"Provides the ECAPA-TDNN speaker encoder whose embeddings guide all four TSE models."},{"cited_title":"ICASSP 2021 deep noise suppression challenge,","cited_arxiv_id":null,"evidence_quote":"Supplies the DNS Challenge noise used for optional noise augmentation in training."}],"review_version":1}