REVIEW 4 major objections 5 minor 55 references
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that target speaker extraction improves when clean LibriTTS target speech is trained against naturally noisy VoxCeleb2 interference and synthetic speakers under a speaker-similarity curriculum, with reported gains of 1.39…
desk verdict A solid, honest dataset contribution whose headline gains are real but partly confounded by data scale and underreported without error bars. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is Libri2Vox, a dataset construction in which each clean LibriTTS target utterance is mixed with a randomly chosen male or female VoxCeleb2 interference utterance at an SNR drawn uniformly from -5 to 5 dB. Two synthetic variants expand the interference side: SynVox2 anonymizes VoxCeleb2 speakers through an orthogonal Householder neural network and reconstructs via HiFi-GAN, while SALT interpolates k-nearest-neighbour WavLM representations of reference speakers and also reconstructs with HiFi-GAN. The training procedure that unlocks these data is a three-stage curriculum: low-similarity real speaker pairs first, then high-similarity real pairs, then real pairs mixed with synthetic interference speakers. This ordering matters because training from scratch on synthetic data alone collapses to 7.17 dB iSDR, whereas the curriculum reaches 16.20 dB with Conformer.
What would settle it
Train the same TSE model on Libri2Vox alone and on Libri2Talker alone, then evaluate both on a held-out set of real overlapping-speaker recordings with ground truth; if the Libri2Vox model does not beat the clean-trained model in iSDR there, the realistic-noise transfer claim fails. A second check is to keep the curriculum fixed but replace synthetic speakers with repeated real speakers, which should remove the reported gains if the synthetic-speaker mechanism is what carries them.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that target speaker extraction models benefit when the interference side of their training mixtures carries real-world recording noise and a much wider speaker population than existing clean-mix datasets provide. Libri2Vox embodies this by fixing target speech as clean LibriTTS utterances and adding VoxCeleb2 interference at SNRs between -5 and 5 dB, so the noise in the interference is authentic rather than synthesized. The paper then shows that synthetic interference speakers from SynVox2 and SALT add further gains, but only when they enter late in a three-stage curriculum ordered by speaker similarity, and that joint training on Libri2Vox and Libri2Talker beats either dataset alone across Conformer, BLSTM, SpeakerBeam, and VoiceFilter. The strongest reported results are a 1.39 dB SDR gain for SpeakerBeam on Libri2Talker and a further 0.78 dB for Conformer when similarity-based curriculum learning replaces random sampling.
Load-bearing premise
The load-bearing premise is that VoxCeleb2's naturally noisy utterances, when used as interference, are close enough to real deployment conditions to improve generalization, which is strained by the paper's own finding that Libri2Vox-only training yields negative iSDR on Libri2Talker and must be rescued by joint training.
Editorial extensions
If this is right
- Joint training on Libri2Vox and Libri2Talker improves iSDR on both test sets over single-dataset training for every architecture the paper tests.
- Synthetic interference speakers add reliable gains only when introduced in the final curriculum stage; the gains surpass simply training longer on real data.
- The synthetic-to-real ratio inside a mini-batch is not monotonic: 0.2 and 0.5 gave the best Conformer results at 16.20 dB, while 90% or 100% synthetic data degraded performance toward 15.82 dB and 13.61 dB.
- Sequentially presenting two different synthetic speaker sets, SALT then SynVox2 or the reverse, outperforms combining both at once, reaching 13.44 dB iSDR for BLSTM.
- Optional DNS-challenge noise augmentation during joint training yields small, consistent robustness gains, for example 0.41 dB for BLSTM on the Libri2Vox test set.
Reading between the lines
- The paper's own Table III shows that Libri2Vox-only training gives negative iSDR on Libri2Talker, for example -1.28 dB for Conformer; an inference is that Libri2Vox is best treated as a complementary training signal for mixed clean and noisy deployment rather than a standalone replacement for clean paired data.
- Because VoxCeleb2 contains multilingual, in-the-wild recordings, the paper's recipe could plausibly extend to target speaker extraction in languages and reverberant settings not represented in Libri2Talker, though the paper does not test that directly.
- A natural ablation the paper does not run is to replace the generative synthetic speakers with pitch-shifted or noise-corrupted real speakers in the same curriculum; if the gains persist, the active ingredient is diversity rather than the generative model's naturalness.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Libri2Vox, a target speaker extraction (TSE) training dataset that mixes clean LibriTTS target utterances with overlapping speech from VoxCeleb2 as natural noisy interference, plus a synthetic extension generated via SynVox2 or SALT. It also proposes a three-stage curriculum learning scheme based on cosine similarity of ECAPA-TDNN speaker embeddings, with synthetic interference speakers introduced at the final stage. Experiments with Conformer, BLSTM, SpeakerBeam, and VoiceFilter compare training on Libri2Talker, Libri2Vox, and their union, with and without DNS noise augmentation, and evaluate curriculum-learning and synthetic-data configurations on the Libri2Vox test set. The main reported findings are that joint training on Libri2Talker plus Libri2Vox improves iSDR on Libri2Talker (e.g., +1.39 dB for SpeakerBeam), and that three-stage curriculum learning with synthetic speakers improves iSDR on Libri2Vox (e.g., +0.78 dB for Conformer over the no-CL real-data baseline).
Significance. If the gains were cleanly attributable to the proposed dataset and training strategy, Libri2Vox would be a useful community resource: it is large (149,691 training mixtures, 250 hours), offers far more interference speakers than existing TSE datasets, and includes real-world acoustic variability. The paper also runs four architectures and three seeds, and includes a useful negative control (synthetic-only training from scratch gives 7.17 dB). However, the experimental design does not currently isolate the causal contribution of the dataset's realistic noisy interference from data volume or speaker count, and the synthetic/curriculum gains are demonstrated only on the matched Libri2Vox test set. With additional controls and clearer reporting, the contribution could be significant; in its current form the central causal claims are not fully supported.
major comments (4)
- [§VII-A, Table III] The headline gain for SpeakerBeam on Libri2Talker (7.65 dB vs. 6.26 dB, i.e., 1.39 dB) compares joint training on Libri2Talker+Libri2Vox with training on Libri2Talker alone, both with noise augmentation. These conditions differ simultaneously in training-set size (Libri2Vox adds roughly 149,691 mixtures and 250 hours), interference-speaker count (about 5,900 additional speakers), and acoustic condition (real-world VoxCeleb2 noise). There is no control that adds an equal amount of clean interference data with a matched speaker count, so the gain cannot be attributed to 'more realistic acoustic conditions' rather than to data volume or speaker diversity. The negative iSDR values for Libri2Vox-only training on Libri2Talker (e.g., -1.28 dB for Conformer) also indicate a large domain gap; the paper needs a matched control or an explicit caveat to support its cross-dataset attribution. The abstract should also state that the 1.39 dB figure is measured against the Libri2Talker-with-noise-augmentation baseline.
- [§VII-B, Table IV and abstract] The 'additional 0.78 dB improvement' (16.20 vs. 15.42 dB) is reported only on the Libri2Vox test set; no curriculum-learning or synthetic-speaker model is evaluated on Libri2Talker. Because the abstract presents this result immediately after the Libri2Talker result, readers may infer cross-dataset gains that have not been measured. Moreover, this bundled figure includes both curriculum learning and synthetic augmentation; the comparison against 'w/ 3-stage CL (Real only)' at 16.01 dB shows a synthetic-only increment of 0.19 dB. The paper should state the test set explicitly in the abstract and report Libri2Talker results for the curriculum/synthetic configurations, or at least qualify the transfer claim.
- [§VI-A vs. §VI-B] The experimental setup states 'No additional data augmentation ... was applied during training,' yet §VI-B1 describes a DNS Challenge noise-augmentation schedule applied with 50% probability. This is a direct contradiction; the text must be reconciled because all Table III rows with noise augmentation depend on this setting.
- [§VI-A and Tables III–V] The paper reports only point averages across three independent seeded runs. No standard deviations, confidence intervals, or per-seed results are given, so the reader cannot assess whether differences such as 0.19 dB (16.01 vs. 16.20 in Table IV) or the ratio-ablation values in Fig. 5 are significant. Please provide variance information for at least the headline comparisons.
minor comments (5)
- [§III-B, Table II] The text says 'For the training set, LibriTTS provides 1,151 speakers, with 8.97 hours of data,' but the training set is later reported as 250 hours; 8.97 hours appears to be the validation duration. Please correct this inconsistency.
- [Abstract and Tables III–V] The abstract and body use 'SDR' while the tables report 'iSDR'; please define iSDR (improvement in SDR, presumably) and use a single consistent term throughout.
- [Dataset availability] No dataset download URL or release plan is given. Since Libri2Vox is a central contribution, please include availability information.
- [§VII-B] The sentence 'To demonstrate the benefits of CL in utilizing synthetic data, starting with 50% real and 50% synthetic data at stage 1 only (e.g., w/o CL (Real+SynVox2)) does not show substantial improvements compared to using synthetic and real data in Stage 3 (e.g., w/ 3-stage CL (Real + SynVox2)), while curriculum learning further enhances the performance of the latter setup' is grammatically tangled and should be split into two sentences.
- [Table V] The table is titled 'different numbers of synthetic speakers' but does not list the actual number of synthetic speakers used in each configuration; please add those counts.
Circularity Check
No significant circularity: the reported SDR/iSDR gains are direct empirical measurements against fixed test sets, not reductions to fitted inputs or self-citation chains.
full rationale
The paper's central claims are empirical measurements, not derivations. The headline 1.39 dB improvement for SpeakerBeam on Libri2Talker is arithmetic on Table III (7.65 vs 6.26 dB with noise augmentation), and the 0.78 dB improvement for Conformer is arithmetic on Table IV (16.20 vs 15.42 dB on the Libri2Vox test set); neither number is defined in terms of the method's own parameters. The curriculum's difficulty criterion uses speaker-embedding cosine similarity, and the same embedding family is used to condition extraction, but the evaluation metric is iSDR rather than similarity, so the reported gains are not forced by construction. The paper cites prior work by the same authors for the Conformer TSE architecture and for the three-stage curriculum learning scheme, but those citations supply heuristics and architectures rather than proof of the present results; the experiments are re-run here and judged against fixed held-out test sets. No uniqueness theorem or fitted quantity is imported to make the conclusion true. The paper also discloses the substantial domain gap (Libri2Vox-only training yields -1.28 dB iSDR on Libri2Talker), and the acknowledged limitations about synthetic data utility are empirical concerns rather than circular steps. Confounds such as the added data volume in joint training and the absence of Libri2Talker evaluation for the synthetic/CL Conformer configuration are experimental-validity issues, not circularity.
Assumptions & free parameters
free parameters (3)
- Cosine similarity threshold for Stage 1 curriculum =
0.5
- Synthetic speaker ratio in mini-batch =
0.5 (0.2 also optimal in ablation)
- Mixing SNR range =
[-5, 5] dB
assumptions (4)
- domain assumption Target speech and interference speech mix additively: m = s + s'.
- domain assumption VoxCeleb2 background noise represents real-world TSE interference conditions.
- domain assumption ECAPA-TDNN embedding cosine similarity is a valid difficulty measure for curriculum learning.
- domain assumption Synthetic speakers from SALT and SynVox2 are sufficiently natural and distinct to act as unseen interference speakers.
Cite this review
Pith. "Pith review of Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data." pith.science (2026). https://pith.science/paper/UYBSLGSU
@misc{pith2026241212512,
author = {Pith},
title = {Pith review of: Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/UYBSLGSU}},
note = {Machine review of arXiv:2412.12512}
}
read the original abstract
Target speaker extraction (TSE) is essential in speech processing applications, particularly in scenarios with complex acoustic environments. Current TSE systems face challenges in limited data diversity and a lack of robustness in real-world conditions, primarily because they are trained on artificially mixed datasets with limited speaker variability and unrealistic noise profiles. To address these challenges, we propose Libri2Vox, a new dataset that combines clean target speech from the LibriTTS dataset with interference speech from the noisy VoxCeleb2 dataset, providing a large and diverse set of speakers under realistic noisy conditions. We also augment Libri2Vox with synthetic speakers generated using state-of-the-art speech generative models to enhance speaker diversity. Additionally, to further improve the effectiveness of incorporating synthetic data, curriculum learning is implemented to progressively train TSE models with increasing levels of difficulty. Extensive experiments across multiple TSE architectures reveal varying degrees of improvement, with SpeakerBeam demonstrating the most substantial gains: a 1.39 dB improvement in signal-to-distortion ratio (SDR) on the Libri2Talker test set compared to baseline training. Building upon these results, we further enhanced performance through our speaker similarity-based curriculum learning approach with the Conformer architecture, achieving an additional 0.78 dB improvement over conventional random sampling methods in which data samples are randomly selected from the entire dataset. These results demonstrate the complementary benefits of diverse real-world data, synthetic speaker augmentation, and structured training strategies in building robust TSE systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Neural target speech extraction: An overview,
K. Zmolikova, M. Delcroix, T. Ochiai, K. Kinoshita, J. ˇCernock´y, and D. Yu, “Neural target speech extraction: An overview,” IEEE Signal Processing Magazine, vol. 40, no. 3, pp. 101–114, 2023
work page 2023
-
[2]
Speech separation with pretrained frontend to minimize domain mismatch,
W. Wang, Z. Pan, X. Li, S. Wang, and H. Li, “Speech separation with pretrained frontend to minimize domain mismatch,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 4184–4198, 2024
work page 2024
-
[3]
Spex: Multi-scale time domain speaker extraction network,
C. Xu, W. Rao, E. S. Chng, and H. Li, “Spex: Multi-scale time domain speaker extraction network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1370–1384, 2020
work page 2020
-
[4]
Spex+: A complete time domain speaker extraction network,
M. Ge, C. Xu, L. Wang, E. S. Chng, J. Dang, and H. Li, “Spex+: A complete time domain speaker extraction network,” in Proc. of INTERSPEECH, 2020, pp. 1406–1410
work page 2020
-
[5]
Target speaker verification with se- lective auditory attention for single and multi-talker speech,
C. Xu, W. Rao, J. Wu, and H. Li, “Target speaker verification with se- lective auditory attention for single and multi-talker speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2696–2709, 2021
work page 2021
-
[6]
V oxCeleb2: Deep speaker recognition,
J. S. Chung, A. Nagrani, and A. Zisserman, “V oxCeleb2: Deep speaker recognition,” in Proceedings of the Annual Conference of the Interna- tional Speech Communication Association (INTERSPEECH) , 2018, pp. 1086–1090
work page 2018
-
[7]
LibriTTS: A corpus derived from librispeech for text-to- speech,
H. Zen, V . Dang, R. Clark, Y . Zhang, R. J. Weiss, Y . Jia, Z. Chen, and Y . Wu, “LibriTTS: A corpus derived from librispeech for text-to- speech,” in Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2019, pp. 1526– 1530
work page 2019
-
[8]
LibriSpeech: an asr corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2015, pp. 5206–5210
work page 2015
Show all 55 references
-
[9]
Machine learning for synthetic data generation: A review,
Y . Lu, M. Shen, H. Wang, X. Wang, C. van Rechem, T. Fu, and W. Wei, “Machine learning for synthetic data generation: A review,” arXiv preprint arXiv:2302.04062 , 2023
2023 arXiv
-
[10]
S. I. Nikolenko, Synthetic Data for Deep Learning. Cham, Switzerland: Springer, 2021
2021
-
[11]
What makes good synthetic training data for learning dis- parity and optical flow estimation?
N. Mayer, E. Ilg, P. Fischer, C. Hazirbas, D. Cremers, A. Dosovitskiy, and T. Brox, “What makes good synthetic training data for learning dis- parity and optical flow estimation?” International Journal of Computer Vision, vol. 126, pp. 942 – 960, 2018
2018
-
[12]
Deep learning-enabled medical computer vision,
A. Esteva, K. Chou, S. Yeung, N. Naik, A. Madani, A. Mottaghi, Y . Liu, E. Topol, J. Dean, and R. Socher, “Deep learning-enabled medical computer vision,” NPJ Digital Medicine , vol. 4, 2021
2021
-
[13]
Data augmentation for low- resource neural machine translation,
M. Fadaee, A. Bisazza, and C. Monz, “Data augmentation for low- resource neural machine translation,” pp. 567–573, 2017
2017
-
[14]
A survey of data augmentation approaches for nlp,
S. Y . Feng, V . Gangal, J. Wei, S. Chandar, S. V osoughi, T. Mitamura, and E. Hovy, “A survey of data augmentation approaches for nlp,” pp. 968–988, 2021
2021
-
[15]
A survey on data synthesis and augmentation for large language models,
K. Wang, J. Zhu, M. Ren, Z. Liu, S. Li, Z. Zhang, C. Zhang, X. Wu, Q. Zhan, Q. Liu, and Y . Wang, “A survey on data synthesis and augmentation for large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12896
2024 arXiv
-
[16]
SYNT++: Utilizing imperfect synthetic data to improve speech recognition,
T.-Y . Hu, M. Armandpour, A. Shrivastava, J.-H. R. Chang, H. Koppula, and O. Tuzel, “SYNT++: Utilizing imperfect synthetic data to improve speech recognition,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 7...
2022
-
[17]
SynthASR: Unlocking synthetic data for speech recogni- tion,
A. Fazel, W. Yang, Y . Liu, R. Barra-Chicote, Y . Meng, R. Maas, and J. Droppo, “SynthASR: Unlocking synthetic data for speech recogni- tion,” in Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2021, pp. 899– 903
2021
-
[18]
Effective data augmentation methods for neural text-to-speech systems,
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y . Zhang, Y . Wang, R. Skerry-Ryanet al., “Effective data augmentation methods for neural text-to-speech systems,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Proces...
2019
-
[19]
TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis using ranking support vector machine with variational autoencoder,
M. Chen, Z. Wu, X. Wang, T. Lee, and H. Meng, “TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis using ranking support vector machine with variational autoencoder,” in ICASSP 2020- 2020 IEEE International Conference on Acoustics, Speech and Signal Processin...
2020
-
[20]
Speaker augmentation for low resource speech recognition,
B. Li, X. Zhang, M. Liu, Y . Liu, Y . Gong, and J. Zhou, “Speaker augmentation for low resource speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018, pp. 5044–5048
2018
-
[21]
Overcoming data scarcity in speaker identification: Dataset augmentation with synthetic mfccs via character-level rnn,
X. Zhang, B. Li, M. Liu, Y . Liu, Y . Gong, and J. Zhou, “Overcoming data scarcity in speaker identification: Dataset augmentation with synthetic mfccs via character-level rnn,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2...
2018
-
[22]
Target speaker extraction with curriculum learning,
Y . Liu, X. Liu, X. Miao, and J. Yamagishi, “Target speaker extraction with curriculum learning,” in INTERSPEECH2024, 2024, pp. 4348– 4352
2024
-
[23]
Curriculum learning: A survey,
P. Soviany, R. T. Ionescu, P. Rota, and N. Sebe, “Curriculum learning: A survey,” International Journal of Computer Vision , vol. 130, no. 6, pp. 1526–1565, 2022
2022
-
[24]
Improving curriculum learning for target speaker extraction with synthetic speakers,
Y . Liu, X. Liu, and J. Yamagishi, “Improving curriculum learning for target speaker extraction with synthetic speakers,” 2024 IEEE Spoken Language Technology Workshop , 12 2024. [Online]. Available: https://arxiv.org/abs/2410.00811
2024 arXiv
-
[25]
SALT: Distinguish- able speaker anonymization through latent space transformation,
Y . Lv, J. Yao, P. Chen, H. Zhou, H. Lu, and L. Xie, “SALT: Distinguish- able speaker anonymization through latent space transformation,” in 2023 IEEE Automatic Speech Recognition and Understanding (ASRU) , 2023
2023
-
[26]
SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,
K. ˇZmol´ıkov´a, M. Delcroix, K. Kinoshita, T. Ochiai, T. Nakatani, L. Burget, and J. ˇCernock´y, “SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 4, pp. 800–814, 2019
2019
-
[27]
V oiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,
Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. Hershey, R. A. Saurous, R. J. Weiss, Y . Jia, and I. L. Moreno, “V oiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,” in Proceedings of the Annual Conference of the International Speech Co...
2019
-
[28]
Deep neural networks for small footprint text-dependent speaker verification,
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez- Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in Proceedings of the IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 4052–4056
2014
-
[29]
Conformer: Convolution- JOURNAL OF LATEX CLASS FILES, VOL. XX, NO. X, XXX 2022 12 augmented transformer for speech recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution- JOURNAL OF LATEX CLASS FILES, VOL. XX, NO. X, XXX 2022 12 augmented transformer for speech recognition,” in Proceedings of the Annual Conference...
2022
-
[30]
ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” in Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2020, pp. 3830–3834
2020
-
[31]
Complex ratio masking for monaural speech separation,
D. S. Williamson, Y . Wang, and D. Wang, “Complex ratio masking for monaural speech separation,”IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 3, pp. 483–492, 2016
2016
-
[32]
CSR-I (WSJ0) Complete,
J. S. Garofolo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Complete,” 1993, lDC93S6A
1993
-
[33]
LibriMix: An open-source dataset for generalizable speech separation,
J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, “LibriMix: An open-source dataset for generalizable speech separation,” arXiv preprint arXiv:2005.11262 , 2020
2005 arXiv
-
[34]
Data augmentation for speech separation,
A. Alex, L. Wang, P. Gastaldo, and A. Cavallaro, “Data augmentation for speech separation,” Speech Communication, vol. 152, p. 102949, 2023
2023
-
[35]
Employing real training data for deep noise suppression,
Z. xu, M. Sach, J. Pirklbauer, and T. Fingscheidt, “Employing real training data for deep noise suppression,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10 731–10 735
2024
-
[36]
Perceptual evaluation of speech quality (pesq)—a new method for speech quality assessment of telephone networks and codecs,
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (pesq)—a new method for speech quality assessment of telephone networks and codecs,” Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) ...
2001
-
[37]
NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,
Z. Ju, Y . Wang, K. Shen, X. Tan, D. Xin, D. Yang, E. Liu, Y . Leng, K. Song, S. Tang, Z. Wu, T. Qin, X. Li, W. Ye, S. Zhang, J. Bian, L. He, J. Li, and S. Zhao, “NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,” in Proceedings of the 41s...
-
[38]
V ALL-E: Neural codec language models are zero-shot text to speech synthesizers,
C. Wang, S. Chen, Y . Wu, Z. Zhang, L. Z. Chen, S. Wang, Z. Liu, S. Ren, J. Liu, Z. Chen, Y . Q. Wu, J. Li, and F. Wei, “V ALL-E: Neural codec language models are zero-shot text to speech synthesizers,” arXiv preprint arXiv:2301.02111, 2023
2023 arXiv
-
[39]
FastSpeech 2: Fast and high-quality end-to-end text to speech,
Y . Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.- Y . Liu, “FastSpeech 2: Fast and high-quality end-to-end text to speech,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2021. [Online]. Available: https: //arxiv.org/abs/2006.04558
2021 arXiv
-
[40]
Synvox2: Towards a privacy-friendly V oxCeleb2 dataset,
X. Miao, X. Wang, E. Cooper, J. Yamagishi, N. Evans, M. Todisco, J.-F. Bonastre, and M. Rouvier, “Synvox2: Towards a privacy-friendly V oxCeleb2 dataset,” inICASSP 2024 - 2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 11 421–11 425
2024
-
[41]
WavLM: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-scale self-supervised pre- training for full stack speech processing,” IEEE Journal of Select...
2022
-
[42]
Nearest neighbor pattern classification,
T. Cover and P. Hart, “Nearest neighbor pattern classification,” IEEE Transactions on Information Theory , vol. 13, no. 1, pp. 21–27, 1967
1967
-
[43]
HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,” in Advances in Neural Information Processing Systems , 2020
2020
-
[44]
Speaker anonymization using orthogonal householder neural network,
X. Miao, X. Wang, E. Cooper, J. Yamagishi, and N. Tomashenko, “Speaker anonymization using orthogonal householder neural network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 3681–3695, 2023
2023
-
[45]
Yet another algorithm for pitch tracking (Y AAPT),
K. Kasi, “Yet another algorithm for pitch tracking (Y AAPT),” Master’s thesis, Old Dominion University, 2002. [Online]. Available: https://digitalcommons.odu.edu/ece etds/388/
2002
-
[46]
HuBERT: self-supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: self-supervised speech representation learning by masked prediction of hidden units,” in Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS), 2021
2021
-
[47]
Attentive statistics pooling for deep speaker embedding,
K. Okabe, T. Koshinaka, and K. Shinoda, “Attentive statistics pooling for deep speaker embedding,” in Proc. Interspeech, 2018, pp. 2252–2256
2018
-
[48]
Additive margin softmax for face verification,
F. Wang, J. Cheng, W. Liu, and H. Liu, “Additive margin softmax for face verification,” IEEE Signal Processing Letters , vol. 25, no. 7, pp. 926–930, 2018
2018
-
[49]
Open-source conversational AI with speechbrain 1.0,
M. Ravanelli, T. Parcollet, A. Moumen, S. de Langen, C. Subakan, P. Plantinga, Y . Wang, P. Mousavi, L. D. Libera, A. Ploujnikov, F. Paissan, D. Borra, S. Zaiem, Z. Zhao, S. Zhang, G. Karakasidis, S.-L. Yeh, P. Champion, A. Rouhe, R. Braun, F. Mai, J. Zuluaga-Gomez, S. M. Mous...
2024 arXiv
-
[50]
V oxCeleb: a large- scale speaker identification dataset,
A. Nagrani, J. S. Chung, and A. Zisserman, “V oxCeleb: a large- scale speaker identification dataset,” in Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH), 2017, pp. 2616–2620
2017
-
[51]
CN-Celeb: multi-genre speaker recognition,
L. Li, R. Liu, J. Kang, Y . Fan, H. Cui, Y . Cai, R. Vipperla, T. F. Zheng, and D. Wang, “CN-Celeb: multi-genre speaker recognition,” Speech Communication, vol. 137, pp. 1–13, 2022
2022
-
[52]
TorchMetrics - measuring reproducibility in pytorch,
N. S. Detlefsen, J. Borovec, J. Schock, A. Harsh, T. Koker, L. D. Liello, D. Stancl, C. Quan, M. Grechkin, and W. Falcon, “TorchMetrics - measuring reproducibility in pytorch,” feb 2022. [Online]. Available: https://github.com/Lightning-AI/torchmetrics
2022
-
[53]
ICASSP 2021 deep noise suppression challenge,
C. K. Reddy, H. Dubey, V . Gopal, R. Cutler, S. Braun, H. Gamper, R. Aichner, and S. Srinivasan, “ICASSP 2021 deep noise suppression challenge,” in ICASSP, 2021
2021
-
[54]
Towards a theoretical understanding of synthetic data in LLM post-training: A reverse-bottleneck perspective,
Z. Gan and Y . Liu, “Towards a theoretical understanding of synthetic data in LLM post-training: A reverse-bottleneck perspective,” arXiv preprint arXiv:2410.01720 , 2024. [Online]. Available: https: //arxiv.org/abs/2410.01720 APPENDIX APPENDIX A DETAILS OF NETWORK ARCHITECTUR...
2024 arXiv
-
[235]
22 605–22 623
PMLR, 21–27 Jul 2024, pp. 22 605–22 623. [Online]. Available: https://proceedings.mlr.press/v235/ju24b.html
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.