Pith. sign in

REVIEW 40 cited by

LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.11262 v1 pith:KNCTHVE6 submitted 2020-05-22 eess.AS

classification eess.AS
keywords librimixspeechwhammodelsseparationwsj0-2mixdatasetdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In recent years, wsj0-2mix has become the reference dataset for single-channel speech separation. Most deep learning-based speech separation models today are benchmarked on it. However, recent studies have shown important performance drops when models trained on wsj0-2mix are evaluated on other, similar datasets. To address this generalization issue, we created LibriMix, an open-source alternative to wsj0-2mix, and to its noisy extension, WHAM!. Based on LibriSpeech, LibriMix consists of two- or three-speaker mixtures combined with ambient noise samples from WHAM!. Using Conv-TasNet, we achieve competitive performance on all LibriMix versions. In order to fairly evaluate across datasets, we introduce a third test set based on VCTK for speech and WHAM! for noise. Our experiments show that the generalization error is smaller for models trained with LibriMix than with WHAM!, in both clean and noisy conditions. Aiming towards evaluation in more realistic, conversation-like scenarios, we also release a sparsely overlapping version of LibriMix's test set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 40 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR

    eess.AS 2024-12 conditional novelty 7.0 of 10

    SQ-Whisper injects trainable speaker queries into Whisper's encoder and decoder, beating prior target-speaker ASR systems with new state-of-the-art WERs on Libri2Mix and WSJ0-2Mix.

  2. Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Cocktail-Talker uses three action tokens and GRPO to make a speech LLM decide whether to respond, keep listening, or ignore audio in noisy multi-speaker conversations.

  3. Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Training three omnimodel speech systems with partial, audio-derived scaffold clues that are removed at test time cuts no-clue mpWER on overlapping noisy speech from 25–71% to 9–15%.

  4. SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

    eess.AS 2026-07 conditional novelty 6.0 of 10

    The REAL-TSE challenge benchmarks target-speaker extraction on real bilingual conversational recordings with online and offline tracks, reporting that top systems beat baselines but no single system wins all quality metrics.

  5. PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Proxy-supervised joint fine-tuning of a BSRNN separator with ASR, speaker-similarity, VAD and DNSMOS losses on a new 71k real-conversation corpus yields the best SIM and timing F1 on REAL-T.

  6. Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

    cs.SD 2026-03 accept novelty 6.0 of 10

    Persistent gated residual cross-attention over onset-ordered talker acoustic memory, refined with LoRA, substantially improves LLM-SOT multi-talker ASR especially on three-talker mixtures.

  7. Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios

    eess.AS 2026-02 conditional novelty 6.0 of 10

    Using the wake-up word as enrollment degrades current target-speech-extraction models; TTS cleanup improves perceived quality but not ASR accuracy.

  8. Autoregressive Speech Enhancement via Acoustic Tokens

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Acoustic tokens outperform semantic tokens on speaker identity in speech enhancement, an autoregressive transducer helps in some settings, but discrete representations still lag continuous ones.

  9. Self-Steering Deep Non-Linear Spatially Selective Filters for Efficient Extraction of Moving Speakers under Weak Guidance

    eess.AS 2025-07 conditional novelty 6.0 of 10

    An autoregressive loop that feeds a spatially selective filter's enhanced output back into a particle filter greatly improves moving-speaker tracking and extraction under only initial-direction guidance.

  10. SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms

    eess.AS 2025-06 conditional novelty 6.0 of 10

    SpeechRefiner, a conformer-based conditional flow matching model, improves SIGMOS perceptual quality scores on speech processed by various front-ends, including unseen systems.

  11. CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models

    eess.AS 2025-05 conditional novelty 6.0 of 10

    An LLM-based ASR system that jointly performs overlapping-speech recognition and rare-word biasing, with a CTC-stage filter that trims large biasing lists, beats the tested baselines on LibriMix and AMI.

  12. AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition

    cs.SD 2025-05 conditional novelty 6.0 of 10

    AISHELL-5 releases 100+ hours of real in-car multi-channel multi-speaker Mandarin speech, 40 hours of noise, and a baseline showing mainstream ASR models still fail badly on this task.

  13. An Investigation on Speaker Augmentation for End-to-End Speaker Extraction

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Speaker augmentation via resampling and rescaling creates pseudo-speakers and hard training samples that reduce target confusion and improve end-to-end speaker extraction.

  14. Steering Deep Non-Linear Spatially Selective Filters for Weakly Guided Extraction of Moving Speakers in Dynamic Scenarios

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A weakly guided speaker extraction pipeline using only the initial direction, with a jointly trained tracker and spatial filter, outperforms a static-trained strongly guided baseline and resolves crossing-speaker ambi...

  15. Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    TFACM achieves separation quality close to TF-GridNet-Causal on WHAM!, WHAMR!, and LibriMix while using only 8.8% of its parameters and 20.4% of its compute.

  16. TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark, TS-SUPERB, evaluates self-supervised speech models on four target-speaker tasks and shows their performance is not predictable from single-speaker benchmarks.

  17. SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A four-stage pipeline (separate, correct in text, re-synthesize, align) improves speech separation quality and out-of-domain noise robustness.

  18. Beyond Speaker Identity: Text Guided Target Speech Extraction

    eess.AS 2025-01 conditional novelty 6.0 of 10

    StyleTSE extracts target speech from mixtures using natural-language speaking style descriptions, optionally combined with reference audio, trained on the new TextrolMix dataset.

  19. AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

    cs.SD 2025-01 conditional novelty 6.0 of 10

    AnCoGen is a single masked autoencoder that maps speech to and from editable attributes, enabling analysis, resynthesis, pitch shifting, and denoising.

  20. Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling

    eess.AS 2024-12 conditional novelty 6.0 of 10

    Speech enhancement quality scales with speaker and noise diversity in training data, not with text or language diversity.

  21. Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A TSE dataset using clean LibriTTS targets, noisy VoxCeleb2 interference, synthetic speaker augmentation and curriculum learning reports iSDR gains of 1.39 dB and 0.78 dB on Libri2Talker and Libri2Vox test sets.

  22. Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

    cs.MM 2026-08 conditional novelty 5.0 of 10

    An audio agent trained with trajectory-based SFT and multi-turn GRPO improves tool-use and reasoning on a new AI-generated audio agent benchmark, including tasks with unseen tools and workflows.

  23. Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

    cs.SD 2026-07 conditional novelty 5.0 of 10

    A flow-matching speech separator with biometric best-of-N candidate selection and chunk-wise channel alignment achieves competitive separation metrics and the best downstream ASR/SV error rates among evaluated systems...

  24. Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers

    eess.AS 2026-03 conditional novelty 5.0 of 10

    Feeding a spatial filter's cleaned output back into Kalman or particle filters dramatically improves moving-speaker tracking with almost no added computation.

  25. Detect, Attend and Extract: Keyword Guided Target Speaker Extraction

    eess.AS 2026-02 conditional novelty 5.0 of 10

    Keyword-guided target speaker extraction (DAE-TSE) uses a few words spoken by the target to detect, localize, and extract that speaker's full utterance from a two-speaker mixture.

  26. GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

    eess.AS 2025-12 conditional novelty 5.0 of 10

    A two-stage decoder-only language model with continuous embeddings and UTMOS-based preference fine-tuning reports improved target-speaker-extraction scores on Libri2Mix.

  27. Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Serialized CTC outputs from auxiliary branches, used as LLM prompts, improve LLM-based multi-talker ASR WER on Libri2Mix and Libri3Mix.

  28. CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CodecBench ranks 14 audio codecs on acoustic fidelity and semantic preservation across 19 datasets and four audio domains, revealing a reconstruction-versus-semantics tradeoff.

  29. SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Conditioning an SOT multi-talker ASR decoder on EEND-EDA speaker embeddings and activity information lowers WER on Libri2Mix and Libri3Mix, provided the diarization branch is accurate.

  30. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A cascaded pipeline of audio compression, latent diffusion extraction, and generative correction achieves state-of-the-art target speech extraction quality and intelligibility on Libri2Mix and out-of-domain data.

  31. SepPrune: Structured Pruning for Efficient Deep Speech Separation

    cs.SD 2025-05 conditional novelty 5.0 of 10

    SepPrune applies differentiable channel masks to compress speech separation models, reporting stronger accuracy than existing pruning baselines at matched FLOPs after fine-tuning.

  32. Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A comprehensive review of end-to-end multi-speaker ASR that contrasts SIMO and SISO architectures and reports that no design wins consistently, with real-world benchmark progress stagnant since 2021.

  33. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.

  34. Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation

    eess.AS 2024-11 conditional novelty 5.0 of 10

    A transformer ASR model with one temporally smoothed attention head per layer yields explicit speaker embeddings that improve diarization without hurting recognition.

  35. Technical Report for MERL's Real-TSE Challenge Submission

    eess.AS 2026-07 accept novelty 4.0 of 10

    Careful multi-stage data preparation and real-mixture adaptation let a baseline TSE model take first place, while DNSMOS and speaker similarity prove easily attackable without harming TER or F1.

  36. Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement

    eess.AS 2025-05 conditional novelty 4.0 of 10

    A speaker-embedding-free enhancement model is extended to do both conventional denoising and target-speaker extraction with a zero-enrollment trick, plus a consistency loss that pairs two enrollment utterances of the ...

  37. Multiple Choice Learning for Efficient Speech Separation with Many Speakers

    cs.SD 2024-11 conditional novelty 4.0 of 10

    Multiple choice learning matches permutation invariant training for speech separation on WSJ0-mix and LibriMix with up to 20 speakers, at lower loss-computation cost.

  38. GhostRNN: Reducing State Redundancy in RNN with Cheap Operations

    cs.CL 2024-11 conditional novelty 4.0 of 10

    GhostRNN compresses RNN hidden states by generating ghost states from a small set of intrinsic states with cheap linear operations, cutting parameters by about 40% with similar accuracy.

  39. Developing an Effective Training Dataset to Enhance the Performance of AI-based Speaker Separation Systems

    cs.SD 2024-11 conditional novelty 4.0 of 10

    A playback-and-record method creates a realistic two-speaker training set that yields up to 1.65 dB SI-SDR improvement over synthetic training.

  40. 30+ Years of Source Separation Research: Achievements and Future Challenges

    eess.AS 2025-01 unverdicted

    A comprehensive review of three decades of audio source separation research, presenting no new technical results.

Pith tools