Pith. sign in

REVIEW 1 cited by

LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00819 v1 pith:JUPX7Y6O submitted 2024-09-01 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords speechdatasetdiarizationseparationsingle-channelfar-fieldmulti-talkerrecognition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges fall into two categories: multi-channel and single-channel solutions. Single-channel approaches, notable for their generality and convenience, do not require specific information about microphone arrays. This paper presents a large-scale far-field overlapping speech dataset, crafted to advance research in speech separation, recognition, and speaker diarization. This dataset is a critical resource for decoding ``Who said What and When'' in multi-talker, reverberant environments, a daunting challenge in the field. Additionally, we introduce a pipeline system encompassing speech separation, recognition, and diarization as a foundational benchmark. Evaluations on the WHAMR! dataset validate the broad applicability of the proposed data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A speaker-aware progressive OSD model using WavLM, Campplus, and VAD-gated masking reports 82.76% F1 on AMI, above the listed prior best of 79.21%.

Pith tools