REVIEW 3 cited by
SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a multi-channel database of overlapping speech for training, evaluation, and detailed analysis of source separation and extraction algorithms: SMS-WSJ -- Spatialized Multi-Speaker Wall Street Journal. It consists of artificially mixed speech taken from the WSJ database, but unlike earlier databases we consider all WSJ0+1 utterances and take care of strictly separating the speaker sets present in the training, validation and test sets. When spatializing the data we ensure a high degree of randomness w.r.t. room size, array center and rotation, as well as speaker position. Furthermore, this paper offers a critical assessment of recently proposed measures of source separation performance. Alongside the code to generate the database we provide a source separation baseline and a Kaldi recipe with competitive word error rates to provide common ground for evaluation.
Forward citations
Cited by 3 Pith papers
-
On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
For joint denoising and dereverberation with diffusion models, cascade quality depends strongly on which distortion is removed first, and a single mixed-objective model gives the best overall compromise.
-
Error Analysis in a Modular Meeting Transcription System
In a modular meeting transcription pipeline, missing speech segments, not primary-to-cross-channel leakage, cause most of the gap to oracle segmentation.
-
Advances in Speech Separation: Techniques, Challenges, and Future Trends
The declared speech separation survey claims a systematic four-part synthesis with fair benchmark comparisons, but its body is not present in the supplied text, so the claims could not be verified.
Discussion (0). Continue with ORCID to comment.