REVIEW 2 cited by
Teacher-Student MixIT for Unsupervised and Semi-supervised Speech Separation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixture invariant training (MixIT) criterion to train a teacher model. The teacher model then estimates separated sources that are used to train a student model with standard permutation invariant training (PIT). The student model can be fine-tuned with supervised data, i.e., paired artificial mixtures and clean speech sources, and further improved via model distillation. Experiments with single and multi channel mixtures show that the teacher-student training resolves the over-separation problem observed in the original MixIT method. Further, the semisupervised performance is comparable to a fully-supervised separation system trained using ten times the amount of supervised data.
Forward citations
Cited by 2 Pith papers
-
Advances in Speech Separation: Techniques, Challenges, and Future Trends
The declared speech separation survey claims a systematic four-part synthesis with fair benchmark comparisons, but its body is not present in the supplied text, so the claims could not be verified.
-
Developing an Effective Training Dataset to Enhance the Performance of AI-based Speaker Separation Systems
A playback-and-record method creates a realistic two-speaker training set that yields up to 1.65 dB SI-SDR improvement over synthetic training.
Discussion (0). Continue with ORCID to comment.