An unsupervised framework jointly trains separation and direction-of-arrival networks by maximizing the evidence lower bound of a complex Gaussian mixture model, then uses the trained network to initialize multichannel EM separation.
Unsupervised training of neural mask-based beamforming
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present an unsupervised training approach for a neural network-based mask estimator in an acoustic beamforming application. The network is trained to maximize a likelihood criterion derived from a spatial mixture model of the observations. It is trained from scratch without requiring any parallel data consisting of degraded input and clean training targets. Thus, training can be carried out on real recordings of noisy speech rather than simulated ones. In contrast to previous work on unsupervised training of neural mask estimators, our approach avoids the need for a possibly pre-trained teacher model entirely. We demonstrate the effectiveness of our approach by speech recognition experiments on two different datasets: one mainly deteriorated by noise (CHiME 4) and one by reverberation (REVERB). The results show that the performance of the proposed system is on par with a supervised system using oracle target masks for training and with a system trained using a model-based teacher.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2019 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Deep Bayesian Unsupervised Source Separation Based on a Complex Gaussian Mixture Model
An unsupervised framework jointly trains separation and direction-of-arrival networks by maximizing the evidence lower bound of a complex Gaussian mixture model, then uses the trained network to initialize multichannel EM separation.