Pith. sign in

REVIEW 2 cited by

Unsupervised training of neural mask-based beamforming

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.01578 v2 pith:JVT2QVI6 submitted 2019-04-02 cs.SD cs.LGstat.ML

classification cs.SDcs.LGstat.ML
keywords trainingapproachneuralsystemtrainedunsupervisedbeamformingmask
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present an unsupervised training approach for a neural network-based mask estimator in an acoustic beamforming application. The network is trained to maximize a likelihood criterion derived from a spatial mixture model of the observations. It is trained from scratch without requiring any parallel data consisting of degraded input and clean training targets. Thus, training can be carried out on real recordings of noisy speech rather than simulated ones. In contrast to previous work on unsupervised training of neural mask estimators, our approach avoids the need for a possibly pre-trained teacher model entirely. We demonstrate the effectiveness of our approach by speech recognition experiments on two different datasets: one mainly deteriorated by noise (CHiME 4) and one by reverberation (REVERB). The results show that the performance of the proposed system is on par with a supervised system using oracle target masks for training and with a system trained using a model-based teacher.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior

    eess.AS 2025-05 conditional novelty 7.0 of 10

    Blind multi-channel speech separation can be solved with a single-speaker diffusion prior plus an estimated likelihood, without knowing array geometry or using paired training data.

  2. Deep Bayesian Unsupervised Source Separation Based on a Complex Gaussian Mixture Model

    cs.SD 2019-08 conditional novelty 6.0 of 10

    An unsupervised framework jointly trains separation and direction-of-arrival networks by maximizing the evidence lower bound of a complex Gaussian mixture model, then uses the trained network to initialize multichanne...

Pith tools