An end-to-end DOA- and beamwidth-conditioned neural network extracts target speech from six-speaker noisy mixtures, reporting SI-SDRi of 18.3 dB and WER reductions on a simulated test set.
Dataset For this study, the speech data is sourced from the LibriSpeech corpus [23], while the background noise is taken from the DE- MAND dataset [24]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
End-to-End DOA-Guided Speech Extraction in Noisy Multi-Talker Scenarios
An end-to-end DOA- and beamwidth-conditioned neural network extracts target speech from six-speaker noisy mixtures, reporting SI-SDRi of 18.3 dB and WER reductions on a simulated test set.