Conditioning a hybrid CTC/attention ASR on speaker embeddings and adding transfer learning from clean speech reduces word error rate on overlapped two-speaker speech to 14.6%, from a prior best of 25.4%.
Overlapped speech – well known in a more general context as the cocktail party problem – remains, however, to be a largely unsolved problem
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning
Conditioning a hybrid CTC/attention ASR on speaker embeddings and adding transfer learning from clean speech reduces word error rate on overlapped two-speaker speech to 14.6%, from a prior best of 25.4%.