Conditioning a hybrid CTC/attention ASR on speaker embeddings and adding transfer learning from clean speech reduces word error rate on overlapped two-speaker speech to 14.6%, from a prior best of 25.4%.
Speaker embeddings inclusion strategies The first set of experiments aims to determine the best strat- egy for inclusion of speaker embeddings in the model’s in- put
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning
Conditioning a hybrid CTC/attention ASR on speaker embeddings and adding transfer learning from clean speech reduces word error rate on overlapped two-speaker speech to 14.6%, from a prior best of 25.4%.