A learned future-audio density ratio, FoCCE, is inserted into the streaming transducer forward recursion during training and modestly reduces word error rates.
Personalization of end-to-end speech recognition on mobile devices for named entities,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
A learned future-audio density ratio, FoCCE, is inserted into the streaming transducer forward recursion during training and modestly reduces word error rates.