Contextual Speech Recognition with Difficult Negative Training Examples

· 2018 · eess.AS · arXiv 1810.12170

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

Improving the representation of contextual information is key to unlocking the potential of end-to-end (E2E) automatic speech recognition (ASR). In this work, we present a novel and simple approach for training an ASR context mechanism with difficult negative examples. The main idea is to focus on proper nouns (e.g., unique entities such as names of people and places) in the reference transcript, and use phonetically similar phrases as negative examples, encouraging the neural model to learn more discriminative representations. We apply our approach to an end-to-end contextual ASR model that jointly learns to transcribe and select the correct context items, and show that our proposed method gives up to $53.1\%$ relative improvement in word error rate (WER) across several benchmarks.

representative citing papers

Cross-Attention End-to-End ASR for Two-Party Conversations

eess.AS · 2019-07-24 · unverdicted · novelty 6.0

End-to-end ASR model with speaker-specific cross-attention for two-party conversations outperforms standard models on the Switchboard corpus.

citing papers explorer

Showing 1 of 1 citing paper.

Cross-Attention End-to-End ASR for Two-Party Conversations eess.AS · 2019-07-24 · unverdicted · none · ref 21 · internal anchor
End-to-end ASR model with speaker-specific cross-attention for two-party conversations outperforms standard models on the Switchboard corpus.

Contextual Speech Recognition with Difficult Negative Training Examples

fields

years

verdicts

representative citing papers

citing papers explorer