By injecting dynamic-phrase tokens into intermediate encoder layers of a self-conditioned CTC model, DYNAC improves biased-phrase WER on LibriSpeech test-clean from 14.1 to 3.2, reaches 2.1 overall WER, and runs at RTF 0.031 versus 0.165 for the autoregressive dynamic-vocabulary model.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
By injecting dynamic-phrase tokens into intermediate encoder layers of a self-conditioned CTC model, DYNAC improves biased-phrase WER on LibriSpeech test-clean from 14.1 to 3.2, reaches 2.1 overall WER, and runs at RTF 0.031 versus 0.165 for the autoregressive dynamic-vocabulary model.