SD-CTC, a CTC extension with per-speaker blank tokens, improves SOT-based two-speaker ASR from 4.7% to 3.5% cpWER on LibriSpeechMix without auxiliary information.
Recognizing multi-talker speech with permutation invariant training,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
SD-CTC, a CTC extension with per-speaker blank tokens, improves SOT-based two-speaker ASR from 4.7% to 3.5% cpWER on LibriSpeechMix without auxiliary information.