With matched data and sizes, CTC+LLM shallow fusion beats tight speech-LLM integration on in-domain ASR, while prefix LLMs win average WER on out-of-domain HuggingFace sets.
Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
dataset 1polarities
use dataset 1representative citing papers
FastTurn fuses streaming CTC semantics with acoustic cues for lower-latency, more accurate turn detection in full-duplex dialogue and releases a real-dialogue test set.
Iterative pseudo-labeling on 22.4k hours of unlabeled code-switching audio reduces Mix Error Rate on SEAME devman to 12.88% and devsge to 18.89%.
citing papers explorer
-
LLMs and Speech: Integration vs. Combination
With matched data and sizes, CTC+LLM shallow fusion beats tight speech-LLM integration on in-domain ASR, while prefix LLMs win average WER on out-of-domain HuggingFace sets.
-
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
FastTurn fuses streaming CTC semantics with acoustic cues for lower-latency, more accurate turn detection in full-duplex dialogue and releases a real-dialogue test set.
-
Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR
Iterative pseudo-labeling on 22.4k hours of unlabeled code-switching audio reduces Mix Error Rate on SEAME devman to 12.88% and devsge to 18.89%.