A transformer ASR model with one temporally smoothed attention head per layer yields explicit speaker embeddings that improve diarization without hurting recognition.
Large-scale self-supervised speech representation learning for automatic speaker verification
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
A transformer ASR model with one temporally smoothed attention head per layer yields explicit speaker embeddings that improve diarization without hurting recognition.