A transformer ASR model with one temporally smoothed attention head per layer yields explicit speaker embeddings that improve diarization without hurting recognition.
Layer-wise analysis of a self-supervised speech rep- resentation model
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
A transformer ASR model with one temporally smoothed attention head per layer yields explicit speaker embeddings that improve diarization without hurting recognition.