In a fixed YourTTS framework, the H/ASP speaker encoder produces higher speaker similarity than x-vector and ECAPA-TDNN encoders.
IEEE Transactions on Audio, Speech, and Language Processing 19(4), 788–798 (2011)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
In a fixed YourTTS framework, the H/ASP speaker encoder produces higher speaker similarity than x-vector and ECAPA-TDNN encoders.