In a fixed YourTTS framework, the H/ASP speaker encoder produces higher speaker similarity than x-vector and ECAPA-TDNN encoders.
In: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
In a fixed YourTTS framework, the H/ASP speaker encoder produces higher speaker similarity than x-vector and ECAPA-TDNN encoders.