Fusing zero-shot TTS-generated speech embeddings with original short-utterance embeddings reduces speaker-verification EER by 10-16% on VoxCeleb1 without retraining.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
Fusing zero-shot TTS-generated speech embeddings with original short-utterance embeddings reduces speaker-verification EER by 10-16% on VoxCeleb1 without retraining.