Confucius4-TTS performs transcript-free, cross-lingual zero-shot voice cloning in 14 languages with competitive intelligibility and speaker similarity.
w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder
Confucius4-TTS performs transcript-free, cross-lingual zero-shot voice cloning in 14 languages with competitive intelligibility and speaker similarity.