Training-free voice conversion via factorized Monge-Kantorovich transport on WavLM embedding blocks matches FACodec-level quality with only 5-10 seconds of reference audio.
Datasets To evaluate any speaker to any speaker voice conversion we conduct our experiments on a LibriSpeech dataset [14], which consist of 40 speakers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Training-Free Voice Conversion with Factorized Optimal Transport
Training-free voice conversion via factorized Monge-Kantorovich transport on WavLM embedding blocks matches FACodec-level quality with only 5-10 seconds of reference audio.