Training-free voice conversion via factorized Monge-Kantorovich transport on WavLM embedding blocks matches FACodec-level quality with only 5-10 seconds of reference audio.
To prevent misuse, it is crucial to develop robust speech detection methods and avoid voice authentication in high-security applications
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Training-Free Voice Conversion with Factorized Optimal Transport
Training-free voice conversion via factorized Monge-Kantorovich transport on WavLM embedding blocks matches FACodec-level quality with only 5-10 seconds of reference audio.