A zero-shot voice conversion model that tunes HuBERT layer weights with adapters and uses a conditional flow-matching decoder achieves higher perceived quality and similarity than kNN-VC, DiffVC, and DDDM-VC.
F0- Consistent Many-To-Many Non-Parallel V oice Conversion Via Condi- tional Autoencoder,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
AdaptVC: High Quality Voice Conversion with Adaptive Learning
A zero-shot voice conversion model that tunes HuBERT layer weights with adapters and uses a conditional flow-matching decoder achieves higher perceived quality and similarity than kNN-VC, DiffVC, and DDDM-VC.