A zero-shot voice conversion model that tunes HuBERT layer weights with adapters and uses a conditional flow-matching decoder achieves higher perceived quality and similarity than kNN-VC, DiffVC, and DDDM-VC.
V oiceMixer: Adver- sarial V oice Style Mixup,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
other 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
AdaptVC: High Quality Voice Conversion with Adaptive Learning
A zero-shot voice conversion model that tunes HuBERT layer weights with adapters and uses a conditional flow-matching decoder achieves higher perceived quality and similarity than kNN-VC, DiffVC, and DDDM-VC.