ClapFM-EVC uses a contrastive emotion encoder and conditional flow matching to convert speech emotion from natural language prompts or reference audio, and reports state-of-the-art quality on a single-speaker Mandarin test set.
System Overview As illustrated in Fig
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
ClapFM-EVC uses a contrastive emotion encoder and conditional flow matching to convert speech emotion from natural language prompts or reference audio, and reports state-of-the-art quality on a single-speaker Mandarin test set.