TES-VC can change both the speaker's voice and the acoustic environment of an audio clip from text prompts while preserving the words, using retrieval of known timbre embeddings and latent diffusion trained on synthetic mixtures.
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and environmental acoustics. TES-VC processes simultaneous text inputs for target voice and environment, accurately generating speech matching described timbre/environment while preserving source content. Trained on synthetic data with decoupled vocal/environment features via latent diffusion modeling, our method eliminates interference between attributes. The Retrieval-Based Timbre Control (RBTC) module enables precise manipulation using abstract descriptions without paired data. Experiments confirm TES-VC effectively generates contextually appropriate speech in both timbre and environment with high content retention and superior controllability which demonstrates its potential for widespread applications.
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
TES-VC can change both the speaker's voice and the acoustic environment of an audio clip from text prompts while preserving the words, using retrieval of known timbre embeddings and latent diffusion trained on synthetic mixtures.