Multi-turn textual directions can iteratively refine the speaking style of synthesized speech through a learned embedding refiner, with modest but measurable alignment to the directions.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-interaction TTS toward professional recording reproduction
Multi-turn textual directions can iteratively refine the speaking style of synthesized speech through a learned embedding refiner, with modest but measurable alignment to the directions.