I2TTS is an end-to-end TTS that conditions on CLIP image features and a frozen reverberation classifier to synthesize scene-matched, speaker-adaptive speech, reporting gains on SRE, MOS, and WER.
Instructtts: Mod- elling expressive tts in discrete latent space with natural language style prompt,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
I2TTS is an end-to-end TTS that conditions on CLIP image features and a frozen reverberation classifier to synthesize scene-matched, speaker-adaptive speech, reporting gains on SRE, MOS, and WER.