A discrete diffusion transformer with curriculum learning generates emotional speech from identity and emotion cues extracted from a face image.
SpeechT5 : Unified -modal encoder-decoder pre-training for spoken language processing
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Emotional Face-to-Speech
A discrete diffusion transformer with curriculum learning generates emotional speech from identity and emotion cues extracted from a face image.