Pith. sign in

Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model of emotions, which represents variety of emotions as derivative products of primary emotions with varying degrees of intensity. In this paper, we propose an emotional text-to-speech design to simulate a wider spectrum of emotions grounded on the structural model. Our proposed design, Daisy-TTS, incorporates a prosody encoder to learn emotionally-separable prosody embedding as a proxy for emotion. This emotion representation allows the model to simulate: (1) Primary emotions, as learned from the training samples, (2) Secondary emotions, as a mixture of primary emotions, (3) Intensity-level, by scaling the emotion embedding, and (4) Emotions polarity, by negating the emotion embedding. Through a series of perceptual evaluations, Daisy-TTS demonstrated overall higher emotional speech naturalness and emotion perceiveability compared to the baseline.

citation-role summary

background 1

citation-polarity summary

fields

eess.AS 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Speech Synthesis along Perceptual Voice Quality Dimensions

eess.AS · 2025-01-15 · conditional · novelty 6.0

A CCNF-based manipulation block inside YourTTS shifts synthesized voices along perceptual voice quality axes, with expert-confirmed changes for breathiness and high roughness, at a cost in speaker similarity.

citing papers explorer

Showing 1 of 1 citing paper.

  • Speech Synthesis along Perceptual Voice Quality Dimensions eess.AS · 2025-01-15 · conditional · none · ref 6 · internal anchor

    A CCNF-based manipulation block inside YourTTS shifts synthesized voices along perceptual voice quality axes, with expert-confirmed changes for breathiness and high roughness, at a cost in speaker similarity.