Pith. sign in

REVIEW 2 cited by

Fine-Grained Quantitative Emotion Editing for Speech Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02002 v2 pith:PC3VPDIE submitted 2024-03-04 cs.SD eess.AS

classification cs.SDeess.AS
keywords emotionhierarchicalintensityspeechgenerationeditingemotionalemotions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulating the rendering of emotions for speech generation. We propose a hierarchical emotion distribution extractor, i.e. Hierarchical ED, that quantifies the intensity of emotions at different levels of granularity. Support vector machines (SVMs) are employed to rank emotion intensity, resulting in a hierarchical emotional embedding. Hierarchical ED is subsequently integrated into the FastSpeech2 framework, guiding the model to learn emotion intensity at phoneme, word, and utterance levels. During synthesis, users can manually edit the emotional intensity of the generated voices. Both objective and subjective evaluations demonstrate the effectiveness of the proposed network in terms of fine-grained quantitative emotion editing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Control of Emotion Rendering in Speech Synthesis

    cs.SD 2024-12 conditional novelty 5.0 of 10

    A flow-matching emotional TTS framework uses hierarchical emotion distributions extracted from reference audio to enable quantitative, user-adjustable emotion intensity at phoneme, word, and utterance levels.

  2. A Review of Human Emotion Synthesis Based on Generative Technology

    cs.LG 2024-12 conditional novelty 3.0 of 10

    A systematic review that taxonomizes roughly 230 papers on generative-model-based emotion synthesis across faces, speech, and text, and catalogs datasets, metrics, and future directions.

Pith tools