Pith. sign in

REVIEW 3 cited by

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.10157 v1 pith:V54Z5UGT submitted 2024-09-16 eess.AS cs.SDeess.SP

classification eess.AScs.SDeess.SP
keywords emotionalemotionemotionsmodelscapabilitiescontrollabledirectemo-dpo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs' in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Differentiable Reward Optimization for LLM based TTS system

    cs.SD 2025-07 conditional novelty 6.0 of 10

    DiffRO optimizes codec-based TTS models directly on differentiable token-level rewards, improving WER and enabling zero-shot emotion control.

  2. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  3. Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

    eess.AS 2025-06 conditional novelty 4.0 of 10

    A prompt-based method lets an LLM-based TTS system synthesize speech with mixed emotions in user-specified proportions without training on mixed-emotion data.

Pith tools