Pith. sign in

REVIEW 6 cited by

EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.12867 v4 pith:XNA3ITEW submitted 2025-04-17 eess.AS cs.AIcs.CL

classification eess.AScs.AIcs.CL
keywords emotionemovoicespeechemotionallanguagemodeldatadataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and chain-of-modality (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints, and demo samples are available at https://github.com/yanghaha0908/EmoVoice.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    FineCombo-TTS learns a unified acoustic representation with a CFM-based Speech Variance Predictor for flexible precise TTS control from reference audio and text descriptions, supported by the new FineEdit paired dataset.

  2. EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    EmoInstruct-TTS uses Emotion2embed and an Instruction-Conditioned Emotion Flow Model (ICE-Flow) to generate acoustically grounded emotion representations from free-form instructions and integrate them into an LLM-base...

  3. UniVocal: Unified Speech-Singing Code-Switching Synthesis

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    UniVocal presents a text-context-only framework for speech-singing code-switching synthesis via two-stage curriculum learning and a synthetic data pipeline, claiming SOTA on a new benchmark.

  4. Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning

    eess.AS 2026-04 unverdicted novelty 5.0 of 10

    A cascaded audio-prompting and ICL-based online RL method improves naturalness and expressivity in conversational TTS with reduced data needs.

  5. MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

    eess.AS 2025-08 conditional novelty 5.0 of 10

    MoE-TTS adds frozen text-expert MoE modules to a Qwen3-based TTS system and reports better out-of-domain description alignment than ElevenLabs and MiniMax on a small hand-built test set.

  6. Semantic-Aware Ship Detection with Vision-Language Integration

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Abstract claims a VLM-plus-adaptive-window framework and a new semantic ship dataset, but the manuscript body is a different voice-timbre paper.

Pith tools