Pith. sign in

REVIEW 3 cited by

Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09524 v1 pith:GQZCFMZW submitted 2024-10-12 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords emphasisconversationalcontextcttsrenderingspeecher-cttsexpression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. While recognizing the significance of the CTTS task, prior studies have not thoroughly investigated speech emphasis expression, which is essential for conveying the underlying intention and attitude in human-machine interaction scenarios, due to the scarcity of conversational emphasis datasets and the difficulty in context understanding. In this paper, we propose a novel Emphasis Rendering scheme for the CTTS model, termed ER-CTTS, that includes two main components: 1) we simultaneously take into account textual and acoustic contexts, with both global and local semantic modeling to understand the conversation context comprehensively; 2) we deeply integrate multi-modal and multi-scale context to learn the influence of context on the emphasis expression of the current utterance. Finally, the inferred emphasis feature is fed into the neural speech synthesizer to generate conversational speech. To address data scarcity, we create emphasis intensity annotations on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in emphasis rendering within a conversational setting. The code and audio samples are available at https://github.com/CodeStoreTTS/ER-CTTS.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

    cs.SD 2026-07 conditional novelty 6.0 of 10

    AuEmoChat's learned 1,000-code emotion token space, combined with emotion-guided token merging and classifier-guided flow matching, yields higher naturalness and emotion scores than four CSS baselines on NCSSD-EmCap.

  2. Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A conversational speech synthesizer that models word-level semantic and prosody interactions with dialogue graphs beats seven baselines on prosody ratings on DailyTalk.

  3. Improving French Synthetic Speech Quality via SSML Prosody Control

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Two fine-tuned LLMs predict SSML prosody tags that raise French TTS naturalness from a 3.20 to 3.87 MOS.

Pith tools