Pith. sign in

REVIEW 1 cited by

Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.06000 v1 pith:MPZNHTQU submitted 2022-07-13 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords stylespeechcontrolcross-speakeremotionalproposetargettransfer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the target style. In many practical situations, users may not have reference speech recorded in target emotion but still be interested in controlling speech style just by typing text description of desired emotional style. In this paper, we propose a text-based interface for emotional style control and cross-speaker style transfer in multi-speaker TTS. We propose the bi-modal style encoder which models the semantic relationship between text description embedding and speech style embedding with a pretrained language model. To further improve cross-speaker style transfer on disjoint, multi-style datasets, we propose the novel style loss. The experimental results show that our model can generate high-quality expressive speech even in unseen style.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

    cs.SD 2025-07 conditional novelty 7.0 of 10

    TTS-CtrlNet adds time-varying emotion control to a frozen flow-matching TTS model using a ControlNet-style trainable copy, improving emotion similarity metrics while preserving the base model's voice cloning.

Pith tools