Pith. sign in

VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a reference speech in a source language and subsequently transferring them to a target language using cross-lingual TTS techniques. While previous approaches have mainly concentrated on controlling voice identity within the cross-lingual TTS framework, there has been limited work on incorporating emotion and voice identity together. To this end, we introduce an end-to-end Voice Identity and Emotional Style Controllable Cross-Lingual (VECL) TTS system using multilingual speakers and an emotion embedding network. Moreover, we introduce content and style consistency losses to enhance the quality of synthesized speech further. The proposed system achieved an average relative improvement of 8.83\% compared to the state-of-the-art (SOTA) methods on a database comprising English and three Indian languages (Hindi, Telugu, and Marathi).

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

background 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Optimizing Multilingual Text-To-Speech with Accents & Emotions cs.LG · 2025-06-19 · reject · none · ref 9 · internal anchor

    A TTS system built on Parler-TTS is claimed to improve accent accuracy and emotional expressiveness for Hindi and Indian English, but the paper lacks detailed architecture and baseline evidence.