Implicit Lombard-style conditioning via a style reconstruction loss gives voice-converted speech intelligibility gains comparable to explicit f0, spectral energy and spectral tilt conditioning on the Audio-Visual Lombard Grid.
Speaking style adaptation in Text-To-Speech synthesis using Sequence-to-sequence models with attention
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Currently, there are increasing interests in text-to-speech (TTS) synthesis to use sequence-to-sequence models with attention. These models are end-to-end meaning that they learn both co-articulation and duration properties directly from text and speech. Since these models are entirely data-driven, they need large amounts of data to generate synthetic speech with good quality. However, in challenging speaking styles, such as Lombard speech, it is difficult to record sufficiently large speech corpora. Therefore, in this study we propose a transfer learning method to adapt a sequence-to-sequence based TTS system of normal speaking style to Lombard style. Moreover, we experiment with a WaveNet vocoder in synthesis of Lombard speech. We conducted subjective evaluations to assess the performance of the adapted TTS systems. The subjective evaluation results indicated that an adaptation system with the WaveNet vocoder clearly outperformed the conventional deep neural network based TTS system in synthesis of Lombard speech.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning
Implicit Lombard-style conditioning via a style reconstruction loss gives voice-converted speech intelligibility gains comparable to explicit f0, spectral energy and spectral tilt conditioning on the Audio-Visual Lombard Grid.