Two open-source dual-track conversational speech corpora (Chinese and English) are introduced and shown to modestly improve a fine-tuned TTS model's naturalness metrics.
DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The majority of current Text-to-Speech (TTS) datasets, which are collections of individual utterances, contain few conversational aspects. In this paper, we introduce DailyTalk, a high-quality conversational speech dataset designed for conversational TTS. We sampled, modified, and recorded 2,541 dialogues from the open-domain dialogue dataset DailyDialog inheriting its annotated attributes. On top of our dataset, we extend prior work as our baseline, where a non-autoregressive TTS is conditioned on historical information in a dialogue. From the baseline experiment with both general and our novel metrics, we show that DailyTalk can be used as a general TTS dataset, and more than that, our baseline can represent contextual information from DailyTalk. The DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis
Two open-source dual-track conversational speech corpora (Chinese and English) are introduced and shown to modestly improve a fine-tuned TTS model's naturalness metrics.