Pith. sign in

REVIEW 2 cited by

DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.01063 v3 pith:FNKXKMDE submitted 2022-07-03 eess.AS cs.AIcs.CL

classification eess.AScs.AIcs.CL
keywords datasetdailytalkbaselineconversationaldialoguegeneralinformationtext-to-speech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The majority of current Text-to-Speech (TTS) datasets, which are collections of individual utterances, contain few conversational aspects. In this paper, we introduce DailyTalk, a high-quality conversational speech dataset designed for conversational TTS. We sampled, modified, and recorded 2,541 dialogues from the open-domain dialogue dataset DailyDialog inheriting its annotated attributes. On top of our dataset, we extend prior work as our baseline, where a non-autoregressive TTS is conditioned on historical information in a dialogue. From the baseline experiment with both general and our novel metrics, we show that DailyTalk can be used as a general TTS dataset, and more than that, our baseline can represent contextual information from DailyTalk. The DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Dialogs is a new 20.6-hour studio-quality Russian conversational speech corpus with emotion labels and a VITS2 proof-of-concept.

  2. Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis

    cs.SD 2025-09 conditional novelty 6.0 of 10

    Two open-source dual-track conversational speech corpora (Chinese and English) are introduced and shown to modestly improve a fine-tuned TTS model's naturalness metrics.

Pith tools