Pith. sign in

REVIEW 4 cited by

LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07969 v1 pith:VLOILJ26 submitted 2024-06-12 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords libritts-pstyleannotationscorpusmodelpromptpromptsspeaker
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce LibriTTS-P, a new corpus based on LibriTTS-R that includes utterance-level descriptions (i.e., prompts) of speaking style and speaker-level prompts of speaker characteristics. We employ a hybrid approach to construct prompt annotations: (1) manual annotations that capture human perceptions of speaker characteristics and (2) synthetic annotations on speaking style. Compared to existing English prompt datasets, our corpus provides more diverse prompt annotations for all speakers of LibriTTS-R. Experimental results for prompt-based controllable TTS demonstrate that the TTS model trained with LibriTTS-P achieves higher naturalness than the model using the conventional dataset. Furthermore, the results for style captioning tasks show that the model utilizing LibriTTS-P generates 2.5 times more accurate words than the model using a conventional dataset. Our corpus, LibriTTS-P, is available at https://github.com/line/LibriTTS-P.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Vo-Ve, a 44-dimension vector of voice-attribute probabilities, offers attribute-level explanations for speaker similarity, but its discrimination accuracy and listener-above-chance validation are modest.

  2. In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion

    cs.SD 2025-06 conditional novelty 6.0 of 10

    TES-VC can change both the speaker's voice and the acoustic environment of an audio clip from text prompts while preserving the words, using retrieval of known timbre embeddings and latent diffusion trained on synthet...

  3. EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations

    cs.MM 2025-05 conditional novelty 6.0 of 10

    EmotionTalk provides 19,250 utterances from 744 Chinese dyadic dialogues with emotion, sentiment, and speaking-style caption annotations.

  4. A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A three-stage training pipeline aligns face and text encoders to a pretrained speech-encoder space, then trains VITS on speech embeddings, and reports gains over single-modal baselines.

Pith tools