REVIEW 4 cited by
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce LibriTTS-P, a new corpus based on LibriTTS-R that includes utterance-level descriptions (i.e., prompts) of speaking style and speaker-level prompts of speaker characteristics. We employ a hybrid approach to construct prompt annotations: (1) manual annotations that capture human perceptions of speaker characteristics and (2) synthetic annotations on speaking style. Compared to existing English prompt datasets, our corpus provides more diverse prompt annotations for all speakers of LibriTTS-R. Experimental results for prompt-based controllable TTS demonstrate that the TTS model trained with LibriTTS-P achieves higher naturalness than the model using the conventional dataset. Furthermore, the results for style captioning tasks show that the model utilizing LibriTTS-P generates 2.5 times more accurate words than the model using a conventional dataset. Our corpus, LibriTTS-P, is available at https://github.com/line/LibriTTS-P.
Forward citations
Cited by 4 Pith papers
-
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
Vo-Ve, a 44-dimension vector of voice-attribute probabilities, offers attribute-level explanations for speaker similarity, but its discrimination accuracy and listener-above-chance validation are modest.
-
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
TES-VC can change both the speaker's voice and the acoustic environment of an audio clip from text prompts while preserving the words, using retrieval of known timbre embeddings and latent diffusion trained on synthet...
-
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
EmotionTalk provides 19,250 utterances from 744 Chinese dyadic dialogues with emotion, sentiment, and speaking-style caption annotations.
-
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
A three-stage training pipeline aligns face and text encoders to a pretrained speech-encoder space, then trains VITS on speech embeddings, and reports gains over single-modal baselines.
Discussion (0). Sign in to comment.