VALL-E is a neural codec language model trained on 60K hours of speech that performs zero-shot TTS, synthesizing natural speech that matches an unseen speaker's voice, emotion, and environment from a 3-second prompt.
Generative spoken language modeling from raw audio
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2verdicts
UNVERDICTED 2representative citing papers
A framework using native-only trained discrete token surprisal and DTW alignment features improves pronunciation assessment PCC to 0.66 on SpeechOcean762, approaching supervised performance.
citing papers explorer
-
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
VALL-E is a neural codec language model trained on 60K hours of speech that performs zero-shot TTS, synthesizing natural speech that matches an unseen speaker's voice, emotion, and environment from a 3-second prompt.
-
Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal
A framework using native-only trained discrete token surprisal and DTW alignment features improves pronunciation assessment PCC to 0.66 on SpeechOcean762, approaching supervised performance.