GRAFT splices a short spoken word sample into a neural codec TTS prompt and uses voice-conversion training so the model copies that pronunciation into any target voice, cutting target-word phoneme error 22-39%.
Byt5 model for massively multilingual grapheme-to-phoneme conversion,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
GRAFT splices a short spoken word sample into a neural codec TTS prompt and uses voice-conversion training so the model copies that pronunciation into any target voice, cutting target-word phoneme error 22-39%.