Pith. sign in

REVIEW 1 cited by

Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.08810 v1 pith:AWVLDGCM submitted 2023-01-20 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords bertmodelsphoneme-levelnaturalnessphonemespredictionstasktext-to-speech
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic prosodic patterns. However, these models are usually word-level or sup-phoneme-level and jointly trained with phonemes, making them inefficient for the downstream TTS task where only phonemes are needed. In this work, we propose a phoneme-level BERT (PL-BERT) with a pretext task of predicting the corresponding graphemes along with the regular masked phoneme predictions. Subjective evaluations show that our phoneme-level BERT encoder has significantly improved the mean opinion scores (MOS) of rated naturalness of synthesized speech compared with the state-of-the-art (SOTA) StyleTTS baseline on out-of-distribution (OOD) texts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prosody Labeling with Phoneme-BERT and Speech Foundation Models

    eess.AS 2025-07 conditional novelty 4.0 of 10

    Combining frozen speech SSL/Whisper features with phoneme-level BERT features improves automatic Japanese prosody label prediction over either input alone.

Pith tools