Pith. sign in

REVIEW 2 cited by

BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.01718 v1 pith:NM7REPF2 submitted 2022-07-04 cs.CL eess.AS

classification cs.CLeess.AS
keywords modelprominencecontrastivefeaturesfocuspredictacousticbert
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Several recent studies have tested the use of transformer language model representations to infer prosodic features for text-to-speech synthesis (TTS). While these studies have explored prosody in general, in this work, we look specifically at the prediction of contrastive focus on personal pronouns. This is a particularly challenging task as it often requires semantic, discursive and/or pragmatic knowledge to predict correctly. We collect a corpus of utterances containing contrastive focus and we evaluate the accuracy of a BERT model, finetuned to predict quantized acoustic prominence features, on these samples. We also investigate how past utterances can provide relevant information for this prediction. Furthermore, we evaluate the controllability of pronoun prominence in a TTS model conditioned on acoustic prominence features.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving French Synthetic Speech Quality via SSML Prosody Control

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Two fine-tuned LLMs predict SSML prosody tags that raise French TTS naturalness from a 3.20 to 3.87 MOS.

  2. WHISTRESS: Enriching Transcriptions with Sentence Stress Detection

    cs.CL 2025-05 conditional novelty 6.0 of 10

    WHISTRESS extends Whisper with a token-level stress classifier trained on a new synthetic dataset, and shows zero-shot transfer to natural speech benchmarks.

Pith tools