Pith. sign in

REVIEW 3 cited by

Learning pronunciation from a foreign language in speech synthesis networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.09364 v4 pith:FHMPQ3XH submitted 2018-11-23 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords languagelanguagesnetworkspeechsynthesispronunciationdifferentacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although there are more than 6,500 languages in the world, the pronunciations of many phonemes sound similar across the languages. When people learn a foreign language, their pronunciation often reflects their native language's characteristics. This motivates us to investigate how the speech synthesis network learns the pronunciation from datasets from different languages. In this study, we are interested in analyzing and taking advantage of multilingual speech synthesis network. First, we train the speech synthesis network bilingually in English and Korean and analyze how the network learns the relations of phoneme pronunciation between the languages. Our experimental result shows that the learned phoneme embedding vectors are located closer if their pronunciations are similar across the languages. Consequently, the trained networks can synthesize the English speakers' Korean speech and vice versa. Using this result, we propose a training framework to utilize information from a different language. To be specific, we pre-train a speech synthesis network using datasets from both high-resource language and low-resource language, then we fine-tune the network using the low-resource language dataset. Finally, we conducted more simulations on 10 different languages to show it is generally extendable to other languages.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Detect, Attend and Extract: Keyword Guided Target Speaker Extraction

    eess.AS 2026-02 conditional novelty 5.0 of 10

    Keyword-guided target speaker extraction (DAE-TSE) uses a few words spoken by the target to detect, localize, and extract that speaker's full utterance from a two-speaker mixture.

  2. Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding

    cs.SD 2025-05 conditional novelty 4.0 of 10

    Masked self-distillation training plus semi-autoregressive decoding improves RNN-T keyword spotting recall at low false alarm rates, especially in noisy conditions.

  3. RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A lip-to-speech model that predicts prosody from an audio prompt and content from lip-reading, then fuses them to synthesize speech, achieving strong benchmark results on LRS2 and LRS3.

Pith tools