Pith. sign in

REVIEW 1 cited by

End to End Bangla Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.00500 v1 pith:7UZGS5WK submitted 2021-08-01 cs.SD cs.MMeess.AS

classification cs.SDcs.MMeess.AS
keywords synthesisspeechsystembanglaevaluationbeendeepmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-Speech (TTS) system is a system where speech is synthesized from a given text following any particular approach. Concatenative synthesis, Hidden Markov Model (HMM) based synthesis, Deep Learning (DL) based synthesis with multiple building blocks, etc. are the main approaches for implementing a TTS system. Here, we are presenting our deep learning-based end-to-end Bangla speech synthesis system. It has been implemented with minimal human annotation using only 3 major components (Encoder, Decoder, Post-processing net including waveform synthesis). It does not require any frontend preprocessor and Grapheme-to-Phoneme (G2P) converter. Our model has been trained with phonetically balanced 20 hours of single speaker speech data. It has obtained a 3.79 Mean Opinion Score (MOS) on a scale of 5.0 as subjective evaluation and a 0.77 Perceptual Evaluation of Speech Quality(PESQ) score on a scale of [-0.5, 4.5] as objective evaluation. It is outperforming all existing non-commercial state-of-the-art Bangla TTS systems based on naturalness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization

    cs.CL 2024-11 conditional novelty 4.0 of 10

    An ASR + MT + TTS pipeline converts Noakhali dialect speech to standard Bangla, with Whisper-large V2 achieving 0.8% CER and BanglaT5 a 41.6 BLEU on the authors' NDD dataset.

Pith tools