REVIEW 3 major objections 5 minor 41 references
The first end-to-end TTS systems for Efik show that MMS-TTS is the strongest of four neural models on a new three-hour single-speaker corpus, with MOS 3.80, yet tonal errors remain.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 18:10 UTC pith:QAEBRV3B
load-bearing objection First real Efik TTS baseline: solid 3-hour single-speaker corpus and a clean four-model MOS bake-off that ranks MMS-TTS highest under genuine low-resource constraints. the 3 major comments →
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under a three-hour single-speaker regime, MMS-TTS is the strongest of the four evaluated neural TTS systems for Efik, achieving the highest MOS (3.80 ± 0.63), Nat-MOS (3.60), and A-MOS (3.04) from five native raters and generating continuous speech up to roughly three minutes without hallucination, whereas VITS, SpeechT5, and Orpheus-TTS score lower and degrade after 20–30 seconds; this constitutes the first documented end-to-end TTS baseline for the language.
What carries the argument
A curated single-speaker Efik corpus of 2,632 manually validated utterances (≈3 hours) used to fine-tune four low-resource neural TTS models (VITS, MMS-TTS initialized from Yoruba, SpeechT5, Orpheus-TTS), with ranking performed by native-speaker MOS, Nat-MOS, and A-MOS ratings.
Load-bearing premise
Subjective scores from only five native listeners on short clips, without objective tone-error or intelligibility metrics, are enough to rank the models and claim relative suitability for long-form and tonal Efik speech.
What would settle it
Collect objective tone-contour error rates or ASR-based intelligibility scores on the same test set, or re-run the MOS study with a larger rater panel and multi-speaker data; if MMS-TTS no longer ranks first or long-form coherence collapses, the central ranking claim fails.
If this is right
- A public, reproducible Efik TTS baseline now exists that later systems can be measured against.
- Multilingual pretraining (as in MMS-TTS) is currently the most practical route for intelligible synthesis of related low-resource tonal languages with only a few hours of data.
- Long-form generation remains unreliable for most architectures under three-hour single-speaker conditions, so practical applications will need additional data or architectural safeguards.
- Tonal and rare-phoneme errors persist even in the best model, confirming that larger corpora and tone-aware modeling are required before synthesis can safely preserve lexical meaning.
Where Pith is reading between the lines
- The same three-hour single-speaker recipe could be applied immediately to neighboring Lower Cross languages that already have small text resources but lack speech synthesis.
- Because tone errors still alter meaning, future evaluation protocols for these languages should include forced-choice lexical-tone discrimination tasks rather than relying solely on overall MOS.
- The observed foreign-accent residual in Orpheus-TTS and SpeechT5 suggests that cross-lingual transfer can introduce speaker-identity leakage that multi-speaker fine-tuning or speaker-embedding conditioning might later suppress.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the first documented end-to-end TTS systems for Efik, a low-resource tonal language of Southeastern Nigeria. It releases a curated single-speaker corpus of 2,632 utterances (~3.08 h) drawn from novels, folktales and educational texts, with manual orthographic transcription, native-speaker/linguist validation, 16 kHz preprocessing, and a 1975/264/393 train/val/test split. Four neural models (VITS, MMS-TTS, SpeechT5, Orpheus-TTS) are fine-tuned under identical low-resource constraints; native-speaker MOS/Nat-MOS/A-MOS (five raters) and qualitative long-form observations are reported. MMS-TTS obtains the highest scores (MOS 3.80 ± 0.63) and the most stable long-form output (~3 min coherent), while the others score lower and degrade after 20–30 s; residual tonal and rare-phoneme errors are noted for all systems. The work is framed as a reproducible baseline that underscores the need for larger multi-speaker data and tone-aware modeling.
Significance. If the corpus and relative ranking hold, the paper supplies a useful first public resource and empirical baseline for Efik speech synthesis, advancing digital preservation of an underrepresented African language. Concrete strengths include the carefully documented data-collection pipeline (consent, acoustic control, manual validation, explicit splits and preprocessing), transparent hyper-parameter listings, multi-metric subjective evaluation, an honest Limitations section, and an ethics statement that restricts release to non-commercial research use. The comparative bake-off under a shared 3-hour single-speaker regime is informative for the low-resource TTS community. Absolute naturalness remains modest and evaluation is limited, yet the existence of the resource and the clear ordering under the stated constraints constitute a genuine contribution.
major comments (3)
- [§6.1 and Table 2] The MOS/Nat-MOS/A-MOS results rest on only five native raters scoring short clips, with no inter-rater reliability statistic and no objective acoustic or tone metrics (MCD, F0 correlation, tone-error rate, ASR intelligibility). While the relative ordering is unambiguous and the limitations are disclosed, this sample size and purely subjective protocol render the absolute scores and the long-form superiority claim for MMS-TTS provisional; expanding the listener pool or adding at least one objective measure would materially strengthen the comparative conclusions that form the paper’s central empirical claim.
- [§4.1–4.2 (Data Labeling / Validation)] The manuscript never states whether the manually produced orthographic transcripts contain tone diacritics. Because Efik is tonal, residual tonal errors are repeatedly highlighted, and the abstract calls for “tone-aware modeling,” the presence or absence of tone marks in the supervision is load-bearing for interpreting model failures and for reproducibility of the baseline. This must be clarified explicitly (and, if marks are absent, the implication for tone learning should be discussed).
- [§6 (long-sequence generation paragraphs)] Long-form stability (MMS-TTS coherent to ~3 min; others collapse after 20–30 s) is asserted only as a qualitative observation. Given that this is presented as a distinguishing advantage of MMS-TTS, a more systematic protocol—e.g., MOS or intelligibility ratings on held-out long utterances, or a simple hallucination/collapse rate—would better support the claim.
minor comments (5)
- [§4.4 / Figures 1–2] Figures 1 and 2 are described (duration and word-length histograms) but their visual content is not present in the supplied text; ensure they appear with clear axis labels, bin widths and sample counts in the final version.
- [§6] The exact procedure used to extend the vocabularies/embeddings of MMS-TTS, VITS and SpeechT5 for the characters ọ and ñ is mentioned only in passing; a short appendix or footnote with the mapping and any random-initialization details would improve reproducibility.
- [References] Several reference entries carry future-dated arXiv identifiers (e.g., 2602.02734, 2603.14873) and the Orpheus-TTS citation is only a GitHub URL; verify metadata and, where possible, supply a more stable bibliographic record.
- [Throughout] Minor orthographic inconsistencies appear (“W AXAL”, “Lagunda”, spacing around “o .”, “Nat-MOS” vs. “Nat MOS”). A careful proof-reading pass will remove them.
- [§4.3] The single-speaker limitation and the 60–100 ms trailing-silence heuristic are well motivated, yet a one-sentence note on whether any automatic silence detection or energy threshold was used would help others replicate the preprocessing exactly.
Circularity Check
No circularity: pure empirical bake-off of four TTS models on a newly curated held-out Efik corpus scored by independent native listeners.
full rationale
The paper's central claims are empirical rankings (MMS-TTS MOS 3.80 ± 0.63 highest; long-form stability to ~3 min vs. 20–30 s degradation for others) obtained by fine-tuning four publicly available architectures on a 3-hour single-speaker corpus (train/val/test split in Table 1) and collecting MOS/Nat-MOS/A-MOS from five native speakers on test utterances (Table 2, §6.1). No equation, parameter, or uniqueness claim is derived from a quantity that is itself defined by the target result. Multilingual pretraining (e.g., Yoruba checkpoint for MMS-TTS) is an external prior, not a circular definition of the Efik MOS numbers. Self-citations ([8], [9]) concern related text corpora/MT work by overlapping authors and supply only background motivation; they do not underwrite the acoustic evaluation or the relative ordering. Limitations (single speaker, five raters, residual tonal errors) are disclosed rather than hidden, so the reported baseline remains self-contained against its own data and protocol.
Axiom & Free-Parameter Ledger
free parameters (5)
- VITS learning rate / epochs / batch size
- MMS-TTS learning rate / epochs / batch size
- SpeechT5 learning rate / max epochs / dropout
- Orpheus-TTS learning rate / epochs / batch size
- Trailing silence retention (60–100 ms)
axioms (4)
- domain assumption Mean Opinion Score (and Nat-MOS / A-MOS) from a small native panel is a valid proxy for naturalness, intelligibility, and accent preservation in Efik TTS.
- domain assumption Multilingual pretraining (especially a Yoruba MMS checkpoint) transfers useful phonetic/tonal knowledge to Efik under three hours of data.
- domain assumption A single native speaker recorded in quiet conditions yields acoustically consistent data sufficient for supervised end-to-end TTS ranking.
- ad hoc to paper Manual orthographic transcription without explicit tone diacritics is adequate paired supervision for the evaluated models.
Cite this review
Pith. "Pith review of Towards Digital Preservation of Efik: TTS for a Low-Resource African Language." pith.science (2026). https://pith.science/paper/QAEBRV3B
@misc{pith2026260704515,
author = {Pith},
title = {Pith review of: Towards Digital Preservation of Efik: TTS for a Low-Resource African Language},
year = {2026},
howpublished = {\url{https://pith.science/paper/QAEBRV3B}},
note = {Machine review of arXiv:2607.04515}
}
read the original abstract
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.
Figures
Reference graph
Works this paper leans on
-
[1]
Modern end-to-end ar- chitectures, often combining attention mechanisms with neural vocoders, achieve high-quality synthesis when trained on large- scale paired text–audio corpora
Introduction Text-to-Speech (TTS) technology has made remarkable progress in producing speech that approaches human natural- ness in high-resource languages [1, 2]. Modern end-to-end ar- chitectures, often combining attention mechanisms with neural vocoders, achieve high-quality synthesis when trained on large- scale paired text–audio corpora. However, th...
-
[2]
analogue,
Efik Language Documentation and Computational Efforts Historically, Efik has enjoyed a robust literary status due to early missionary documentation, such as the dictionaries of Hugh Goldie and R.G.F. Adams, and the extensive cultural records of indigenous historians like E. U. Aye. However, these works remained largely “analogue,” restricted to hard copie...
-
[3]
establishes the foundational need to transition Efik cultural and historical records from deteriorating analogue manuscripts into digital repositories, Kalejaiye [8] demonstrate the techni- cal application of such digitized data through the Ibom NLP project, creating the essential text corpora required for Anang, Efik, Ibibio, and Oro to participate in gl...
Pith/arXiv arXiv 2026
-
[4]
Several studies have explored TTS for relatively higher- resource African languages, including Yoruba, Igbo, Hausa, Swahili, Akan, Lagunda, Lingala, and Ewe [15, 1, 16, 17, 18, 19]
Related Work Recent advances in neural text-to-speech (TTS) have stimu- lated growing interest in speech technologies for African lan- guages. Several studies have explored TTS for relatively higher- resource African languages, including Yoruba, Igbo, Hausa, Swahili, Akan, Lagunda, Lingala, and Ewe [15, 1, 16, 17, 18, 19]. Large-scale multilingual efforts...
-
[5]
Recordings were cap- tured using a wireless microphone in a quiet indoor environment to minimize background noise and ensure acoustic consistency
Dataset Creation We curated approximately three hours of high-quality speech data from a single native Efik speaker. Recordings were cap- tured using a wireless microphone in a quiet indoor environment to minimize background noise and ensure acoustic consistency. The final corpus consists of 2,632 utterances segmented at the sentence level. The recording ...
-
[6]
TTS Models We fine-tuned four state-of-the-art (SOTA) Text-to-Speech (TTS) models for Efik: VITS, MMS-TTS, SpeechT5, and Orpheus-TTS. Each model was selected based on its ability to perform high-quality single-speaker TTS under low-resource conditions, which is critical for our dataset of 3 hours of single- speaker speech. VITS (Variational Inference with...
-
[7]
VITS was trained for 50 epochs with a learning rate of 2e-4, a batch size of 4, and the Adam optimizer
Experiments The four TTS models were fine-tuned with hyperparameters op- timized for low-resource, single-speaker Efik TTS. VITS was trained for 50 epochs with a learning rate of 2e-4, a batch size of 4, and the Adam optimizer. MMS-TTS was trained for 50 epochs with a learning rate of 2e-5, a batch size of 16, and the AdamW optimizer, leveraging multiling...
-
[8]
MMS-TTS achieved the highest MOS (3.80 ± 0.63), producing the most natural and intelligible speech
Results and Discussion The MOS evaluation in Table 2 highlights clear performance differences among the four TTS models for Efik speech syn- thesis. MMS-TTS achieved the highest MOS (3.80 ± 0.63), producing the most natural and intelligible speech. Its strong performance is likely due to multilingual pretraining, which en- ables the model to leverage cros...
-
[9]
We fine-tuned four state-of-the-art end-to-end TTS models, VITS, MMS-TTS, SpeechT5, and Orpheus-TTS, us- ing a single-speaker, three-hour dataset
Conclusion This work presents the first systematic effort to develop Text- to-Speech systems for Efik, a low-resource tonal African lan- guage. We fine-tuned four state-of-the-art end-to-end TTS models, VITS, MMS-TTS, SpeechT5, and Orpheus-TTS, us- ing a single-speaker, three-hour dataset. Evaluation with native speakers shows that MMS-TTS achieved the hi...
-
[10]
Limitations This study is limited by the use of a single-speaker dataset and a total of only 3 hours of audio, which constrained prosodic variation and long-sequence modeling, particularly for models like VITS. Rare phonemes such as ˜n posed challenges across all models, and while MMS-TTS handled tonal patterns reason- ably well, Orpheus-TTS and SpeechT5 ...
-
[11]
All collected materials, including narrative and educational texts, were used in accordance with copyright and cultural guidelines
Ethics Statement The single native Efik speaker provided informed consent for the recording and use of their voice for research purposes in text-to-speech development. All collected materials, including narrative and educational texts, were used in accordance with copyright and cultural guidelines. We recognize the potential risks associated with speech s...
-
[12]
We also extend our appreciation to Dr
Acknowledgements We sincerely thank the native Efik speakers for their invalu- able contribution, and the linguist who carefully validated the dataset. We also extend our appreciation to Dr. David Ade- lani, Luel Hagos, Saheed Azeez, Abraham Owodunni, Gideon george and Steven Kolawole for their guidance, insightful ad- vice, and support throughout this pr...
-
[13]
Any use of generative AI was limited to grammar correction and language editing
Generative AI Use Disclosure No generative AI tools were used to generate scientific con- tent, analyses, results, figures, or conclusions in this manuscript. Any use of generative AI was limited to grammar correction and language editing
-
[14]
P. Ogayo, G. Neubig, and A. W. Black, “Building african voices,” inProceedings of Interspeech 2022, 2022. [Online]. Available: https://arxiv.org/abs/2207.00688
Pith/arXiv arXiv 2022
-
[15]
Speech generation for indigenous language education,
A. Pine, E. Cooper, D. Guzm ´an, E. Joanis, A. Kazantseva, R. Krekoski, R. Kuhn, S. Larkin, P. Littell, D. Lothian, A. Martin, K. Richmond, M. Tessier, C. Valentini-Botinhao, D. Wells, and J. Yamagishi, “Speech generation for indigenous language education,”Computer Speech & Language, vol. 90, p. 101723, 2025. [Online]. Available: https://www.sciencedirect...
2025
-
[16]
Low-resource languages: A review of past work and future challenges,
A. Magueresse, V . Carles, and E. Heetderks, “Low-resource languages: A review of past work and future challenges,” arXiv preprint arXiv:2006.07264, 2020. [Online]. Available: https://arxiv.org/abs/2006.07264
Pith/arXiv arXiv 2006
-
[17]
Waxal: A large-scale multilingual african language speech corpus,
A. Diack, P. Nelson, K. Agbesi, A. Nakalembe, M. Mo- hamedKhair, V . Dube, T. Siyavora, S. Venugopalan, J. Hickey, U. Okonkwo, A. Bapna, I. Wiafe, R. D. Helegah, E. D. Atsakpo, C. Nutrokpor, F. B. P. Winful, K. K. Solaga, J.-D. Abdulai, A. O. Ekpezu, A. Niyonkuru, S. Rutunda, B. Ishimwe, M. Melese, E. Bainomugisha, J. Nakatumba-Nabende, A. Ka- tumba, C. B...
arXiv 2026
-
[18]
Speak, read and prompt: High-fidelity text-to-speech with mini- mal supervision,
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour, “Speak, read and prompt: High-fidelity text-to-speech with mini- mal supervision,”Transactions of the Association for Computa- tional Linguistics, vol. 11, pp. 1703–1718, 2023. [Online]. Avail- able: https://direct.mit.edu/tacl/article...
doi:10.1162/tacl 2023
-
[19]
The adaptation of english consonants by efik learners of english,
E. Mensah and E. Mensah, “The adaptation of english consonants by efik learners of english,”English Language Teaching, vol. 7, no. 3, pp. 38–45, 2013. [Online]. Available: http://dx.doi.org/10.5539/elt.v7n3p38
-
[20]
Digital language death,
A. Kornai, “Digital language death,”PLoS One, vol. 8, no. 10, p. e77056, 2013
2013
-
[21]
Ibom nlp: A step toward inclusive natural language processing for nigeria’s minority languages,
O. Kalejaiye, L. H. Beyene, D. I. Adelani, M.-m. G. Edet, A. D. Akpan, E.-A. Urua, and A. Andy, “Ibom nlp: A step toward inclusive natural language processing for nigeria’s minority languages,” inProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for C...
2025
-
[22]
Developing an english–efik corpus and machine translation system for digitization inclusion,
O. B. Edet, M. S. Awak, E. Oyo-Ita, B. O. Nyong, and I. E. Bassey, “Developing an english–efik corpus and machine translation system for digitization inclusion,”arXiv preprint arXiv:2603.14873, 2026, accepted at AfricaNLP 2026 (co-located with EACL). [Online]. Available: https: //arxiv.org/abs/2603.14873
arXiv 2026
-
[23]
Building text-to-speech models for low-resourced languages from crowdsourced data,
A. Katumba, S. Kagumire, J. Nakatumba-Nabende, J. Quinn, and S. Murindanyi, “Building text-to-speech models for low-resourced languages from crowdsourced data,”Applied AI Letters, vol. 6, no. 2, p. e117,
-
[24]
[Online]. Available: https://www.researchgate.net/ publication/391230042 Building Text-to-Speech Models for Low-Resourced Languages From Crowdsourced Data
-
[25]
Lexical tone as a distinctive cue in tonal languages,
M. Yip, “Lexical tone as a distinctive cue in tonal languages,”Cambridge lexical tone studies, 2002, describes how pitch variation conveys lexical contrast in tonal languages. [Online]. Available: https://www.cambridge. org/core/journals/bilingualism-language-and-cognition/article/ lexical-tone-as-a-cue-in-statistical-word-learning-from-bilingual-input/ 7...
2002
-
[26]
E. E. Ita,Efik Proverbs. Their Beauty, Interpretation and Applica- tion. (Book 1-4). Alpha Grafix, 2016
2016
-
[27]
From analogue to digital: Using digital technology to preserve and promote efik language, culture and history in the 21st century,
W. E. Akoda, “From analogue to digital: Using digital technology to preserve and promote efik language, culture and history in the 21st century,”Bassey Andah Journal, vol. 15, pp. 38–48, 2022. [Online]. Available: https://www. academicexcellencesociety.com/from analogue to digital using digital technology to preserve and promote efik.pdf
2022
-
[28]
A multimodal dataset for automating language vitality and endangerment assessment in south-south nigeria,
M. Ekpenyong, I. Udoh, E.-A. Urua, N. Udoh, E. Obikudo, O. Anyanwu, A. Shehu, E. Sylvanus, R. Bassey, U. Saturday, T. Fakiyesi, C.-P. Kekai, E. Udoh, S. Ansa, E. Ifesieh, G. Ikhimwin, U. Udoeyo, E. Alexander, E. Okon, M. Ekpe, B. O. Nyong, M. Darah, A. Diffre-Odiete, L. Ejobee, W. Aigbedo, F. Imoudu, C. Manda, M.-e. Kiine, D. Ugwu, and A. Akpan, “A multim...
2025
-
[29]
Building a luganda text-to-speech model from crowdsourced data,
S. Kagumire, A. Katumba, J. Nakatumba-Nabende, and J. Quinn, “Building a luganda text-to-speech model from crowdsourced data,”arXiv preprint arXiv:2405.10211, 2024, presented at the AfricaNLP Workshop at ICLR 2024. [Online]. Available: https://arxiv.org/abs/2405.10211
Pith/arXiv arXiv 2024
-
[30]
Bibletts: a large, high-fidelity, multilingual, and uniquely african speech corpus,
J. Meyer, D. I. Adelani, E. Casanova, A. ¨Oktem, D. Whitenack, J. Weber, S. Kabongo, E. Salesky, I. Orife, C. Leong, P. Ogayo, C. Emezue, J. Mukiibi, S. Osei, A. Agbolo, V . Akinode, B. Opoku, S. Olanrewaju, J. Alabi, and S. Muhammad, “Bibletts: a large, high-fidelity, multilingual, and uniquely african speech corpus,” inProceedings of Interspeech 2022, 2...
Pith/arXiv arXiv 2022
-
[31]
`Ir`oy`ınspeech: A multi-purpose yor `ub´a speech corpus,
A. A. Tol ´ul´op`e ´Og´unr`em´ı, K `ol´a T ´ub`os´un, “ `Ir`oy`ınspeech: A multi-purpose yor `ub´a speech corpus,” inProceedings of LREC 2024, 2024. [Online]. Available: https://aclanthology.org/2024. lrec-main.812/
2024
-
[32]
A cross-lingual text-to-speech system for hausa using dnn-based approach,
A. A. Aliero, D. Muhammed, M. B. Binti Peer Mustafa, M. A. Saidu, and M. Garba, “A cross-lingual text-to-speech system for hausa using dnn-based approach,”International Journal of Advanced Research in Computer Science and Software Engi- neering, vol. 10, no. 35, pp. 4460–4470, 2020, available online Jan. 1, 2020. [Online]. Available: https://www.researchg...
arXiv 2020
-
[33]
Edge-based speech transcription and synthesis for kinyarwanda and swahili languages,
P. S. Mbonimpa, D. Tuyizere, A. A. Biyabani, and O. K. Tonguz, “Edge-based speech transcription and synthesis for kinyarwanda and swahili languages,”arXiv preprint arXiv:2510.16497, 2025. [Online]. Available: https://arxiv.org/abs/2510.16497
arXiv 2025
-
[34]
Requirements and motivations of low-resource speech synthesis for language revitalization,
A. Pine, D. Wells, N. Brinklow, P. Littell, and K. Richmond, “Requirements and motivations of low-resource speech synthesis for language revitalization,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL). Dublin, Ireland: Association for Computational Linguistics, 2022, pp. 7346–7359. [Online]. Available: http...
2022
-
[35]
Robust speech recognition via large-scale weak supervision,
A. Radford and et al., “Robust speech recognition via large-scale weak supervision,” https://github.com/openai/whisper, 2023
2023
-
[36]
Xls-r: Self-supervised cross- lingual speech representation learning at scale,
A. Babu and W. Team, “Xls-r: Self-supervised cross- lingual speech representation learning at scale,” inarXiv preprint arXiv:2111.09296, 2021. [Online]. Available: https: //arxiv.org/abs/2111.09296
Pith/arXiv arXiv 2021
-
[37]
Conditional variational autoencoder with adversarial learning for end-to-end text-to- speech,
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to- speech,”arXiv preprint arXiv:2106.06103, 2021. [Online]. Available: https://arxiv.org/abs/2106.06103
Pith/arXiv arXiv 2021
-
[38]
Scaling speech technology to 1,000+ languages,
V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevski, Y . Adi, X. Zhang, W.-N. Hsu, A. Conneau, and M. Auli, “Scaling speech technology to 1,000+ languages,” arXiv preprint arXiv:2305.13516, 2023. [Online]. Available: https://arxiv.org/abs/2305.13516
Pith/arXiv arXiv 2023
-
[39]
Speecht5: Unified-modal encoder-decoder pre- training for spoken language processing,
J. Ao, R. Wang, L. Zhou, C. Wang, S. Ren, Y . Wu, S. Liu, T. Ko, Q. Li, Y . Zhang, Z. Wei, Y . Qian, J. Li, and F. Wei, “Speecht5: Unified-modal encoder-decoder pre- training for spoken language processing,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Ling...
2022
-
[40]
Orpheus-tts,
CanopyAI, “Orpheus-tts,” 2023, gitHub repository. [Online]. Available: https://github.com/canopyai/Orpheus-TTS
2023
-
[41]
Measuring speech qual- ity for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,
M. Viswanathan and M. Viswanathan, “Measuring speech qual- ity for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,”Computer Speech and Language, vol. 19, no. 1, pp. 55–83, 2005
2005
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.