Pith. sign in

REVIEW 3 major objections 5 minor 41 references

The first end-to-end TTS systems for Efik show that MMS-TTS is the strongest of four neural models on a new three-hour single-speaker corpus, with MOS 3.80, yet tonal errors remain.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-07-11 18:10 UTC pith:QAEBRV3B

load-bearing objection First real Efik TTS baseline: solid 3-hour single-speaker corpus and a clean four-model MOS bake-off that ranks MMS-TTS highest under genuine low-resource constraints. the 3 major comments →

arxiv 2607.04515 v1 pith:QAEBRV3B submitted 2026-07-05 cs.CL

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

classification cs.CL
keywords text-to-speechlow-resource languagesEfik languagetonal languagesspeech synthesisAfrican languagesMOS evaluationMMS-TTS
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Efik is a tonal language of Southeastern Nigeria with millions of speakers but no prior public speech-synthesis systems. This paper builds the first documented end-to-end text-to-speech pipeline for it by releasing a carefully recorded, manually transcribed single-speaker corpus of 2,632 utterances (about three hours) and fine-tuning four modern neural models under that severe data constraint. Native listeners rate the systems with MOS, naturalness, and accent scores; MMS-TTS leads with MOS 3.80 and can produce coherent speech lasting several minutes, while VITS, SpeechT5, and Orpheus-TTS score lower and often collapse after 20–30 seconds. The work supplies a reproducible baseline that shows current architectures can already generate intelligible Efik, yet also makes plain that larger multi-speaker data and explicit tone modeling will be required before synthesis can reliably preserve lexical meaning and cultural sound.

Core claim

Under a three-hour single-speaker regime, MMS-TTS is the strongest of the four evaluated neural TTS systems for Efik, achieving the highest MOS (3.80 ± 0.63), Nat-MOS (3.60), and A-MOS (3.04) from five native raters and generating continuous speech up to roughly three minutes without hallucination, whereas VITS, SpeechT5, and Orpheus-TTS score lower and degrade after 20–30 seconds; this constitutes the first documented end-to-end TTS baseline for the language.

What carries the argument

A curated single-speaker Efik corpus of 2,632 manually validated utterances (≈3 hours) used to fine-tune four low-resource neural TTS models (VITS, MMS-TTS initialized from Yoruba, SpeechT5, Orpheus-TTS), with ranking performed by native-speaker MOS, Nat-MOS, and A-MOS ratings.

Load-bearing premise

Subjective scores from only five native listeners on short clips, without objective tone-error or intelligibility metrics, are enough to rank the models and claim relative suitability for long-form and tonal Efik speech.

What would settle it

Collect objective tone-contour error rates or ASR-based intelligibility scores on the same test set, or re-run the MOS study with a larger rater panel and multi-speaker data; if MMS-TTS no longer ranks first or long-form coherence collapses, the central ranking claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • A public, reproducible Efik TTS baseline now exists that later systems can be measured against.
  • Multilingual pretraining (as in MMS-TTS) is currently the most practical route for intelligible synthesis of related low-resource tonal languages with only a few hours of data.
  • Long-form generation remains unreliable for most architectures under three-hour single-speaker conditions, so practical applications will need additional data or architectural safeguards.
  • Tonal and rare-phoneme errors persist even in the best model, confirming that larger corpora and tone-aware modeling are required before synthesis can safely preserve lexical meaning.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same three-hour single-speaker recipe could be applied immediately to neighboring Lower Cross languages that already have small text resources but lack speech synthesis.
  • Because tone errors still alter meaning, future evaluation protocols for these languages should include forced-choice lexical-tone discrimination tasks rather than relying solely on overall MOS.
  • The observed foreign-accent residual in Orpheus-TTS and SpeechT5 suggests that cross-lingual transfer can introduce speaker-identity leakage that multi-speaker fine-tuning or speaker-embedding conditioning might later suppress.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces the first documented end-to-end TTS systems for Efik, a low-resource tonal language of Southeastern Nigeria. It releases a curated single-speaker corpus of 2,632 utterances (~3.08 h) drawn from novels, folktales and educational texts, with manual orthographic transcription, native-speaker/linguist validation, 16 kHz preprocessing, and a 1975/264/393 train/val/test split. Four neural models (VITS, MMS-TTS, SpeechT5, Orpheus-TTS) are fine-tuned under identical low-resource constraints; native-speaker MOS/Nat-MOS/A-MOS (five raters) and qualitative long-form observations are reported. MMS-TTS obtains the highest scores (MOS 3.80 ± 0.63) and the most stable long-form output (~3 min coherent), while the others score lower and degrade after 20–30 s; residual tonal and rare-phoneme errors are noted for all systems. The work is framed as a reproducible baseline that underscores the need for larger multi-speaker data and tone-aware modeling.

Significance. If the corpus and relative ranking hold, the paper supplies a useful first public resource and empirical baseline for Efik speech synthesis, advancing digital preservation of an underrepresented African language. Concrete strengths include the carefully documented data-collection pipeline (consent, acoustic control, manual validation, explicit splits and preprocessing), transparent hyper-parameter listings, multi-metric subjective evaluation, an honest Limitations section, and an ethics statement that restricts release to non-commercial research use. The comparative bake-off under a shared 3-hour single-speaker regime is informative for the low-resource TTS community. Absolute naturalness remains modest and evaluation is limited, yet the existence of the resource and the clear ordering under the stated constraints constitute a genuine contribution.

major comments (3)
  1. [§6.1 and Table 2] The MOS/Nat-MOS/A-MOS results rest on only five native raters scoring short clips, with no inter-rater reliability statistic and no objective acoustic or tone metrics (MCD, F0 correlation, tone-error rate, ASR intelligibility). While the relative ordering is unambiguous and the limitations are disclosed, this sample size and purely subjective protocol render the absolute scores and the long-form superiority claim for MMS-TTS provisional; expanding the listener pool or adding at least one objective measure would materially strengthen the comparative conclusions that form the paper’s central empirical claim.
  2. [§4.1–4.2 (Data Labeling / Validation)] The manuscript never states whether the manually produced orthographic transcripts contain tone diacritics. Because Efik is tonal, residual tonal errors are repeatedly highlighted, and the abstract calls for “tone-aware modeling,” the presence or absence of tone marks in the supervision is load-bearing for interpreting model failures and for reproducibility of the baseline. This must be clarified explicitly (and, if marks are absent, the implication for tone learning should be discussed).
  3. [§6 (long-sequence generation paragraphs)] Long-form stability (MMS-TTS coherent to ~3 min; others collapse after 20–30 s) is asserted only as a qualitative observation. Given that this is presented as a distinguishing advantage of MMS-TTS, a more systematic protocol—e.g., MOS or intelligibility ratings on held-out long utterances, or a simple hallucination/collapse rate—would better support the claim.
minor comments (5)
  1. [§4.4 / Figures 1–2] Figures 1 and 2 are described (duration and word-length histograms) but their visual content is not present in the supplied text; ensure they appear with clear axis labels, bin widths and sample counts in the final version.
  2. [§6] The exact procedure used to extend the vocabularies/embeddings of MMS-TTS, VITS and SpeechT5 for the characters ọ and ñ is mentioned only in passing; a short appendix or footnote with the mapping and any random-initialization details would improve reproducibility.
  3. [References] Several reference entries carry future-dated arXiv identifiers (e.g., 2602.02734, 2603.14873) and the Orpheus-TTS citation is only a GitHub URL; verify metadata and, where possible, supply a more stable bibliographic record.
  4. [Throughout] Minor orthographic inconsistencies appear (“W AXAL”, “Lagunda”, spacing around “o .”, “Nat-MOS” vs. “Nat MOS”). A careful proof-reading pass will remove them.
  5. [§4.3] The single-speaker limitation and the 60–100 ms trailing-silence heuristic are well motivated, yet a one-sentence note on whether any automatic silence detection or energy threshold was used would help others replicate the preprocessing exactly.

Circularity Check

0 steps flagged

No circularity: pure empirical bake-off of four TTS models on a newly curated held-out Efik corpus scored by independent native listeners.

full rationale

The paper's central claims are empirical rankings (MMS-TTS MOS 3.80 ± 0.63 highest; long-form stability to ~3 min vs. 20–30 s degradation for others) obtained by fine-tuning four publicly available architectures on a 3-hour single-speaker corpus (train/val/test split in Table 1) and collecting MOS/Nat-MOS/A-MOS from five native speakers on test utterances (Table 2, §6.1). No equation, parameter, or uniqueness claim is derived from a quantity that is itself defined by the target result. Multilingual pretraining (e.g., Yoruba checkpoint for MMS-TTS) is an external prior, not a circular definition of the Efik MOS numbers. Self-citations ([8], [9]) concern related text corpora/MT work by overlapping authors and supply only background motivation; they do not underwrite the acoustic evaluation or the relative ordering. Limitations (single speaker, five raters, residual tonal errors) are disclosed rather than hidden, so the reported baseline remains self-contained against its own data and protocol.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

The central ranking rests on standard neural TTS practice plus a small set of domain assumptions about subjective evaluation and transfer from multilingual pretraining. No new physical or mathematical entities are postulated. Free parameters are ordinary training hyperparameters chosen for the low-resource regime; they affect absolute quality but not the existence of the baseline itself.

free parameters (5)
  • VITS learning rate / epochs / batch size
    Set to 2e-4, 50 epochs, batch 4; chosen for low-resource single-speaker training rather than derived.
  • MMS-TTS learning rate / epochs / batch size
    2e-5, 50 epochs, batch 16; hand-selected under low-resource constraints.
  • SpeechT5 learning rate / max epochs / dropout
    1e-5, up to 2500 epochs with early stopping, dropout 0.1; schedule chosen to stabilize limited data.
  • Orpheus-TTS learning rate / epochs / batch size
    2e-5, 50 epochs, batch 8; again a hand-tuned low-resource setting.
  • Trailing silence retention (60–100 ms)
    Heuristic chosen to protect sentence-final tonal cues during preprocessing.
axioms (4)
  • domain assumption Mean Opinion Score (and Nat-MOS / A-MOS) from a small native panel is a valid proxy for naturalness, intelligibility, and accent preservation in Efik TTS.
    Invoked throughout §6.1 and Table 2 as the sole quantitative ranking criterion; standard in TTS but fragile with n=5.
  • domain assumption Multilingual pretraining (especially a Yoruba MMS checkpoint) transfers useful phonetic/tonal knowledge to Efik under three hours of data.
    Used to justify MMS-TTS initialization and vocabulary extension in §6; explains the reported performance edge.
  • domain assumption A single native speaker recorded in quiet conditions yields acoustically consistent data sufficient for supervised end-to-end TTS ranking.
    Dataset design premise in §4; acknowledged as a limitation in §9.
  • ad hoc to paper Manual orthographic transcription without explicit tone diacritics is adequate paired supervision for the evaluated models.
    §4.1–4.2 describe full manual labeling after ASR failure; tone is not separately annotated, yet tonal errors are later reported as residual failures.

reviewed 2026-07-11 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Digital Preservation of Efik: TTS for a Low-Resource African Language." pith.science (2026). https://pith.science/paper/QAEBRV3B

@misc{pith2026260704515,
  author       = {Pith},
  title        = {Pith review of: Towards Digital Preservation of Efik: TTS for a Low-Resource African Language},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QAEBRV3B}},
  note         = {Machine review of arXiv:2607.04515}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.

Figures

Figures reproduced from arXiv: 2607.04515 by Archibong Okon Archibong, David Effanga Bassey, Emmanuel Oyo-Ita, Mbuotidem Sunday Awak, Offiong Bassey Edet.

Figure 1
Figure 1. Figure 1: shows the distribution of audio clip durations, where the y-axis represents the number of clips and the x-axis repre￾sents duration in seconds [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Distribution of transcript length per audio clip. 5. TTS Models We fine-tuned four state-of-the-art (SOTA) Text-to-Speech (TTS) models for Efik: VITS, MMS-TTS, SpeechT5, and Orpheus-TTS. Each model was selected based on its ability to perform high-quality single-speaker TTS under low-resource conditions, which is critical for our dataset of 3 hours of single￾speaker speech. VITS (Variational Inference with… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 1 canonical work pages

  1. [1]

    Modern end-to-end ar- chitectures, often combining attention mechanisms with neural vocoders, achieve high-quality synthesis when trained on large- scale paired text–audio corpora

    Introduction Text-to-Speech (TTS) technology has made remarkable progress in producing speech that approaches human natural- ness in high-resource languages [1, 2]. Modern end-to-end ar- chitectures, often combining attention mechanisms with neural vocoders, achieve high-quality synthesis when trained on large- scale paired text–audio corpora. However, th...

  2. [2]

    analogue,

    Efik Language Documentation and Computational Efforts Historically, Efik has enjoyed a robust literary status due to early missionary documentation, such as the dictionaries of Hugh Goldie and R.G.F. Adams, and the extensive cultural records of indigenous historians like E. U. Aye. However, these works remained largely “analogue,” restricted to hard copie...

  3. [3]

    establishes the foundational need to transition Efik cultural and historical records from deteriorating analogue manuscripts into digital repositories, Kalejaiye [8] demonstrate the techni- cal application of such digitized data through the Ibom NLP project, creating the essential text corpora required for Anang, Efik, Ibibio, and Oro to participate in gl...

  4. [4]

    Several studies have explored TTS for relatively higher- resource African languages, including Yoruba, Igbo, Hausa, Swahili, Akan, Lagunda, Lingala, and Ewe [15, 1, 16, 17, 18, 19]

    Related Work Recent advances in neural text-to-speech (TTS) have stimu- lated growing interest in speech technologies for African lan- guages. Several studies have explored TTS for relatively higher- resource African languages, including Yoruba, Igbo, Hausa, Swahili, Akan, Lagunda, Lingala, and Ewe [15, 1, 16, 17, 18, 19]. Large-scale multilingual efforts...

  5. [5]

    Recordings were cap- tured using a wireless microphone in a quiet indoor environment to minimize background noise and ensure acoustic consistency

    Dataset Creation We curated approximately three hours of high-quality speech data from a single native Efik speaker. Recordings were cap- tured using a wireless microphone in a quiet indoor environment to minimize background noise and ensure acoustic consistency. The final corpus consists of 2,632 utterances segmented at the sentence level. The recording ...

  6. [6]

    TTS Models We fine-tuned four state-of-the-art (SOTA) Text-to-Speech (TTS) models for Efik: VITS, MMS-TTS, SpeechT5, and Orpheus-TTS. Each model was selected based on its ability to perform high-quality single-speaker TTS under low-resource conditions, which is critical for our dataset of 3 hours of single- speaker speech. VITS (Variational Inference with...

  7. [7]

    VITS was trained for 50 epochs with a learning rate of 2e-4, a batch size of 4, and the Adam optimizer

    Experiments The four TTS models were fine-tuned with hyperparameters op- timized for low-resource, single-speaker Efik TTS. VITS was trained for 50 epochs with a learning rate of 2e-4, a batch size of 4, and the Adam optimizer. MMS-TTS was trained for 50 epochs with a learning rate of 2e-5, a batch size of 16, and the AdamW optimizer, leveraging multiling...

  8. [8]

    MMS-TTS achieved the highest MOS (3.80 ± 0.63), producing the most natural and intelligible speech

    Results and Discussion The MOS evaluation in Table 2 highlights clear performance differences among the four TTS models for Efik speech syn- thesis. MMS-TTS achieved the highest MOS (3.80 ± 0.63), producing the most natural and intelligible speech. Its strong performance is likely due to multilingual pretraining, which en- ables the model to leverage cros...

  9. [9]

    We fine-tuned four state-of-the-art end-to-end TTS models, VITS, MMS-TTS, SpeechT5, and Orpheus-TTS, us- ing a single-speaker, three-hour dataset

    Conclusion This work presents the first systematic effort to develop Text- to-Speech systems for Efik, a low-resource tonal African lan- guage. We fine-tuned four state-of-the-art end-to-end TTS models, VITS, MMS-TTS, SpeechT5, and Orpheus-TTS, us- ing a single-speaker, three-hour dataset. Evaluation with native speakers shows that MMS-TTS achieved the hi...

  10. [10]

    Limitations This study is limited by the use of a single-speaker dataset and a total of only 3 hours of audio, which constrained prosodic variation and long-sequence modeling, particularly for models like VITS. Rare phonemes such as ˜n posed challenges across all models, and while MMS-TTS handled tonal patterns reason- ably well, Orpheus-TTS and SpeechT5 ...

  11. [11]

    All collected materials, including narrative and educational texts, were used in accordance with copyright and cultural guidelines

    Ethics Statement The single native Efik speaker provided informed consent for the recording and use of their voice for research purposes in text-to-speech development. All collected materials, including narrative and educational texts, were used in accordance with copyright and cultural guidelines. We recognize the potential risks associated with speech s...

  12. [12]

    We also extend our appreciation to Dr

    Acknowledgements We sincerely thank the native Efik speakers for their invalu- able contribution, and the linguist who carefully validated the dataset. We also extend our appreciation to Dr. David Ade- lani, Luel Hagos, Saheed Azeez, Abraham Owodunni, Gideon george and Steven Kolawole for their guidance, insightful ad- vice, and support throughout this pr...

  13. [13]

    Any use of generative AI was limited to grammar correction and language editing

    Generative AI Use Disclosure No generative AI tools were used to generate scientific con- tent, analyses, results, figures, or conclusions in this manuscript. Any use of generative AI was limited to grammar correction and language editing

  14. [14]

    Building african voices,

    P. Ogayo, G. Neubig, and A. W. Black, “Building african voices,” inProceedings of Interspeech 2022, 2022. [Online]. Available: https://arxiv.org/abs/2207.00688

  15. [15]

    Speech generation for indigenous language education,

    A. Pine, E. Cooper, D. Guzm ´an, E. Joanis, A. Kazantseva, R. Krekoski, R. Kuhn, S. Larkin, P. Littell, D. Lothian, A. Martin, K. Richmond, M. Tessier, C. Valentini-Botinhao, D. Wells, and J. Yamagishi, “Speech generation for indigenous language education,”Computer Speech & Language, vol. 90, p. 101723, 2025. [Online]. Available: https://www.sciencedirect...

  16. [16]

    Low-resource languages: A review of past work and future challenges,

    A. Magueresse, V . Carles, and E. Heetderks, “Low-resource languages: A review of past work and future challenges,” arXiv preprint arXiv:2006.07264, 2020. [Online]. Available: https://arxiv.org/abs/2006.07264

  17. [17]

    Waxal: A large-scale multilingual african language speech corpus,

    A. Diack, P. Nelson, K. Agbesi, A. Nakalembe, M. Mo- hamedKhair, V . Dube, T. Siyavora, S. Venugopalan, J. Hickey, U. Okonkwo, A. Bapna, I. Wiafe, R. D. Helegah, E. D. Atsakpo, C. Nutrokpor, F. B. P. Winful, K. K. Solaga, J.-D. Abdulai, A. O. Ekpezu, A. Niyonkuru, S. Rutunda, B. Ishimwe, M. Melese, E. Bainomugisha, J. Nakatumba-Nabende, A. Ka- tumba, C. B...

  18. [18]

    Speak, read and prompt: High-fidelity text-to-speech with mini- mal supervision,

    E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour, “Speak, read and prompt: High-fidelity text-to-speech with mini- mal supervision,”Transactions of the Association for Computa- tional Linguistics, vol. 11, pp. 1703–1718, 2023. [Online]. Avail- able: https://direct.mit.edu/tacl/article...

  19. [19]

    The adaptation of english consonants by efik learners of english,

    E. Mensah and E. Mensah, “The adaptation of english consonants by efik learners of english,”English Language Teaching, vol. 7, no. 3, pp. 38–45, 2013. [Online]. Available: http://dx.doi.org/10.5539/elt.v7n3p38

  20. [20]

    Digital language death,

    A. Kornai, “Digital language death,”PLoS One, vol. 8, no. 10, p. e77056, 2013

  21. [21]

    Ibom nlp: A step toward inclusive natural language processing for nigeria’s minority languages,

    O. Kalejaiye, L. H. Beyene, D. I. Adelani, M.-m. G. Edet, A. D. Akpan, E.-A. Urua, and A. Andy, “Ibom nlp: A step toward inclusive natural language processing for nigeria’s minority languages,” inProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for C...

  22. [22]

    Developing an english–efik corpus and machine translation system for digitization inclusion,

    O. B. Edet, M. S. Awak, E. Oyo-Ita, B. O. Nyong, and I. E. Bassey, “Developing an english–efik corpus and machine translation system for digitization inclusion,”arXiv preprint arXiv:2603.14873, 2026, accepted at AfricaNLP 2026 (co-located with EACL). [Online]. Available: https: //arxiv.org/abs/2603.14873

  23. [23]

    Building text-to-speech models for low-resourced languages from crowdsourced data,

    A. Katumba, S. Kagumire, J. Nakatumba-Nabende, J. Quinn, and S. Murindanyi, “Building text-to-speech models for low-resourced languages from crowdsourced data,”Applied AI Letters, vol. 6, no. 2, p. e117,

  24. [24]

    Available: https://www.researchgate.net/ publication/391230042 Building Text-to-Speech Models for Low-Resourced Languages From Crowdsourced Data

    [Online]. Available: https://www.researchgate.net/ publication/391230042 Building Text-to-Speech Models for Low-Resourced Languages From Crowdsourced Data

  25. [25]

    Lexical tone as a distinctive cue in tonal languages,

    M. Yip, “Lexical tone as a distinctive cue in tonal languages,”Cambridge lexical tone studies, 2002, describes how pitch variation conveys lexical contrast in tonal languages. [Online]. Available: https://www.cambridge. org/core/journals/bilingualism-language-and-cognition/article/ lexical-tone-as-a-cue-in-statistical-word-learning-from-bilingual-input/ 7...

  26. [26]

    E. E. Ita,Efik Proverbs. Their Beauty, Interpretation and Applica- tion. (Book 1-4). Alpha Grafix, 2016

  27. [27]

    From analogue to digital: Using digital technology to preserve and promote efik language, culture and history in the 21st century,

    W. E. Akoda, “From analogue to digital: Using digital technology to preserve and promote efik language, culture and history in the 21st century,”Bassey Andah Journal, vol. 15, pp. 38–48, 2022. [Online]. Available: https://www. academicexcellencesociety.com/from analogue to digital using digital technology to preserve and promote efik.pdf

  28. [28]

    A multimodal dataset for automating language vitality and endangerment assessment in south-south nigeria,

    M. Ekpenyong, I. Udoh, E.-A. Urua, N. Udoh, E. Obikudo, O. Anyanwu, A. Shehu, E. Sylvanus, R. Bassey, U. Saturday, T. Fakiyesi, C.-P. Kekai, E. Udoh, S. Ansa, E. Ifesieh, G. Ikhimwin, U. Udoeyo, E. Alexander, E. Okon, M. Ekpe, B. O. Nyong, M. Darah, A. Diffre-Odiete, L. Ejobee, W. Aigbedo, F. Imoudu, C. Manda, M.-e. Kiine, D. Ugwu, and A. Akpan, “A multim...

  29. [29]

    Building a luganda text-to-speech model from crowdsourced data,

    S. Kagumire, A. Katumba, J. Nakatumba-Nabende, and J. Quinn, “Building a luganda text-to-speech model from crowdsourced data,”arXiv preprint arXiv:2405.10211, 2024, presented at the AfricaNLP Workshop at ICLR 2024. [Online]. Available: https://arxiv.org/abs/2405.10211

  30. [30]

    Bibletts: a large, high-fidelity, multilingual, and uniquely african speech corpus,

    J. Meyer, D. I. Adelani, E. Casanova, A. ¨Oktem, D. Whitenack, J. Weber, S. Kabongo, E. Salesky, I. Orife, C. Leong, P. Ogayo, C. Emezue, J. Mukiibi, S. Osei, A. Agbolo, V . Akinode, B. Opoku, S. Olanrewaju, J. Alabi, and S. Muhammad, “Bibletts: a large, high-fidelity, multilingual, and uniquely african speech corpus,” inProceedings of Interspeech 2022, 2...

  31. [31]

    `Ir`oy`ınspeech: A multi-purpose yor `ub´a speech corpus,

    A. A. Tol ´ul´op`e ´Og´unr`em´ı, K `ol´a T ´ub`os´un, “ `Ir`oy`ınspeech: A multi-purpose yor `ub´a speech corpus,” inProceedings of LREC 2024, 2024. [Online]. Available: https://aclanthology.org/2024. lrec-main.812/

  32. [32]

    A cross-lingual text-to-speech system for hausa using dnn-based approach,

    A. A. Aliero, D. Muhammed, M. B. Binti Peer Mustafa, M. A. Saidu, and M. Garba, “A cross-lingual text-to-speech system for hausa using dnn-based approach,”International Journal of Advanced Research in Computer Science and Software Engi- neering, vol. 10, no. 35, pp. 4460–4470, 2020, available online Jan. 1, 2020. [Online]. Available: https://www.researchg...

  33. [33]

    Edge-based speech transcription and synthesis for kinyarwanda and swahili languages,

    P. S. Mbonimpa, D. Tuyizere, A. A. Biyabani, and O. K. Tonguz, “Edge-based speech transcription and synthesis for kinyarwanda and swahili languages,”arXiv preprint arXiv:2510.16497, 2025. [Online]. Available: https://arxiv.org/abs/2510.16497

  34. [34]

    Requirements and motivations of low-resource speech synthesis for language revitalization,

    A. Pine, D. Wells, N. Brinklow, P. Littell, and K. Richmond, “Requirements and motivations of low-resource speech synthesis for language revitalization,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL). Dublin, Ireland: Association for Computational Linguistics, 2022, pp. 7346–7359. [Online]. Available: http...

  35. [35]

    Robust speech recognition via large-scale weak supervision,

    A. Radford and et al., “Robust speech recognition via large-scale weak supervision,” https://github.com/openai/whisper, 2023

  36. [36]

    Xls-r: Self-supervised cross- lingual speech representation learning at scale,

    A. Babu and W. Team, “Xls-r: Self-supervised cross- lingual speech representation learning at scale,” inarXiv preprint arXiv:2111.09296, 2021. [Online]. Available: https: //arxiv.org/abs/2111.09296

  37. [37]

    Conditional variational autoencoder with adversarial learning for end-to-end text-to- speech,

    J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to- speech,”arXiv preprint arXiv:2106.06103, 2021. [Online]. Available: https://arxiv.org/abs/2106.06103

  38. [38]

    Scaling speech technology to 1,000+ languages,

    V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevski, Y . Adi, X. Zhang, W.-N. Hsu, A. Conneau, and M. Auli, “Scaling speech technology to 1,000+ languages,” arXiv preprint arXiv:2305.13516, 2023. [Online]. Available: https://arxiv.org/abs/2305.13516

  39. [39]

    Speecht5: Unified-modal encoder-decoder pre- training for spoken language processing,

    J. Ao, R. Wang, L. Zhou, C. Wang, S. Ren, Y . Wu, S. Liu, T. Ko, Q. Li, Y . Zhang, Z. Wei, Y . Qian, J. Li, and F. Wei, “Speecht5: Unified-modal encoder-decoder pre- training for spoken language processing,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Ling...

  40. [40]

    Orpheus-tts,

    CanopyAI, “Orpheus-tts,” 2023, gitHub repository. [Online]. Available: https://github.com/canopyai/Orpheus-TTS

  41. [41]

    Measuring speech qual- ity for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,

    M. Viswanathan and M. Viswanathan, “Measuring speech qual- ity for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,”Computer Speech and Language, vol. 19, no. 1, pp. 55–83, 2005

This paper was first reviewed by grok-4.5 on July 11, 2026.