{"id":"2c8a2259-9db9-4586-8376-8f24cb583ff2","arxiv_id":"2607.04515","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"First end-to-end Efik TTS baseline: a 3-hour single-speaker corpus and four fine-tuned models, with MMS-TTS best at MOS 3.80±0.63 but residual tonal errors.","lead":"Researchers built the first text-to-speech systems for Efik, a tonal Nigerian language, using a new three-hour single-speaker corpus and four neural models. The best system (MMS-TTS) scored 3.80 MOS and stayed coherent for longer audio, giving a practical baseline for digital preservation of under-resourced African languages.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged evaluation limits; the baseline ranking claim holds under the paper's stated scope.","rationale":"The paper's strongest claim is carefully scoped as a first low-resource baseline rather than a definitive ranking of architectures for all tonal African languages. The evaluation design (five raters, subjective MOS only, single speaker) is a genuine limitation, but it is the same limitation the reader already elevated to weakest_assumption and used to justify CONDITIONAL rather than ACCEPT. No stronger load-bearing flaw (e.g., non-reproducible splits, unacknowledged data contamination, or claims that exceed the evidence) appears on re-reading §4–§7. Therefore the reader's verdict, confidence, and diagnosis stand; no adjustment is required.","tokens_in":10812,"tokens_out":494,"duration_ms":5835,"concrete_test":"Independently re-score the same held-out test subset (Table 1: 393 utterances) with ≥15 native raters plus a simple tone-error rate (manual high/low/mid mis-realization count on a 50-utterance stratified sample); if MMS-TTS remains highest on both MOS and tone accuracy and still produces the longest coherent generations, the ranking claim is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is modest and internally consistent: under a 3-hour single-speaker Efik regime, MMS-TTS is the strongest of the four systems by native MOS (3.80 ± 0.63) and long-form stability (~3 min coherent vs. 20–30 s degradation for the others). The reader's weakest_assumption correctly identifies the softest point—only five raters, short clips, no objective/tone-error metrics, single speaker—but these are already disclosed in §6.1, Table 2, §9 Limitations, and the abstract itself (\"tonal errors persisted\"). They do not create an internal contradiction or reverse the reported ordering; they simply make the absolute scores provisional. No hidden assumption, data-leakage risk, or architectural misapplication undermines the existence of the first documented Efik TTS baseline or the relative ranking under the stated constraints.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces the first documented end-to-end TTS systems for Efik, a low-resource tonal language of Southeastern Nigeria. It releases a curated single-speaker corpus of 2,632 utterances (~3.08 h) drawn from novels, folktales and educational texts, with manual orthographic transcription, native-speaker/linguist validation, 16 kHz preprocessing, and a 1975/264/393 train/val/test split. Four neural models (VITS, MMS-TTS, SpeechT5, Orpheus-TTS) are fine-tuned under identical low-resource constraints; native-speaker MOS/Nat-MOS/A-MOS (five raters) and qualitative long-form observations are reported. MMS-TTS obtains the highest scores (MOS 3.80 ± 0.63) and the most stable long-form output (~3 min coherent), while the others score lower and degrade after 20–30 s; residual tonal and rare-phoneme errors are noted for all systems. The work is framed as a reproducible baseline that underscores the need for larger multi-speaker data and tone-aware modeling.","tokens_in":11023,"tokens_out":1312,"duration_ms":36192,"significance":"If the corpus and relative ranking hold, the paper supplies a useful first public resource and empirical baseline for Efik speech synthesis, advancing digital preservation of an underrepresented African language. Concrete strengths include the carefully documented data-collection pipeline (consent, acoustic control, manual validation, explicit splits and preprocessing), transparent hyper-parameter listings, multi-metric subjective evaluation, an honest Limitations section, and an ethics statement that restricts release to non-commercial research use. The comparative bake-off under a shared 3-hour single-speaker regime is informative for the low-resource TTS community. Absolute naturalness remains modest and evaluation is limited, yet the existence of the resource and the clear ordering under the stated constraints constitute a genuine contribution.","major_comments":[{"comment":"The MOS/Nat-MOS/A-MOS results rest on only five native raters scoring short clips, with no inter-rater reliability statistic and no objective acoustic or tone metrics (MCD, F0 correlation, tone-error rate, ASR intelligibility). While the relative ordering is unambiguous and the limitations are disclosed, this sample size and purely subjective protocol render the absolute scores and the long-form superiority claim for MMS-TTS provisional; expanding the listener pool or adding at least one objective measure would materially strengthen the comparative conclusions that form the paper’s central empirical claim.","section":"§6.1 and Table 2"},{"comment":"The manuscript never states whether the manually produced orthographic transcripts contain tone diacritics. Because Efik is tonal, residual tonal errors are repeatedly highlighted, and the abstract calls for “tone-aware modeling,” the presence or absence of tone marks in the supervision is load-bearing for interpreting model failures and for reproducibility of the baseline. This must be clarified explicitly (and, if marks are absent, the implication for tone learning should be discussed).","section":"§4.1–4.2 (Data Labeling / Validation)"},{"comment":"Long-form stability (MMS-TTS coherent to ~3 min; others collapse after 20–30 s) is asserted only as a qualitative observation. Given that this is presented as a distinguishing advantage of MMS-TTS, a more systematic protocol—e.g., MOS or intelligibility ratings on held-out long utterances, or a simple hallucination/collapse rate—would better support the claim.","section":"§6 (long-sequence generation paragraphs)"}],"minor_comments":[{"comment":"Figures 1 and 2 are described (duration and word-length histograms) but their visual content is not present in the supplied text; ensure they appear with clear axis labels, bin widths and sample counts in the final version.","section":"§4.4 / Figures 1–2"},{"comment":"The exact procedure used to extend the vocabularies/embeddings of MMS-TTS, VITS and SpeechT5 for the characters ọ and ñ is mentioned only in passing; a short appendix or footnote with the mapping and any random-initialization details would improve reproducibility.","section":"§6"},{"comment":"Several reference entries carry future-dated arXiv identifiers (e.g., 2602.02734, 2603.14873) and the Orpheus-TTS citation is only a GitHub URL; verify metadata and, where possible, supply a more stable bibliographic record.","section":"References"},{"comment":"Minor orthographic inconsistencies appear (“W AXAL”, “Lagunda”, spacing around “o .”, “Nat-MOS” vs. “Nat MOS”). A careful proof-reading pass will remove them.","section":"Throughout"},{"comment":"The single-speaker limitation and the 60–100 ms trailing-silence heuristic are well motivated, yet a one-sentence note on whether any automatic silence detection or energy threshold was used would help others replicate the preprocessing exactly.","section":"§4.3"}],"recommendation":"minor_revision","confidential_remarks":"Evaluation is thin by contemporary TTS standards (five raters, no objective metrics). The paper is therefore a stronger fit for a low-resource or African-language special issue / workshop track than for a top-tier venue that expects larger listening tests. Novelty of the language and the carefully documented corpus remain high, so I do not recommend rejection; minor revision with the clarifications above should suffice."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first documented end-to-end TTS study for Efik. They built a carefully validated single-speaker corpus (2,632 utterances, ~3 h) from novels, folktales and educational texts, with manual transcription after forced alignment failed, native-speaker plus linguist checks for tone and alignment, and a clean train/val/test split. Then they fine-tuned VITS, MMS-TTS (from a Yoruba checkpoint with vocabulary extension), SpeechT5 and Orpheus-TTS and had five native listeners score MOS / Nat-MOS / A-MOS. MMS-TTS wins at 3.80 ± 0.63, stays coherent for ~3 minutes, and the others degrade after 20–30 s; VITS collapses. That ranking is new for this language and matches what we already know about multilingual pretraining helping low-resource tonal settings.\n\nWhat they did well is the data work and the honesty. The preprocessing notes (trailing silence kept for final tones, orthographic normalization, rare-phoneme struggles with ñ) are practical and useful. Limitations section and abstract both flag that tonal errors persist and that the corpus is tiny and single-speaker. Citations cover the relevant African TTS literature (BibleTTS, WAXAL, Yoruba/Hausa/Swahili work) without padding. No circularity; it is a straightforward empirical bake-off.\n\nSoft spots are real but already disclosed and proportionate: only five raters, no inter-rater stats, no objective metrics (MCD, tone error rate, ASR intelligibility), no multi-speaker prosody, and data/code not yet public at submission. Absolute MOS numbers are therefore provisional; the relative ordering under the stated 3-hour regime is still credible. Hyperparameters are listed, so the experiment is at least describable.\n\nThis is for people working on low-resource African speech or digital language preservation who need a concrete starting point rather than a new architecture. It is not a methods breakthrough, but it is a usable community baseline. I would send it to peer review; a serious referee can push for more raters, tone-specific analysis and data release without the paper needing to be reinvented. Worth engaging if you care about the language or the low-resource TTS pattern.","headline":"First real Efik TTS baseline: solid 3-hour single-speaker corpus and a clean four-model MOS bake-off that ranks MMS-TTS highest under genuine low-resource constraints.","tokens_in":11712,"tokens_out":578,"would_cite":true,"duration_ms":5856,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"The first end-to-end TTS systems for Efik show that MMS-TTS is the strongest of four neural models on a new three-hour single-speaker corpus, with MOS 3.80, yet tonal errors remain.","keywords":["text-to-speech","low-resource languages","Efik language","tonal languages","speech synthesis","African languages","MOS evaluation","MMS-TTS"],"falsifier":"Collect objective tone-contour error rates or ASR-based intelligibility scores on the same test set, or re-run the MOS study with a larger rater panel and multi-speaker data; if MMS-TTS no longer ranks first or long-form coherence collapses, the central ranking claim fails.","tokens_in":11675,"feed_emoji":"🔊","tokens_out":989,"duration_ms":9797,"temperature":0.7,"pith_summary":"Efik is a tonal language of Southeastern Nigeria with millions of speakers but no prior public speech-synthesis systems. This paper builds the first documented end-to-end text-to-speech pipeline for it by releasing a carefully recorded, manually transcribed single-speaker corpus of 2,632 utterances (about three hours) and fine-tuning four modern neural models under that severe data constraint. Native listeners rate the systems with MOS, naturalness, and accent scores; MMS-TTS leads with MOS 3.80 and can produce coherent speech lasting several minutes, while VITS, SpeechT5, and Orpheus-TTS score lower and often collapse after 20–30 seconds. The work supplies a reproducible baseline that shows current architectures can already generate intelligible Efik, yet also makes plain that larger multi-speaker data and explicit tone modeling will be required before synthesis can reliably preserve lexical meaning and cultural sound.","feed_headline":"First Efik TTS: MMS-TTS leads at MOS 3.80 on 3-hour data","feed_subtitle":"Four neural models ranked by native listeners; tonal accuracy and long-form speech still need larger corpora.","key_machinery":"A curated single-speaker Efik corpus of 2,632 manually validated utterances (≈3 hours) used to fine-tune four low-resource neural TTS models (VITS, MMS-TTS initialized from Yoruba, SpeechT5, Orpheus-TTS), with ranking performed by native-speaker MOS, Nat-MOS, and A-MOS ratings.","core_discovery":"Under a three-hour single-speaker regime, MMS-TTS is the strongest of the four evaluated neural TTS systems for Efik, achieving the highest MOS (3.80 ± 0.63), Nat-MOS (3.60), and A-MOS (3.04) from five native raters and generating continuous speech up to roughly three minutes without hallucination, whereas VITS, SpeechT5, and Orpheus-TTS score lower and degrade after 20–30 seconds; this constitutes the first documented end-to-end TTS baseline for the language.","pith_inferences":["The same three-hour single-speaker recipe could be applied immediately to neighboring Lower Cross languages that already have small text resources but lack speech synthesis.","Because tone errors still alter meaning, future evaluation protocols for these languages should include forced-choice lexical-tone discrimination tasks rather than relying solely on overall MOS.","The observed foreign-accent residual in Orpheus-TTS and SpeechT5 suggests that cross-lingual transfer can introduce speaker-identity leakage that multi-speaker fine-tuning or speaker-embedding conditioning might later suppress."],"forward_implications":["A public, reproducible Efik TTS baseline now exists that later systems can be measured against.","Multilingual pretraining (as in MMS-TTS) is currently the most practical route for intelligible synthesis of related low-resource tonal languages with only a few hours of data.","Long-form generation remains unreliable for most architectures under three-hour single-speaker conditions, so practical applications will need additional data or architectural safeguards.","Tonal and rare-phoneme errors persist even in the best model, confirming that larger corpora and tone-aware modeling are required before synthesis can safely preserve lexical meaning."],"fun_headline_variants":["MMS-TTS leads Efik TTS with MOS 3.80 on 3-hour single-speaker data","First Efik end-to-end TTS: MMS-TTS tops four models at MOS 3.80","Efik low-resource TTS baseline ranks MMS-TTS highest by native MOS","MMS-TTS strongest of four neural systems for tonal Efik speech","Three-hour Efik corpus: MMS-TTS yields best MOS and longer speech"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Subjective scores from only five native listeners on short clips, without objective tone-error or intelligibility metrics, are enough to rank the models and claim relative suitability for long-form and tonal Efik speech.","fun_headline_variants_meta":{"raw":{"variants":["MMS-TTS leads Efik TTS with MOS 3.80 on 3-hour single-speaker data","First Efik end-to-end TTS: MMS-TTS tops four models at MOS 3.80","Efik low-resource TTS baseline ranks MMS-TTS highest by native MOS","MMS-TTS strongest of four neural systems for tonal Efik speech","Three-hour Efik corpus: MMS-TTS yields best MOS and longer speech"]},"model":"grok-4.5","effort":"low","cost_usd":0.00593,"raw_usage":{"total_tokens":1523,"prompt_tokens":754,"num_sources_used":0,"completion_tokens":118,"cost_in_usd_ticks":59300000,"prompt_tokens_details":{"text_tokens":754,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":651,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":754,"tokens_out":118,"duration_ms":5704,"temperature":1.0,"reasoning_tokens":651,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T18:10:24.761888+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Collect objective tone-contour error rates or ASR-based intelligibility scores on the same test set, or re-run the MOS study with a larger rater panel and multi-speaker data; if MMS-TTS no longer ranks first or long-form coherence collapses, the central ranking claim fails.","supporting_citations":[],"review_version":1}