{"id":"c9770541-2d09-46b0-a117-15cffbb81da7","arxiv_id":"2607.21540","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.","lead":"DONDO releases 26 open ASR models (21 monolingual, 5 multilingual) covering 27 African language varieties, fine-tuned from w2v-BERT 2.0 on read religious speech. The models are Apache-2.0 licensed, with multilingual checkpoints reaching 10–13% average WER after learning-rate annealing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"WER averages in Tables 2–6 are not reproducible from per-language values; without reported weights or corrected arithmetic, the headline 10–13% claim is unsupported.","rationale":"The reader's verdict is REJECT and I do not think it should change. The reader's stated weakest assumption was limited/under-reported test sets, but the most immediately load-bearing problem is internal: the Avg columns in the central results tables are inconsistent with the per-language numbers displayed alongside them. This is not a matter of disagreement with external baselines or a contested evaluation protocol; it is a checkable arithmetic/weighting issue within the paper's own evidence. The missing test-set details (which the reader emphasized) are closely related: if the Avg is weighted by test-set size, those sizes are essential and absent; if it is unweighted, the arithmetic is wrong. Either way, the 10–13% headline is not reproducible. The Sierra Leone family's missing Step-2 evaluation is a separate but compounding gap: the paper's main methodological contribution (two-step annealing) is claimed across five families but shown for only four. The released Apache-2.0 models are a tangible contribution and the authors' intent to provide an open test bed is valuable, but the paper must supply corrected tables, complete evaluations, and test-set metadata before its central quantitative claims can be accepted. Since the reader already reached REJECT, my assessment leaves the verdict unchanged.","tokens_in":7214,"tokens_out":6087,"duration_ms":49046,"concrete_test":"Recompute the Avg column in Tables 2–6 from the per-language WERs using the per-language test-set sizes, which must be reported in a revision. If unweighted, verify arithmetic; e.g., Table 2 Step 2 should show 11.98 (excluding English) or 13.91 (including English), not 11.4. If weighted, state the weights and confirm the weighted means exactly reproduce the printed Avg values. Also run the Step-2 annealing for the Sierra Leone family and report per-language WERs; if no table is provided, the claim that annealing recovers the gap remains unestablished for that family.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — that annealed multilingual models reach average WERs of 10–13% and close most of the gap to monolingual baselines — depends entirely on the Avg columns in Tables 2–6. These columns do not match the per-language WERs shown in the same tables under any obvious unweighted mean. For example, Table 2 Step 2 lists Eng 27.4, Adangme 8.78, Ewe 6.57, Fante 13.4, French 3.64, Ga 16.0, Nzema 20.8, As. Twi 14.7: the mean excluding English is 11.98, including English 13.91, neither is 11.4. Step 1 in the same table gives 15.94 (excl. Eng) or 18.31 (incl. Eng), not 14.7. Table 3 Step 2 gives 12.30/13.63, not 10.3; Table 4 Step 1 gives 9.47/12.61, not 6.06; Table 6 Step 3 gives 11.21/17.02, not 12.5. These discrepancies could be explained only by a weighted average over per-language test-set sizes, but no weights or test-set sizes are reported. Additionally, Table 4 (Sierra Leone) reports only Step 1; the annealing step that is the paper's main recipe is never evaluated for one of the five families. As written, the headline results cannot be reconstructed from the evidence presented, so the central claim is not supported by the paper's own data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DONDO, a family of open Apache-2.0-licensed ASR base models for 27 African language varieties, built by fine-tuning the w2v-BERT 2.0 encoder. It releases 21 monolingual and 5 multilingual checkpoints, and proposes two experimental techniques: a two- or three-step learning-rate schedule for multilingual fine-tuning, and a prefix-frame language-conditioning mechanism that prepends a one-hot language vector to the acoustic features. The central quantitative claim is that the annealed multilingual models reach average WERs of 10–13% across five regional families, closing most of the gap to monolingual baselines and occasionally surpassing them. The paper also gives a conservative estimate of the potential human reach of the covered languages.","tokens_in":7631,"tokens_out":7898,"duration_ms":74745,"significance":"If the results are reproducible, this is a valuable community resource: permissively licensed ASR base models for low-resource African languages, with a simple and portable fine-tuning recipe and a lightweight language-conditioning method. The explicit release of model checkpoints, the focus on license-clear religious read-speech data, and the honest limitations discussion are strengths. The reach analysis is clearly framed as an estimate. However, the paper's empirical backbone is weakened by internal inconsistencies in the reported average WERs and by missing evaluation metadata; these issues must be resolved before the engineering contributions can be assessed.","major_comments":[{"comment":"The Avg columns are not reproducible from the per-language WERs in the same tables under any obvious unweighted mean. For example, Table 2 Step 2 lists per-language WERs whose mean is 13.9% including English and 12.0% excluding English, not 11.4%; Table 3 Step 2 gives 13.6%/12.3%, not 10.3%; Table 4 Step 1 gives 12.6%/9.5%, not 6.06%; Table 6 Step 3 gives 17.0%/11.2%, not 12.5%. Some rows (e.g., Table 5 Step 2) are close to the unweighted excl-English mean, but the pattern is not consistent. If the Avg is a weighted mean over per-language test-set sizes, those weights and sizes must be reported; otherwise the headline 10–13% claim is unsupported.","section":"Section 5, Tables 2–6"},{"comment":"Only the Step 1 checkpoint is evaluated for the Sierra Leone family; the two-step annealing recipe that constitutes the paper's main methodological contribution is never applied or evaluated for one of the five families. The abstract and Section 5.1 nonetheless state that annealed models across the five multilingual families reach 10–13% WER. Either run and report the Step 2 (and Step 3, if used) evaluation for Sierra Leone, or explicitly restrict the claim to the four families with annealed results.","section":"Section 5, Table 4"},{"comment":"The evaluation protocol is underspecified: no corpus names, total hours per language, train/test split procedure, speaker overlap policy, or test-set sizes are given. The paper itself states in Section 7 that 'evaluation for the smallest languages rests on limited test sets.' Without this information, the WER numbers cannot be independently reproduced or compared across languages, and the claimed gap closure against monolingual baselines may be an artifact of test-set composition. Full evaluation metadata is required.","section":"Sections 3.2 and 5; Section 7 Limitations"},{"comment":"The prefix-frame language-conditioning mechanism is described only loosely: a one-hot vector is 'map[ped] into the feature dimension' and repeated over p prefix frames, but the projection (learned linear layer? fixed embedding? normalization?) is not defined, and the value of p is not reported. As this is one of the two stated contributions, the paper should specify the implementation and provide an ablation comparing the prefix-conditioned model against the same model without prefixes or with random prefixes, to demonstrate that the mechanism is actually responsible for the reported cross-lingual steering.","section":"Section 4.3"}],"minor_comments":[{"comment":"The title and abstract contain spacing errors: 'Openw2v-BERTSpeech-Recognition' should be 'Open w2v-BERT Speech-Recognition'.","section":"Title/Abstract"},{"comment":"The column header is typeset as 'A vg' instead of 'Avg'.","section":"Tables 2–6"},{"comment":"The Step 1 row reports an average (21.7%) but states per-language WERs were not computed for that step; it is unclear how an average can be formed without per-language values.","section":"Table 6"},{"comment":"The schedule is called 'learning-rate annealing' but actually uses two or three fixed learning rates in sequence; the term annealing usually implies a continuous decay within a run. The number of steps/epochs per stage is also not reported, making the comparison between Step 1 and Step 2 hard to interpret.","section":"Section 4.2"},{"comment":"The paper says models are released on Hugging Face, but Table 1 lists model names without direct URLs; full repository links should be included. The w2v-BERT 2.0 backbone is also cited only via the original w2v-BERT paper; the specific 2.0 release should be cited.","section":"Section 3.1 / Table 1"},{"comment":"Population figures are said to be compiled from 'standard references,' but no sources or citations are provided for the L1/L2 estimates in Table 7.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The arithmetic inconsistencies in the Avg columns are severe enough that I would not support publication without a full re-analysis and disclosure of test-set sizes/weights. The missing Step 2 evaluation for Sierra Leone is also a substantive gap, not a presentation issue. I am recommending major_revision rather than reject because these problems appear fixable if the underlying logs and evaluation sets exist; however, if the authors cannot supply the missing metadata and corrected numbers, the paper should not be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth knowing: this paper ships a large, permissive set of ASR base models for 27 African languages, which is a genuinely useful contribution. But the main quantitative claim — that annealed multilingual models hit 10–13% average WER — is not supported by the paper's own tables. The 'Avg' columns are not the arithmetic mean of the per-language WERs in the same rows. For example, Table 2 Step 2 lists values whose mean (including English) is 13.9%, not 11.4%; Table 3 Step 2's mean is 13.6%, not 10.3%. The mismatch is systematic across all tables except one where it could be rounding. There is no reported weighting that would explain these numbers, and no test-set sizes given. On top of that, the Sierra Leone family (Table 4) only has Step 1 results, so it never tests the annealing recipe that is the paper's core idea.\n\nWhat's good: the model release is real and valuable — 21 monolingual and 5 multilingual checkpoints on Hugging Face under Apache-2.0, covering languages that mostly have no other public ASR. The fine-tuning recipe is simple and reproducible, and the prefix-frame language conditioning is a lightweight alternative to adapters. The paper honestly discusses the limits of using read religious texts and acknowledges the small test-set problem.\n\nWhere it's soft: beyond the average arithmetic, there's no information on test-set construction, size, or speaker overlap with training data. The per-language numbers still look plausible for the better-resourced languages, but for the smallest ones we have no way to assess reliability. The English column is consistently worse in multilingual models, which the paper mentions but doesn't interrogate.\n\nBottom line: this is a resource paper whose models are likely useful, but the evaluation is sloppy enough that the headline claims can't be verified. A corrected revision with accurate tables, complete evaluations, and test-set details would be worth publishing. As it stands, I wouldn't cite the WER numbers, though I might cite the model release itself. I'd send it to peer review because the contribution is substantial and the flaws are fixable. A careful referee should ask for the actual averages to be recomputed and reported correctly.","headline":"Releases many useful open ASR models, but the headline WER averages don't match its own tables; don't trust the numbers without a revision.","tokens_in":8107,"tokens_out":6121,"would_cite":false,"duration_ms":50354,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By fine-tuning a shared speech encoder first at a high learning rate on a multilingual mixture and then at a ten-fold lower rate, the authors produce open automatic speech recognition models for 27 African language varieties whose word erro","keywords":["automatic speech recognition","African languages","w2v-BERT","multilingual ASR","low-resource speech","learning-rate annealing","language conditioning","open models"],"falsifier":"For one of the smallest languages (e.g., Waali or Kasem), build an independent test set from a different source of the same language — ideally spontaneous speech or a different reader — transcribe it with human annotation, and compute the released model's WER. If the WER is substantially higher than the reported single-digit figure (or the gap to the monolingual baseline widens sharply), the claim that the recipe recovers most of the gap on usable speech is falsified for that language.","tokens_in":7118,"feed_emoji":"🗣️","tokens_out":5938,"duration_ms":53914,"temperature":0.7,"pith_summary":"The paper sets out to show that a deliberately simple fine-tuning recipe can turn a widely used self-supervised speech encoder into usable, openly licensed automatic speech recognition models for 27 African language varieties. The authors argue that training on license-clear read speech from religious texts is a practical route to coverage where transcribed audio is scarce, and that a two-step learning-rate annealing schedule recovers most of the error-rate gap between a shared multilingual model and per-language monolingual baselines. Across the five regional multilingual families, annealed models reach average word error rates of 10–13 percent, and in several languages beat the monolingual models. A lightweight prefix-frame language-conditioning mechanism lets one checkpoint be steered to a target language at inference without changing the encoder. If the numbers hold, this gives the community a reproducible, commercially usable base for further fine-tuning in languages spoken by roughly a hundred million people.","feed_headline":"Two-step anneal closes the gap to monolingual African-language ASR","feed_subtitle":"A 10x lower learning rate after coarse multilingual training recovers most of the error gap across 27 language varieties.","key_machinery":"The load-bearing object is the two-step learning-rate annealing schedule applied to the w2v-BERT 2.0 self-supervised speech encoder (a Conformer backbone with a contrastive plus masked-language-modelling pre-training objective), followed by a CTC head. Step 1 (LR 5e-5) coarsely adapts the shared encoder to the multilingual mixture; Step 2 (LR 5e-6) anneals and sharpens per-language decoding; an optional Step 3 (LR 5e-7) helps the East/Southern Africa family. The second mechanism is prefix-frame language conditioning: a one-hot language vector is mapped to the feature dimension, repeated over a short block of frames, and prepended to the acoustic input, giving the encoder a language 'prompt'","core_discovery":"The paper's central claim is that a deliberately plain recipe — fine-tune the w2v-BERT 2.0 Conformer encoder with a CTC head on read religious-text audio, first at 5e-5 on a multilingual mix, then at 5e-6 — is enough to bring five regional multilingual ASR families to average WERs of 10–13%, closing most of the gap to monolingual models and, for languages like Fante and Meru, beating them. The annealing step, not the initial adaptation, is what recovers performance: Step 1 alone inflates WER two- to three-fold on the hardest languages, while Step 2 recovers the bulk of that gap. A one-hot language identity injected as prefix frames into the acoustic feature stream lets a single checkpoint be","pith_inferences":["Beyond the paper: the same 'coarse multi-task adaptation then low-rate sharpening' pattern may be a general low-resource transfer heuristic, worth testing on other self-supervised encoders and non-speech sequence models.","Because evaluation is in-domain read speech, the released WERs are best read as an upper bound on clean-speech quality; the models' real-world utility depends on how much fine-tuning data downstream users can supply, which the paper does not quantify.","The prefix-frame conditioning sensitivity (how many frames, what happens with a learned embedding instead of a one-hot prompt) is unexplored; a small ablation study would pin down whether the mechanism is a true soft prompt or just a fixed bias.","The paper's reach estimate counts speakers, not users; if adoption is the goal, a more decision-relevant metric would be the number of people who could realistically obtain a working ASR pipeline, which depends on devices, connectivity, and downstream ecosystem — none of which the paper models."],"forward_implications":["A single annealed multilingual checkpoint can replace several monolingual models in deployment, covering a whole region at comparable WER.","The same two-step annealing schedule, without per-language hyperparameter search, is presented as a reproducible recipe that can be applied to any new language group built on this encoder.","Prefix-frame language conditioning provides a parameter-free way to steer a model to a target language at inference, which the paper argues makes a single checkpoint operationally attractive.","The open, attribution-only licence means the base models can be fine-tuned commercially, which the authors frame as lowering every barrier to domain-specific adaptation.","For lower-resource languages such as Meru (26.5% monolingual WER to 16.9% in the multilingual model), the results indicate positive cross-lingual transfer even in modest data settings."],"fun_headline_variants":["Step-down LR recovers ASR accuracy for 27 African languages","One-hot prefix steers a single ASR checkpoint to 27 African tongues","Open ASR for 100M African speakers: anneal to 10-13% WER","Religious texts enable open ASR for 27 African language varieties"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The WER numbers for the smallest languages rest on in-domain test sets whose size and construction are not reported; if those test sets are too small or not representative, the reported 10–13% averages and the comparison to monolingual baselines may not hold outside the test distribution.","fun_headline_variants_meta":{"raw":{"variants":["Step-down LR recovers ASR accuracy for 27 African languages","One-hot prefix steers a single ASR checkpoint to 27 African tongues","Open ASR for 100M African speakers: anneal to 10-13% WER","Religious texts enable open ASR for 27 African language varieties"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000629,"raw_usage":{"total_tokens":2806,"prompt_tokens":868,"completion_tokens":1938,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":1866}},"tokens_in":612,"tokens_out":1938,"duration_ms":15866,"temperature":1.0,"reasoning_tokens":1866,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:06:09.107147+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For one of the smallest languages (e.g., Waali or Kasem), build an independent test set from a different source of the same language — ideally spontaneous speech or a different reader — transcribe it with human annotation, and compute the released model's WER. If the WER is substantially higher than the reported single-digit figure (or the gap to the monolingual baseline widens sharply), the claim that the recipe recovers most of the gap on usable speech is falsified for that language.","supporting_citations":[],"review_version":1}