{"id":"2165772e-29d3-4bc8-bb9a-ef96f976881d","arxiv_id":"2506.17459","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Fine-tuned MMS outperforms XLS-R on fieldwork ASR with less than one hour of training data, while XLS-R reaches parity beyond one hour.","lead":"This paper benchmarks two multilingual speech recognition models, MMS and XLS-R, fine-tuned on five low-resource fieldwork languages using only 10 to 120 minutes of transcribed audio. It finds that MMS performs better with under an hour of data, while XLS-R catches up once about an hour is available, a practical result for field linguists facing a severe transcription bottleneck.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The crossover claim rests on single runs with no confidence intervals; at 10–60 min the MMS advantage could be sampling noise, and the one-hour threshold is extrapolated from a single 120-min point.","rationale":"The reader's stated weakest assumption is that random subsets from ELAR archives represent true fieldwork scarcity. I agree this matters, but it is not the first thing that would falsify the central claim, because both models are trained and evaluated on the same subsets; a non-representative draw would inflate or deflate absolute CER for both models and only indirectly perturb the model comparison. The more direct threat is that the comparison itself is reported as single point estimates. Figure 1 is the entire evidence for the crossover, and the caption disclaims interpolation between the plotted durations, yet the one-hour threshold is exactly such an interpolation. Section 7 lists several limitations but never states that no variance estimate or significance test was computed. A conditional accept is still appropriate: the work is a useful benchmark with honest data description, and the required fix is feasible—add confidence intervals or bootstrap/permutation evidence, and either add an intermediate duration or soften the precise 'one hour' wording. If the re-runs show a stable, significant MMS advantage below one hour and a stable crossover near one hour, the paper's recommendation stands; if not, the headline needs to be re-framed as exploratory.","tokens_in":13548,"tokens_out":8122,"duration_ms":100788,"concrete_test":"Re-run the full comparison with 5 independent random subset draws at each duration and report bootstrap 95% CIs on the MMS−XLS-R CER difference (resampling test utterances within each language, then languages), plus a permutation test of whether the under-one-hour MMS advantage is significant. If at 10, 30, or 60 minutes the difference is not significant, or if the crossover point moves by more than 30 minutes across draws, the headline 'one hour threshold' recommendation should be weakened or re-framed as exploratory.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 and Figure 1 base the headline recommendation on one fine-tuning run per model per language per duration: 5 languages × 4 durations, with no error bars, confidence intervals, or significance tests in the paper. The per-language CER differences could be within run-to-run, subset-sampling, or test-set noise, especially since each test set is a fixed 10-minute sample (§2.2) and no held-out speaker is used. The plotted crossover is therefore not established as a systematic effect. The claimed one-hour threshold is even less directly supported: training durations are 10, 30, 60, and 120 minutes, so 'parity once data exceed one hour' is inferred from a single 120-minute condition; the Figure 1 caption itself states the connecting lines do not imply intermediate performance. Section 7 acknowledges shared speakers and limited languages but does not address the absence of variance estimates. This is a robustness gap, not an internal inconsistency: the trends may be real, but the central practical claim currently rests on point estimates that could shift with a different subset draw or test set. The subset-representativeness concern raised in §2.2 is real but secondary, because both models train on identical subsets; a biased draw would affect absolute CER more than the model comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares fine-tuned MMS-1B-l1107 and XLS-R-300m on ASR for five low-resource fieldwork languages from ELAR, using controlled amounts of 10, 30, 60, and 120 minutes of training data and a fixed 10-minute test set per language. It reports CER and WER, concludes that MMS is preferable below one hour while XLS-R reaches parity above one hour, and adds phonologically informed error analysis for tone, nasality, and vowel/consonant length in Cicipu and Mocho’.\n\nThe central quantitative claim is plausible but not yet fully supported: the crossover conclusion rests on single-run point estimates with no variance quantification, and the one-hour threshold is extrapolated from a single 120-minute condition. The paper is a useful empirical contribution for the fieldwork community if these robustness gaps are addressed.","tokens_in":13722,"tokens_out":6211,"duration_ms":73494,"significance":"Strengths of the paper include its use of real fieldwork recordings with environmental noise and spontaneous speech, its coverage of five typologically diverse languages, its choice of CER as the primary metric for languages without standardized orthographies, its candid limitations section, and its attempt to check for a possible Kichwa pre-training advantage with a linear mixed-effects model. The phonological error analysis also goes beyond the usual global error rates and gives field linguists information about tone, nasality, and length contrasts.\n\nIf the crossover trend is reproducible, the paper would provide a simple, actionable rule for practitioners and a useful benchmark for future low-resource ASR work. However, the headline recommendation is currently based on a small number of point estimates without error bars, confidence intervals, or significance tests, and the abstract/conclusion wording is stronger than the displayed evidence. The practical guidance therefore needs additional statistical support or a more cautious framing before it can be adopted with confidence.","major_comments":[{"comment":"The central claim that MMS “consistently achieves lower error rates” below one hour and that XLS-R reaches parity above one hour is based on a single fine-tuning run per model per language per duration (5 languages × 4 durations × 2 models) with no confidence intervals, error bars, or significance tests. Because the test set is a fixed 10-minute sample (Section 2.2) and no held-out speaker is used, the observed per-language differences could be within run-to-run or subset-sampling noise. Section 7 acknowledges shared speakers and the small number of languages but does not mention the absence of variance estimates. Please provide variance estimates (e.g., multiple seeds, bootstrap over test segments, or a significance test) or soften the claim to describe the observed runs; as written, the practical recommendation is not statistically supported.","section":"Section 4.1, Figure 1; Section 7"},{"comment":"The “one-hour threshold” is not directly measured. Training durations are 10, 30, 60, and 120 minutes, and the Figure 1 caption explicitly states that the connecting lines do not imply performance at intermediate durations. The abstract and Section 4.1 nevertheless conclude that “once training data exceed one hour” XLS-R reaches parity, and Section 4.1 further says XLS-R “becomes a more effective option” with approximately one hour or more. The only evidence past one hour is the single 120-minute condition. Please add intermediate points (e.g., 90 minutes) or rephrase the conclusion as “at two hours in these runs” to avoid over-generalizing from one data point.","section":"Abstract; Section 4.1; Figure 1 caption"},{"comment":"The low-resource scenario is simulated by drawing 10–120 minute subsets from archives that contain up to 22.84 hours for Toratán and several cleaned hours for the other languages (Table 6). The paper does not specify how the subsets were selected: random draws, first N utterances, or balanced-by-speaker/genre sampling. If the subsets are random draws from a larger archive, the experimental condition is not equivalent to having only 10 minutes of newly collected fieldwork audio, because the full archive’s speaker and genre coverage is known to the experimenter. This does not invalidate the model comparison, since both models train on identical subsets, but it does affect the external validity of the “extremely low-resource fieldwork” recommendation. Please specify the selection procedure and discuss how this affects the practical guidance.","section":"Section 2.2; Table 6"},{"comment":"The abstract says XLS-R “shows parity performance once training data exceed one hour,” while Section 4.1 says XLS-R “becomes a more effective option” when approximately one hour or more is obtainable. Moreover, in Figure 1, MMS appears still lower than XLS-R in several languages at the 120-minute point. The paper should align the abstract, Section 4.1, the conclusion, and the figure, and should avoid claiming a clear advantage for XLS-R beyond one hour when the displayed point estimates do not consistently show it.","section":"Abstract vs. Section 4.1; Figure 1"}],"minor_comments":[{"comment":"The phrase “for further provide insights towards practical guidelines” is grammatically incomplete and should be revised to “and provide insights toward practical guidelines.”","section":"Abstract"},{"comment":"The fine-tuning procedure is credited only to “von Platen” via a footnote; please add a full bibliographic reference for the MMS adapter recipe.","section":"Section 3.2"},{"comment":"No numeric CER or WER table is provided in the main text. Adding a compact table of per-language CER at each duration would make the trends in Figure 1 auditable and facilitate comparison with future work.","section":"Section 4; Figure 1"},{"comment":"The linear mixed-effects model is underspecified: the text does not define the “time” variable, whether it is log-transformed, or whether random slopes or interactions beyond the reported one were considered. With only five languages, the non-significant interaction (β = 0.021, p = 0.896) should be described as an underpowered check rather than strong evidence of no Kichwa-specific benefit.","section":"Section 4.2"},{"comment":"Table 3 does not state which training duration the phonological error rates correspond to, whereas Table 4 explicitly reports the XLS-R 120-minute model. If Table 3 is also 120-minute-only, the claim that “both models struggle” with these phonological categories should be restricted to that condition and not generalized across all data sizes.","section":"Section 4.3.1; Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is an empirical benchmark with no code release or seed specification; for a reproducibility-conscious venue, the editor may want to request code, random seeds, and exact subset-selection scripts. The comparison is also limited to two specific model checkpoints (MMS-1B-l1107 and XLS-R-300m), which is reasonable for the fieldwork context but should not be presented as a general model-selection law beyond the tested conditions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful, honestly written empirical benchmark comparing MMS and XLS-R for fieldwork ASR on five languages with controlled training durations. The headline claim—MMS is better under an hour, XLS-R reaches parity after an hour—is plausible but not yet established, because the comparison rests on single fine-tuning runs with no variance estimates.\n\nWhat is actually new: the cross-language controlled-duration comparison is a real step beyond the single-language studies that dominate this niche, and using spontaneous, noisy ELAR recordings rather than clean read speech is the right call. The choice of CER over WER is well argued for languages without stable orthographies, and the phonological error analysis for Cicipu and Mocho' gives useful detail on where these models fail. The limitations section is also candid about shared speakers and the small language sample.\n\nThe soft spots are exactly where the stress-test note points. Figure 1 and the Section 4.1 prose assert a systematic crossover, but there are no error bars, confidence intervals, or significance tests. Each language×duration cell is one fine-tuning run on one subset draw, and test sets are fixed 10-minute samples from the same speakers. At the 10–60 minute range the MMS advantage could easily be sampling noise. The \"one hour\" threshold is even thinner: it is inferred from a single 120-minute point, and the figure caption itself warns that connecting lines do not predict intermediate behavior. The linear mixed-effects model in Section 4.2 only tests the Kichwa pre-training interaction; it does not test model×duration, so the central claim remains unsupported statistically. The subset-representativeness concern from the reader's report is real but secondary for the model comparison, since both models see identical subsets; it matters more for absolute CER and generalization to new field recordings.\n\nThis paper deserves a serious referee: the question matters, the data are real, and the design is a step up from prior single-language work. But I would ask for a variance analysis—repeated runs or bootstrap CIs on the main comparison—and a more careful statement of the crossover claim before publication. Useful reading for field linguists and low-resource ASR people; treat the one-hour threshold as a rule of thumb, not a measured fact.","headline":"Useful five-language benchmark of MMS vs XLS-R for fieldwork ASR, but the crossover claim needs variance estimates; worth reviewing with revisions.","tokens_in":14324,"tokens_out":2462,"would_cite":true,"duration_ms":27099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning experiments on five fieldwork languages show MMS beats XLS-R under one hour of transcribed speech, with XLS-R reaching parity once data pass one hour.","keywords":["automatic speech recognition","low-resource ASR","fieldwork languages","MMS","XLS-R","fine-tuning","endangered language documentation","character error rate"],"falsifier":"Fine-tune MMS and XLS-R on 10, 30, 60, and 120 minutes of a new fieldwork language with a test set of speakers never heard during training; if XLS-R no longer trails below one hour, or if MMS's lead at 10 minutes disappears, the paper's ranking and threshold claim is refuted.","tokens_in":13255,"feed_emoji":"🎙️","tokens_out":7169,"duration_ms":67370,"temperature":0.7,"pith_summary":"A single hour of fieldwork audio can require up to 50 hours of manual transcription, so choosing an ASR model that works with almost no transcribed data matters for language documentation. The paper tests two fine-tunable multilingual models, MMS and XLS-R, on five typologically varied languages using real, noisy, spontaneous field recordings and controlled training durations of 10, 30, 60, and 120 minutes. Its central claim is a decision rule: MMS gives lower character error rates when less than one hour of transcribed data is available, while XLS-R matches MMS once training data reach about one hour. The paper also shows that both models continue to miss tone, nasality, and vowel-length contrasts, so those categories need targeted treatment beyond model choice.","feed_headline":"One hour is the tipping point between two low-resource ASR models","feed_subtitle":"With under an hour of transcribed speech, MMS wins; with an hour or more, XLS-R matches it","key_machinery":"The comparison is carried by two wav2vec 2.0-based models adapted in different ways: MMS-1B-l1107, a 1-billion-parameter model with frozen base weights and a 2-million-parameter trainable adapter, and XLS-R-300m, a 300-million-parameter model that is fully fine-tuned. Training is controlled by constructing superset splits of 10, 30, 60, and 120 minutes per language, with a fixed 10-minute test set and Character Error Rate as the primary metric, since CER tracks phoneme-level accuracy better than word error rate for languages without standardized orthographies. This design isolates data duration as the variable that separates MMS's advantage from XLS-R's parity.","core_discovery":"With under one hour of transcribed fieldwork speech, fine-tuned MMS achieves lower Character Error Rate than fine-tuned XLS-R across the five test languages; with one hour or more, XLS-R reaches parity. The authors attribute the early advantage to MMS's pre-training on more than a thousand languages and to its built-in ASR fine-tuning plus adapter layers, and XLS-R's catch-up to its more conversationally diverse pre-training corpus. They further show that the advantage is not explained by MMS having seen related Kichwa dialects, since a linear mixed-effects model finds no significant extra gain for Upper Napo Kichwa. Phonologically informed error analysis on Cicipu and Mocho' finds both models struggle with tone, nasality, and consonant/vowel length, with deletions dominating nasality errors and substitutions dominating length errors.","pith_inferences":["The one-hour threshold is probably not a universal constant; languages with larger orthographic inventories (Cicipu's 93 characters) or with no related variety in the pretraining data may need more data before XLS-R catches up, and languages closer to pretraining data may need less.","Because training and test segments draw on the same speakers, the reported parity point may be optimistic; a held-out-speaker evaluation could push both models' error rates up and possibly change which model leads.","A testable extension is continued pre-training on untranscribed field audio: the paper's own reading of MMS's adapter advantage suggests that adding in-language unlabeled data could lower the one-hour threshold for both models.","The superset training design means the 120-minute condition contains the 10-minute condition; comparing models under strictly disjoint data draws would more directly test whether the advantage is about quantity or about which utterances are included."],"forward_implications":["Field linguists with under an hour of transcribed material should fine-tune MMS first, since it gives lower character error rates under that threshold.","The one-hour mark is a practical decision point: beyond it, XLS-R is competitive and may be preferable because its full fine-tuning uses a more conversationally diverse pre-training corpus.","Character error rate is the right evaluation target for documentation work, where phonetic accuracy matters more than word-level output.","Tone and nasality transcription will not be fixed by model choice alone; targeted augmentation, adapter design, or loss functions are needed for these categories.","The observed plateau after roughly one hour suggests that additional transcription effort has diminishing returns for fine-tuning these models."],"supporting_citations":[{"why":"Supplies the MMS model and its adapter-based ASR fine-tuning, the model that wins below one hour.","marker":"[Pratap et al., 2024]"},{"why":"Supplies XLS-R, the 128-language self-supervised baseline that reaches parity after one hour.","marker":"[Babu et al., 2021]"},{"why":"Provides the wav2vec 2.0 self-supervised framework on which both MMS and XLS-R are built.","marker":"[Baevski et al., 2020]"},{"why":"Introduces the adapter layers that let MMS fine-tune on tiny data without updating the full base model.","marker":"[Houlsby et al., 2019]"},{"why":"Reports MMS outperforming XLS-R on Mvskoke, the direct precedent for the paper's ranking.","marker":"[Mainzinger and Levow, 2024]"},{"why":"Documents diminishing returns in low-resource fine-tuning, supporting the one-hour plateau interpretation.","marker":"[Guillaume et al., 2022a]"},{"why":"Finds no single ASR architecture consistently wins under extreme scarcity, motivating the controlled comparative benchmark.","marker":"[Jimerson et al., 2023]"},{"why":"Documents Napo Quichua varieties, used to test whether MMS's advantage for Upper Napo Kichwa could be explained by pretraining overlap.","marker":"[Eberhard et al., 2024]"}],"fun_headline_variants":["MMS beats XLS-R under an hour of fieldwork speech","One hour of data flips the best ASR model for fieldwork","Fieldwork ASR: MMS wins with <1 hour, XLS-R ties after","Tipping point: 1 hour decides best low-resource ASR model","Under 1 hour, MMS beats XLS-R; at 1 hour, tie"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim assumes that a 10-to-120-minute slice of an existing archive behaves like all the data a field linguist would have for a new language, even though the full archive (up to 22 hours) and its speakers were available when the slice was selected.","fun_headline_variants_meta":{"raw":{"variants":["MMS beats XLS-R under an hour of fieldwork speech","One hour of data flips the best ASR model for fieldwork","Fieldwork ASR: MMS wins with <1 hour, XLS-R ties after","Tipping point: 1 hour decides best low-resource ASR model","Under 1 hour, MMS beats XLS-R; at 1 hour, tie"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000632,"raw_usage":{"total_tokens":2873,"prompt_tokens":856,"completion_tokens":2017,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":1925}},"tokens_in":472,"tokens_out":2017,"duration_ms":16048,"temperature":1.0,"reasoning_tokens":1925,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:29:21.556186+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fine-tune MMS and XLS-R on 10, 30, 60, and 120 minutes of a new fieldwork language with a test set of speakers never heard during training; if XLS-R no longer trails below one hour, or if MMS's lead at 10 minutes disappears, the paper's ranking and threshold claim is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MMS model and its adapter-based ASR fine-tuning, the model that wins below one hour."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the wav2vec 2.0 self-supervised framework on which both MMS and XLS-R are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the adapter layers that let MMS fine-tune on tiny data without updating the full base model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports MMS outperforming XLS-R on Mvskoke, the direct precedent for the paper's ranking."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Finds no single ASR architecture consistently wins under extreme scarcity, motivating the controlled comparative benchmark."}],"review_version":1}