{"id":"43aa1c3d-f242-4a4d-b491-e4192916e762","arxiv_id":"2608.10670","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On a five-seed Garhwali ASR benchmark using the official VAANI splits, standard CTC beats Focal CTC, a matra-weighted objective, and Hindi transfer, while speed augmentation and encoder choice provide the only robust gains.","lead":"A new study benchmarks Garhwali speech recognition across five training seeds and finds that common improvements such as focal loss, vowel-diacritic weighting, and Hindi transfer do not reliably beat standard training. The authors argue that multi-seed evaluation should be standard for low-resource dialectal ASR, since a single lucky run can manufacture gains that vanish on replication.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Core multi-seed null results are well supported, but the abstract's 'pretraining design, not parameter count' claim is not identified: the four-model comparison in §5.2/Table 7 varies architecture, pretraining data, tokenizer, and scale simultaneously.","rationale":"I read the paper as a careful, transparent negative-result benchmark. The central methodological claim is well supported: per-seed values are released, paired tests and effect sizes are reported, the five-seed Wilcoxon floor is acknowledged in §3.2, and the post-hoc power analysis in §4.5 correctly states that the objective gaps are not reliably established at this seed budget rather than demonstrably zero. I therefore do not see a load-bearing attack on the multi-seed null results.\n\nThe one place where the paper reaches past its evidence is the causal attribution of encoder quality to pretraining design rather than parameter count. The reader's weakest assumption identifies this same spot. Table 7 compares four off-the-shelf checkpoints that differ in architecture, pretraining data, tokenizer, and scale simultaneously; no controlled scaling within one architecture is performed. The 'design, not parameter count' wording in the abstract and §5.3 is thus an inference from an uncontrolled comparison. It is an overclaim but is separable from the main benchmark, so I retain the reader's CONDITIONAL verdict; 'UNCHANGED' means no adjustment to the existing conditional is needed.","tokens_in":19924,"tokens_out":6443,"duration_ms":65464,"concrete_test":"Run a scale-controlled contrast on the same VAANI splits and five-seed protocol: fine-tune XLS-R 1B (same architecture and pretraining family as XLS-R 300M) and, if available, a 580M MMS checkpoint under the identical recipe. If XLS-R 1B lowers WER toward w2v-BERT 2.0 and MMS-1B is close to XLS-R 1B, parameter count (or its correlate) explains much of the gap; if XLS-R 1B stays near 50.2 while w2v-BERT 2.0 remains near 47.0, the design/coverage attribution is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central methodological argument—that multi-seed evaluation on official splits is necessary and that Focal CTC, the matra-weighted objective, and Hindi transfer show no reliable gain over standard CTC—is well supported by the per-seed tables, paired Wilcoxon results, and the explicit power analysis. The five-seed test floor is disclosed in §3.2, and the conclusions are worded as statistical rather than absolute, so the null results are not overstated.\n\nThe load-bearing weakness is the third contribution and the abstract's 'pretraining design, not parameter count, drives performance' (also §5.2 and Fig. 2). The evidence is a four-model comparison (w2v-BERT 2.0 580M, MMS-1B, XLS-R 300M, HuBERT Large) under one protocol. These checkpoints differ simultaneously in architecture (w2v-BERT vs. wav2vec2-style vs. HuBERT), pretraining corpus and language coverage, tokenizer or quantization, and parameter count. No axis is varied alone: there is no 580M baseline from the XLS-R/MMS family, no 1B w2v-BERT-style model, and no two checkpoints matched on everything except pretraining design. The observed ordering could be driven by pretraining data, architecture, fine-tuning recipe, or scale; the 'design, not parameter count' attribution is therefore not identified. This is an overclaim in the abstract and §5.3, but it does not threaten the objective/transfer null results.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that at low-resource dialect corpus sizes, single-run ASR comparisons are unreliable, and it supports this by building a five-seed Garhwali benchmark on the official VAANI splits. Under this protocol, the paper reports that Focal CTC, a matra-weighted CTC objective, and Hindi-to-Garhwali transfer do not reliably improve over standard CTC, that speed augmentation gives a small and largely consistent mean gain, and that w2v-BERT 2.0 (580M) reaches the lowest WER among the compared encoders (47.0% over five seeds). The paper also provides per-seed outputs, Holm-corrected paired Wilcoxon tests, a bootstrap cross-check, a post-hoc power analysis, and a per-category error analysis indicating that matra and conjunct errors dominate.","tokens_in":20274,"tokens_out":7010,"duration_ms":68537,"significance":"If accepted with the proposed correction, this is a valuable methodological contribution to low-resource ASR evaluation. The study is exemplary in transparency: per-seed results are reported in the appendix, the five-seed Wilcoxon floor is disclosed in Section 3.2, the power analysis in Section 4.5 is honest about what the seed budget can and cannot establish, and code plus per-seed outputs are released. The central null results for the three objective/transfer interventions are empirically grounded and are worded as statistical rather than absolute. The main weakness is the encoder-comparison claim, which is not identified by the evidence as currently presented.","major_comments":[{"comment":"The claim that \"pretraining design, not parameter count, drives performance\" is not identified by the evidence offered. The four models in Table 7 and Figure 2 — w2v-BERT 2.0 (580M), MMS-1B (1B), XLS-R 300M, and HuBERT Large — differ simultaneously in architecture (w2v-BERT versus wav2vec2-style versus HuBERT), pretraining corpus and language coverage, tokenizer or quantization details, fine-tuning recipe, and parameter count. No axis is varied alone: there is no same-family checkpoint at a different scale, and no two checkpoints are matched on everything except pretraining design. The observed ordering could plausibly be driven by pretraining data, architecture, or scale rather than by pretraining design per se. A concrete test would be to run a same-family scaling comparison (e.g., XLS-R 300M versus XLS-R 2B under the same protocol, or a 1B w2v-BERT-style checkpoint) or to match parameter counts across architectures. Until such a comparison is added, the claim in the abstract and in Section 5.3 should be softened to, for example, \"the best-performing checkpoint in our set is not the largest\" or \"pretraining design is a plausible factor, but scale and design are confounded in our comparison.\" This overclaim does not affect the multi-seed objective and transfer null results, which rest on Tables 2 and 3 and are supported by the per-seed data.","section":"Abstract; §5.2, Table 7, Fig. 2"}],"minor_comments":[{"comment":"The citation \"(K et al., 2025)\" should be expanded to the full author name used in the reference list (\"Thennal D K et al., 2025\") so the reference is unambiguous.","section":"Appendix E"},{"comment":"The text repeatedly uses the nonstandard ligature \"oﬀicial\"; please replace it with \"official\" for consistency with standard notation.","section":"Throughout"},{"comment":"The broken y-axis is not explicitly labeled as broken in the figure; adding a clear break marker or a caption note would prevent misreading the vertical distances near WER 46–48.","section":"Fig. 2"},{"comment":"The caption says percentages are \"of error chunks,\" but the table lists both chunk counts and percentages; clarifying the denominator in the caption or a footnote would help readers avoid interpreting the percentages as row-level rates.","section":"Table 10"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically careful, and the core null results are solid. The main required revision is to correct the encoder-comparison overclaim in the abstract and in Section 5.3, either by adding a controlled scaling comparison or by weakening the claim to match the confounded design. The per-seed data release and the explicit discussion of the Wilcoxon floor and power are exemplary practices that should be highlighted in the final version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this paper if you care about low-resource dialectal ASR evaluation. It does what it advocates: all primary systems run over five seeds on the official VAANI splits, per-seed WER/CER values are in the appendix, paired Wilcoxon with Holm correction is applied, a bootstrap cross-check is included with its limitations spelled out, and the post-hoc power analysis is reported with the five-seed p-value floor disclosed. The null results for Focal CTC, matra-weighted CTC, and Hindi-to-Garhwali transfer are well supported by the per-seed tables; the authors even note that five-seed Wilcoxon cannot reach p<0.05 and rest their conclusions on paired differences, effect sizes, and the t-based power analysis. That is honest and reproducible. The divergence from the prior single-run Garhwali study, where encoder rankings swap between the two papers, makes the methodological point effectively. Speed augmentation is direction-consistent (13 of 15 seeds) and they correctly describe it as small and not statistically certified. The benchmark itself is genuinely useful; I would cite it. The main soft spot is contribution 3 and the abstract sentence 'pretraining design, not parameter count, drives performance.' The evidence is four checkpoints - w2v-BERT 2.0 (580M), MMS-1B, XLS-R 300M, and HuBERT Large - that differ simultaneously in architecture, pretraining corpus and coverage, tokenizer or quantization, and scale. No axis is varied alone. The observed ordering could be driven by pretraining data, architecture, or fine-tuning recipe just as easily as by 'design' versus 'parameter count'; the attribution is not identified. This overclaim does not damage the core negative results about objectives and transfer, which stand independently, but the abstract should be reworded to say that w2v-BERT 2.0 was the best encoder among those tested and that larger scale did not help across these checkpoints. A controlled scaling study within one model family would be needed to support the causal sentence. Other soft spots are minor and mostly self-flagged: the layer-wise probing is single-seed and exploratory; the Whisper reference is unstable with two failed seeds; the LLM-based error categorization is single-seed and only directional. They are labeled as such. The focal and matra hyperparameters were fixed before the multi-seed comparison, so there is no fitting-to-significance concern. Final take: this is a careful, well-scoped empirical paper with a real contribution as a reproducible benchmark and a useful caution about single-run reporting in this regime. The central null results hold up. The pretraining-design claim needs softening. I would send it to serious peer review with the expectation of minor-to-moderate revision; my main request as referee would be to fix the abstract and contribution 3, and to acknowledge the confound or add one controlled comparison.","headline":"Solid multi-seed Garhwali benchmark with honest null results for Focal CTC, matra weighting, and Hindi transfer; the 'pretraining design, not parameter count' claim is overreach from a confounded four-model comparison.","tokens_in":796,"tokens_out":1952,"would_cite":true,"duration_ms":37407,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that single-run comparisons are unreliable in low-resource dialectal ASR and proves it for Garhwali: under a five-seed protocol, Focal CTC, a matra-weighted objective, and Hindi-to-Garhwali transfer all fail to beat…","keywords":["low-resource ASR","dialectal speech recognition","Garhwali","multi-seed evaluation","CTC","focal loss","cross-lingual transfer","reproducibility"],"falsifier":"Run the three interventions and standard CTC under the same official splits but with 20 seeds; if any intervention's Holm-corrected paired test gives a significantly lower mean WER than standard CTC, the paper's central claim that these objectives do not reliably help is falsified.","tokens_in":19733,"feed_emoji":"🎲","tokens_out":9301,"duration_ms":80498,"temperature":0.7,"pith_summary":"At the small corpus sizes typical of low-resource dialects, a single training run cannot tell whether an ASR change helps. The paper shows this for Garhwali, building the first reproducible multi-seed benchmark on the official released splits and testing three plausible interventions: Focal CTC (a loss that down-weights easy examples), a matra-weighted objective that up-weights vowel-diacritic errors, and Hindi-to-Garhwali transfer. All three dissolve into seed noise: once results are averaged over five seeds with paired significance testing, none beats standard CTC, and the matra-weighted loss does not even reduce the vowel-diacritic errors it was built to attack. What survives is ordinary: a 580M-parameter w2v-BERT 2.0 encoder reaches 47.0% word error rate while beating larger multilingual models, and speed augmentation gives a small, largely consistent gain. The broader point a reader can take away is that multi-seed reporting with per-seed outputs and significance tests is the minimum needed to trust claimed gains in low-resource dialectal ASR.","feed_headline":"Five seeds show speech-recognition 'gains' vanish as seed noise","feed_subtitle":"On official Garhwali splits, standard CTC at 47.0% word error rate beats all three proposed improvements.","key_machinery":"The central mechanism is the multi-seed paired evaluation protocol: every primary system is fine-tuned over five fixed seeds (42, 123, 777, 2025, 1234) on the same official train/validation/test splits, with per-seed WER and CER reported, paired Wilcoxon signed-rank tests with Holm-Bonferroni correction, and a post-hoc power analysis that states how many seeds a real effect would require. This is what carries the argument, because the within-system seed spread (1.6 WER points for standard CTC) is comparable to or larger than the gaps between objectives (0.4 to 0.8 points). The design makes visible what a single lucky seed can hide: the best Focal seed and the best standard seed are within noise of one another, while the multi-seed means are not.","core_discovery":"Stated on the paper's own terms: on the official splits, with five seeds per system, standard CTC is the floor that the trained objectives cannot beat. The five-seed mean is 47.0% WER for standard CTC, 47.83% for Focal CTC, 47.42% for the matra-weighted objective, and 47.22% for Hindi-to-Garhwali transfer; the matra objective's targeted error rate is essentially unchanged (22.26% versus 22.31%), and the only seed-consistent gap, standard versus focal, favors the baseline. Across encoders, w2v-BERT 2.0 (580M) beats MMS-1B (49.0%), XLS-R (50.2%), and HuBERT (60.9%), which the paper reads as evidence that pretraining design, not parameter count, drives performance. Speed augmentation lowers WER by 1.1 to 1.5 points under every objective, making it the only dependable intervention tested. The paper concludes that at this corpus size the residual bottleneck is representational, involving fine phonological contrasts such as vowel length and conjunct marking, rather than procedural, and that single-seed numbers in this regime should not be trusted as evidence of an effect.","pith_inferences":["Beyond the paper: if the representational-bottleneck reading is right, the next gains for Garhwali should come from better acoustic representations or additional unlabeled audio, not from further loss engineering; the paper leaves this as future work, but its own evidence points that way.","Beyond the paper: the same multi-seed protocol could be applied to other low-resource dialect benchmarks to check whether their reported single-run gains replicate; the paper argues the protocol transfers, but the specific null results are Garhwali-only.","Beyond the paper: the power analysis implies that a real effect smaller than about one WER point would need roughly 11 to 20 seeds to detect, so any future study claiming a small gain in this regime should budget its seed count accordingly.","Beyond the paper: the silence hallucination finding, where all eight true-silence segments receive spurious text across all five seeds, suggests a cheap and testable mitigation: voice-activity gating before decoding."],"forward_implications":["Future low-resource dialectal ASR results should report multi-seed means, per-seed values, and significance tests before claiming an objective or transfer benefit; single-run gains should be treated as hypotheses rather than results.","At Garhwali-scale data, objective engineering such as focal or class-weighted losses is unlikely to yield reliable gains, so effort should move to pretrained encoder choice, augmentation, and richer acoustic or data evidence.","Model comparisons from single runs should not be trusted: the same encoders swapped relative order between the two studies, so rankings need variance reporting to replicate.","Speed perturbation is a cheap, consistent improvement of roughly one to one and a half WER points across all objectives and should be part of the default training recipe.","Hindi-to-Garhwali transfer is not a dependable lever; direct fine-tuning is as good or better, so transfer studies need the same multi-seed scrutiny before claiming gains."],"supporting_citations":[{"why":"Provides the only prior Garhwali ASR benchmark (49.3% WER on a single-run split) whose numbers the paper re-examines; the two studies' encoder rankings diverge, illustrating split and seed sensitivity.","marker":"Dhasmana et al. (2026)"},{"why":"Supplies the variance-accounting methodology that the paper's multi-seed protocol is built on.","marker":"Bouthillier et al. (2021)"},{"why":"Documents how strongly random seeds can affect deep-learning results, motivating the caution against single-run comparisons.","marker":"Picard (2021)"},{"why":"Quantifies seed effects on fine-tuning and is cited as evidence that seed-induced variance must be reported explicitly.","marker":"Bui et al. (2025)"},{"why":"Defines the CTC loss that is the standard baseline and the base of the two modified objectives.","marker":"Graves et al. (2006)"},{"why":"Introduces focal loss, which the paper adapts into Focal CTC for utterance-level weighting.","marker":"Lin et al. (2017)"},{"why":"Establishes speed perturbation as the augmentation technique that produces the paper's only consistent mean WER gain.","marker":"Ko et al. (2015)"},{"why":"Introduces w2v-BERT 2.0, the primary encoder that reaches 47.0% WER.","marker":"Chung et al. (2021)"},{"why":"Supplies XLS-R, one of the comparison encoders whose lower performance supports the design-not-scale argument.","marker":"Babu et al. (2022)"},{"why":"Supplies MMS-1B, the larger model that the smaller w2v-BERT 2.0 outperforms.","marker":"Pratap et al. (2024)"}],"fun_headline_variants":["Standard CTC beats all proposed fixes in multi-seed Garhwali test","Garhwali ASR: seed-level testing exposes fragile gains","Five-seed test: Focal CTC, matra, transfer fail to beat baseline","Pretraining design beats model size for Garhwali speech recognition","Speed augmentation is the only reliable gain for low-resource ASR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that five seeds on one 8.8-hour corpus capture enough run-to-run noise to judge what is real; the paper's own power analysis shows the seed-level tests cannot reach $p < .05$ at that budget, so the negative results are statements about reliability rather than proof of absence.","fun_headline_variants_meta":{"raw":{"variants":["Standard CTC beats all proposed fixes in multi-seed Garhwali test","Garhwali ASR: seed-level testing exposes fragile gains","Five-seed test: Focal CTC, matra, transfer fail to beat baseline","Pretraining design beats model size for Garhwali speech recognition","Speed augmentation is the only reliable gain for low-resource ASR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000773,"raw_usage":{"total_tokens":3443,"prompt_tokens":988,"completion_tokens":2455,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":2361}},"tokens_in":604,"tokens_out":2455,"duration_ms":17873,"temperature":1.0,"reasoning_tokens":2361,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:42:17.694823+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the three interventions and standard CTC under the same official splits but with 20 seeds; if any intervention's Holm-corrected paired test gives a significantly lower mean WER than standard CTC, the paper's central claim that these objectives do not reliably help is falsified.","supporting_citations":[],"review_version":1}