{"id":"772909bc-8a71-4646-8eb8-5a13fb0b4d80","arxiv_id":"2508.09957","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new Badini Kurdish speech corpus and a comparison of Wav2Vec2 versus Whisper show Wav2Vec2 achieves 82.67% accuracy versus Whisper's 53.17%.","lead":"This paper builds a 15-hour Badini Kurdish speech corpus from children's stories and uses it to fine-tune two speech-to-text models. It reports that Wav2Vec2 outperforms Whisper on this dialect, with higher readability and accuracy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Whisper-small may not have been fine-tuned on the same Badini data as Wav2Vec2, making the reported accuracy gap an artifact of unequal training rather than model superiority.","rationale":"The reader's verdict is UNVERDICTED, and I agree. The abstract alone provides insufficient detail to evaluate the central claim. My specific, load-bearing concern is the fairness of the comparison: without explicit confirmation that Whisper-small was fine-tuned on the same Badini data and under the same conditions as Wav2Vec2, the reported accuracy gap may be an artifact of unequal training. This is precisely one of the two components of the reader's weakest_assumption ('compared under equivalent conditions'), so I mark agreement as 'agree'. The full text is garbled mojibake, making it impossible to check the methodology directly. A controlled experiment with equal fine-tuning budgets and a shared test set would resolve the concern. Since the necessary evidence is absent, the verdict should remain UNVERDICTED; my concern does not shift it to ACCEPT/REJECT because it is not a demonstrated flaw, only a plausible confound that cannot be assessed from the provided text.","tokens_in":3359,"tokens_out":3080,"duration_ms":37253,"concrete_test":"Recover the experimental section (or contact the authors) and determine whether Whisper-small underwent fine-tuning on the same Badini training segments as Wav2Vec2, with the same data split, hyperparameters, and number of epochs. If the paper omits this, run a controlled comparison: fine-tune both models on the same 80/10/10 split of the 15-hour corpus for equal epochs, decode the identical held-out test set, and compute word error rate plus the paper's 'accuracy' and 'readability' metrics with 95% confidence intervals. If the accuracy gap persists (e.g., >20 percentage points), the central claim survives; if it narrows substantially or reverses, the original comparison was confounded by unequal training conditions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a comparative superiority statement: Wav2Vec2-Large-XLSR-53 'performs better' than Whisper-small on Badini Kurdish STT, with 82.67% vs 53.17% accuracy. Such a claim is only valid if the two models were trained and evaluated under identical conditions: same training/validation/test split, same fine-tuning data, same optimization budget, and identical metric definitions. The abstract reports only that both models were 'used to develop the language models' and gives the accuracy/readability numbers; it does not state whether Whisper-small was fine-tuned on the 15-hour Badini corpus or simply evaluated in zero-shot mode. If Whisper-small was not fine-tuned (or was fine-tuned with a different recipe), the comparison is confounded: Wav2Vec2-Large-XLSR-53 is a cross-lingual model that benefits strongly from target-domain fine-tuning, while Whisper-small, a general-purpose model with limited exposure to low-resource dialects, would be expected to do poorly zero-shot. The reported gap could then reflect training-effort disparity, not an inherent performance advantage. Because Badini is a low-resource dialect and the paper's purpose is to build a language model, the experimental setup must be known. The garbled full text prevents verifying this, so the central claim remains unsubstantiated until the training protocol is disclosed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a new Badini Kurdish speech corpus constructed from 78 children's stories from eight books, read by six narrators, with approximately 17 hours of audio reduced to nearly 15 hours after cleaning and segmentation (19,193 segments, 25,221 words). Using this corpus, the authors compare Wav2Vec2-Large-XLSR-53 and Whisper-small for speech-to-text. The abstract reports that Wav2Vec2 outperforms Whisper-small on readability (90.38% vs 65.45%) and accuracy (82.67% vs 53.17%), concluding that Wav2Vec2 performs better for Badini Kurdish STT. The paper's central claim is empirical, but the submitted text provides almost no experimental detail, and the body text is largely unreadable due to encoding corruption.","tokens_in":3658,"tokens_out":3801,"duration_ms":44917,"significance":"If the corpus is released and the comparison is properly controlled, this would be a valuable contribution to low-resource Kurdish dialect technology and a useful apples-to-apples comparison of two widely used pretrained models. The resource-creation effort (15 hours, 78 stories, six narrators) is a concrete strength. However, the significance cannot be assessed from the submitted version: the final performance numbers are not backed by any experimental protocol, and the manuscript body is garbled. The claim of Wav2Vec2 superiority depends entirely on a controlled comparison that the paper does not document.","major_comments":[{"comment":"The central comparative claim rests on two numbers, but the paper reports no training or evaluation protocol. Specify the data split (train/validation/test), fine-tuning hyperparameters (learning rate, batch size, epochs, optimizer, warmup), the exact fine-tuning data for each model, and whether both models were trained on the same Badini segments. If Whisper-small was evaluated zero-shot or with a different fine-tuning budget, the reported gap is not evidence of model superiority. This is load-bearing because the paper's conclusion is exactly that Wav2Vec2 performs better.","section":"Abstract / Experiments"},{"comment":"The term 'readability' is not defined. The abstract reports 90.38% and 65.45% readability and 82.67% and 53.17% accuracy, but no formula, annotation protocol, or scoring procedure is given. Clearly define both metrics (e.g., word/character error rate complement, token-level accuracy, human-rated readability) and state how each was computed.","section":"Metric definitions"},{"comment":"No error bars, confidence intervals, or significance tests are reported. With a single corpus and a single split, the reader cannot judge whether the 82.67% vs 53.17% gap is stable or an artifact of one split. Provide repeated runs, bootstrap intervals, or at least the test-set size and a statistical significance test.","section":"Statistical reliability"},{"comment":"The body of the manuscript as provided is visually corrupted (mojibake/encoding errors), so the methodology, table, and results sections are not readable. A clean version is required for review. This is not a scientific flaw in the work itself, but it prevents verification of any claim beyond the abstract.","section":"Full text / readability of submission"}],"minor_comments":[{"comment":"'Bandin' appears to be a typo for 'Badini'; also 'SST' and 'STT' are used inconsistently.","section":"Abstract"},{"comment":"Numbers such as '19193 segments' and '25221 words' should be formatted with thousands separators for readability.","section":"Data statistics"},{"comment":"Ensure full citations for Wav2Vec2-Large-XLSR-53 and Whisper-small are present; the reference list is not visible in the submitted version due to the encoding corruption.","section":"References"},{"comment":"The corpus consists entirely of read children's stories; state explicitly that the reported performance may not generalize to conversational or spontaneous speech, which the introduction seems to motivate.","section":"Scope / generalization"}],"recommendation":"major_revision","confidential_remarks":"The corpus is potentially valuable, but the submitted version is not reviewable because the main text is garbled. Please ensure a clean PDF is provided for any revision; the core scientific request should be for a transparent, controlled comparison of the two models."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real asset is the corpus: 15 hours of Badini Kurdish read speech from six narrators, aligned to story transcripts. That is a genuinely useful resource for a dialect with essentially no STT support, and the authors are clear that they are filling a gap. The comparison of Wav2Vec2-Large-XLSR-53 and Whisper-small is a reasonable practical question. But the abstract is all we have, and it lacks the details that would make the comparison interpretable: no train/test split, no hyperparameters, no fine-tuning protocol, no definition of 'readability,' no error bars. More concerning, the 82.67% vs 53.17% gap could simply mean Whisper-small was evaluated differently—maybe zero-shot, maybe with less fine-tuning. That is not a minor caveat; the central claim is a superiority claim that only holds under matched training conditions. The stress-test note is right to flag it. I cannot confirm it from the garbled full text, but the burden is on the authors to show the two models were treated fairly. In a field where pretrained model comparisons are routine, this matters. The corpus itself is also limited to read children's stories, so generalizing to conversational Badini is an assumption. The authors say they chose stories for grammatical confidence and ready transcriptions, which is defensible, but the paper should say it outright. I would send this to review, not desk-reject it: the dataset is worth refereeing, and a patient reviewer can ask for the missing controls and a release of the corpus. If the code and data ship, this becomes a citable benchmark for a low-resource dialect. Recommended: accept with major revision after honest experimental details and a matched fine-tuning protocol.","headline":"Useful new Badini Kurdish corpus; the model comparison is under-specified and possibly confounded.","tokens_in":4114,"tokens_out":1622,"would_cite":false,"duration_ms":18552,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Wav2Vec2-Large-XLSR-53, fine-tuned on roughly 15 hours of Badini Kurdish read speech, transcribes the dialect with 82.67% accuracy and 90.38% readability, clearly outperforming Whisper-small at 53.17% and 65.45%.","keywords":["Badini Kurdish","speech-to-text","Wav2Vec2","Whisper","low-resource language","corpus construction","dialect ASR","fine-tuning"],"falsifier":"Take the two fine-tuned models and evaluate them on a new held-out set of spontaneous Badini conversation recorded from adult speakers, none of whose voices appear in the 15-hour corpus. If Whisper-small matches or exceeds Wav2Vec2 on accuracy or readability there, the paper's conclusion that Wav2Vec2 'performs better' is reversed. A simpler check: release the exact train/validation/test splits and re-run the evaluation with test speakers completely excluded; the reported 82.67%-vs-53.17% gap should persist.","tokens_in":3267,"feed_emoji":"🎙️","tokens_out":4198,"duration_ms":39567,"temperature":0.7,"pith_summary":"This paper tries to establish that Wav2Vec2-Large-XLSR-53, a self-supervised speech model, outperforms Whisper-small when both are fine-tuned for Badini Kurdish speech-to-text, a dialect that previously had no working STT system. To make the test possible, the authors built a new corpus from 78 children's stories in eight books, narrated by six speakers, yielding nearly 15 hours of audio split into 19,193 segments. On this corpus, Wav2Vec2 reaches 82.67% accuracy and 90.38% readability, while Whisper-small reaches 53.17% and 65.45%. If these numbers hold, the paper provides a first usable baseline for Badini STT and identifies which pretrained model is the better starting point for a low-resource dialect.","feed_headline":"Wav2Vec2 beats Whisper on Badini Kurdish speech-to-text","feed_subtitle":"Fine-tuned on 15 hours of Badini children's stories, Wav2Vec2 hits 83% accuracy vs Whisper's 53%.","key_machinery":"The load-bearing machinery is the pair of pretrained models and the new corpus. Wav2Vec2-Large-XLSR-53 is a self-supervised speech encoder pre-trained on 53 languages, producing representations that are then fine-tuned on the Badini audio; Whisper-small is a transformer-based speech-recognition model trained on large-scale multilingual data with weak supervision. The corpus — 78 Badini children's stories, spoken by six narrators, cleaned and segmented into 19,193 segments containing 25,221 words — is what makes the comparison possible, because it gives both models the same target domain and evaluation data. The comparison itself is the argument: same data, same task, two models, and a report","core_discovery":"The central claim is that Wav2Vec2-Large-XLSR-53 is the better model for Badini Kurdish speech recognition: fine-tuned on roughly 15 hours of read children's stories, it produces transcriptions that are both more accurate and more readable than Whisper-small under the same conditions. Specifically, the paper reports 82.67% accuracy and 90.38% readability for Wav2Vec2, versus 53.17% accuracy and 65.45% readability for Whisper-small. The authors interpret the gap as evidence that self-supervised speech representations can adapt to an under-resourced dialect more effectively than a large multitask model, and they frame the corpus itself as a contribution that closes a gap for Badini speakers, w","pith_inferences":["The comparison probably favors Wav2Vec2 more than a deployment would: read children's stories have clear articulation and limited vocabulary, so the gap on spontaneous adult conversation could be smaller or even reversed, especially for Whisper, which was trained on diverse audio.","The paper does not report whether the six narrators overlap between training and test sets; if the test set includes narrators already heard in training, the reported accuracy may overstate speaker-independent performance.","Readability is the more striking difference, but without a precise definition of the readability metric it is hard to know whether it measures human judgment, word-order preservation, or some automatic proxy."],"forward_implications":["If the ranking holds, Badini speakers get a functional speech-to-text baseline, opening the way for dictation, subtitles, and voice-controlled applications in the dialect.","The result suggests that for very low-resource dialects, a self-supervised speech encoder fine-tuned on a few dozen hours may beat a much larger multitask model trained for general speech recognition.","The corpus construction recipe — written children's stories, multiple narrators, segmentation into short clips — can be reused for other under-resourced dialects such as Hawrami, which the paper names as still lacking STT.","The reported readability gap (90.38% vs 65.45%) is larger than the accuracy gap, implying Whisper's errors are not just more frequent but also harder for a human to understand, which affects real-world usability.","Future work can test whether larger Whisper variants or newer models close the gap, and whether conversational Badini speech behaves like the read-story corpus."],"supporting_citations":[],"fun_headline_variants":["Wav2Vec2 outperforms Whisper for Badini Kurdish STT","On Badini Kurdish, Wav2Vec2 beats Whisper in STT accuracy","Wav2Vec2 tops Whisper for Badini speech recognition","Fine-tuned Wav2Vec2 beats Whisper on Badini Kurdish"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that roughly 15 hours of read children's stories from six narrators is representative enough of Badini speech to generalize the model ranking, and that the two models were fine-tuned and evaluated under genuinely equivalent conditions — including the same train/test split, hyperparameters, and tuning effort.","fun_headline_variants_meta":{"raw":{"variants":["Wav2Vec2 outperforms Whisper for Badini Kurdish STT","On Badini Kurdish, Wav2Vec2 beats Whisper in STT accuracy","Wav2Vec2 tops Whisper for Badini speech recognition","Fine-tuned Wav2Vec2 beats Whisper on Badini Kurdish"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000406,"raw_usage":{"total_tokens":2007,"prompt_tokens":863,"completion_tokens":1144,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1061}},"tokens_in":607,"tokens_out":1144,"duration_ms":8857,"temperature":1.0,"reasoning_tokens":1061,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:39:55.535827+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the two fine-tuned models and evaluate them on a new held-out set of spontaneous Badini conversation recorded from adult speakers, none of whose voices appear in the 15-hour corpus. If Whisper-small matches or exceeds Wav2Vec2 on accuracy or readability there, the paper's conclusion that Wav2Vec2 'performs better' is reversed. A simpler check: release the exact train/validation/test splits and re-run the evaluation with test speakers completely excluded; the reported 82.67%-vs-53.17% gap should persist.","supporting_citations":[],"review_version":1}