{"id":"fe384030-510d-4d15-880d-0dc97b034892","arxiv_id":"2411.14493","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of Urdu ASR that catalogs datasets, toolkits, and techniques from HMM-GMM to pre-trained multilingual models, concluding that data scarcity and limited transfer learning hold the field back.","lead":"This paper reviews research on automatic speech recognition for Urdu, a language spoken by more than 100 million people but with very little transcribed speech data. It catalogs Urdu speech datasets, tools, and methods from HMM-GMM to fine-tuned transformers, and argues that data scarcity and underused transfer learning are the main barriers to progress.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central comparison tables contradict the text and each other on key Urdu ASR numbers (Whisper WER 17.5 vs 22.6/24.2; FLEURS hours 1.4 vs 8.5 vs 12; XLSR '0.49' with no metric), so the 'comprehensive' claim is not yet supportable.","rationale":"We read the paper in good faith as a survey whose value depends on the accuracy and comparability of its compiled results. The reader's weakest assumption concerned the reproducibility and completeness of the Section 3 literature search; that is a real weakness, since no databases, dates, screening counts, or inclusion criteria are given. However, a more immediate, text-verifiable failure is that the paper's central comparison tables contradict each other and the prose. The purpose of a survey is to let a reader trust the summarized numbers; when the same system's WER is 17.5 in one section and 22.6/24.2 in tables, or FLEURS Urdu hours are 1.4, 8.5, and 12 in different places, the 'comprehensive' claim fails at the data level. This is internal inconsistency, not a disagreement with external consensus, and it is directly visible in the manuscript. We also give credit where it is due: the qualitative narrative correctly sketches the arc from HMM-GMM baselines to DNNs and fine-tuned multilingual pre-trained models, identifies private small datasets as the binding constraint, and assembles a broad reference list. The duplicated paragraphs and duplicate citation keys are editorial flaws that do not by themselves sink the central claim. The concrete test above would settle whether the contradictory table entries are typographical errors that can be corrected or evidence that the compilation is unreliable. Because the paper is likely salvageable by reconciling these numbers and documenting the search, the existing CONDITIONAL verdict remains appropriate rather than a hard reject.","tokens_in":23539,"tokens_out":4465,"duration_ms":44964,"concrete_test":"Trace the contradictory entries to their primary sources and report the exact Urdu-specific values. Specifically: (1) From Radford et al. (2023) or the Whisper model card, determine the Urdu WER on the benchmark referenced in Sec. 4.4, and verify whether 17.5%, 22.6%, or 24.2% is correct for the cited evaluation conditions (FLEURS vs Common Voice 9). (2) From Conneau et al. (2023), resolve whether FLEURS contains 1.4, 8.5, or 12 hours of Urdu read speech, and correct Table 1, Sec. 4.1.3, and Table 5 accordingly. (3) From Mohiuddin et al. (2023), identify the metric for the '0.49' XLSR result in Table 2 and report it explicitly. If any value cannot be traced, mark that row as unresolved and remove the unsupported claim of comprehensiveness until corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is the first comprehensive systematic review of Urdu ASR (Secs. 1.2, 7.2, 8), whose deliverable is a trustworthy comparison of datasets, tools, and WERs. That deliverable is undermined by direct internal contradictions. Whisper's Urdu WER is reported as 17.5% in Sec. 4.4 but as 22.6% on FLEURS and 24.2% on Common Voice 9 in Table 4, and as 22.6 in Table 2. FLEURS Urdu audio is described as '1.4 hours' in Sec. 4.1.3, '8.5 hours' in Table 1, and '12-hour dataset per language' in Table 5. The XLSR-Wav2Vec2 result in Table 2 is '0.49' with no metric given. These are not cosmetic slips: a survey built to give researchers comparable figures cannot do so when the headline numbers disagree within the same paper. A reader consulting Sec. 4.4 would cite a different Whisper Urdu WER than one consulting Table 4, and there is no indication of which source or evaluation condition is correct. Until every contradictory entry is traced to its primary source and reconciled, the 'comprehensive systematic review' claim fails on data quality grounds, independently of the under-documented search protocol in Sec. 3. The editorial issues (duplicated paragraphs in Sec. 2.2 and Sec. 6, duplicate citation keys) compound this but are secondary.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys automatic speech recognition (ASR) for Urdu, covering the transition from HMM-GMM systems to deep neural networks and to multilingual pre-trained models such as Whisper and XLSR. The manuscript is organized into sections on motivation, background, methodology, Urdu datasets, tools, monolingual and multilingual ASR approaches, comparison tables, discussion, open challenges, and conclusion. Its central claim, repeated in Sections 1.2, 7.2, and 8, is that it is the first comprehensive systematic review of Urdu ASR. The qualitative storyline is plausible and consistent with the cited literature in outline, but the quantitative comparison tables contain unresolved internal contradictions, and Section 3 does not document a reproducible literature search.","tokens_in":23676,"tokens_out":6853,"duration_ms":63586,"significance":"If the internal inconsistencies are corrected and the search protocol is documented, the survey would be a useful starting point for researchers entering Urdu ASR: it collects scattered dataset descriptions, toolkit choices, and accuracy figures into one document, and it identifies open problems such as lack of standardized orthography, limited transfer learning, and missing end-to-end systems. The claimed novelty as the first Urdu-specific ASR survey is plausible but not yet evidenced, because the authors do not demonstrate that the literature collection was exhaustive or that no prior survey covers the same ground. I therefore view the contribution as defensible in outline but currently blocked by data-quality problems in the central tables.","major_comments":[{"comment":"The Whisper Urdu WER is internally inconsistent: Section 4.4 reports 17.5% WER, Table 2 row 15 reports 22.6 for 'Whisper Large-V2', and Table 4 row 4 reports 22.6 on FLEURS and 24.2 on Common Voice 9. Because Whisper is the strongest pre-trained model discussed and this number is the principal evidence for its suitability, the authors must trace each value to its primary source, state the evaluation condition (dataset split, use of language model, fine-tuning), and reconcile the text with the tables.","section":"Sec. 4.4; Table 2; Table 4"},{"comment":"The FLEURS Urdu audio size is reported in three incompatible ways: 1.4 hours in Section 4.1.3, 8.5 hours in Table 1 row 23, and 12 hours per language in Table 5. This affects how readers judge the resource status of Urdu within FLEURS, so it is not a formatting slip. The authors should consult the FLEURS source and give a single figure with explicit split information (e.g., train/validation/test) rather than three different totals.","section":"Sec. 4.1.3; Table 1; Table 5"},{"comment":"Table 2 row 14 reports the XLSR-Wav2Vec2 result as '0.49' without specifying whether this is WER, CER, accuracy, or some other metric, and Section 4.5 describes the same study without providing any quantitative result. A unitless value cannot be compared with the other rows in the table; the authors must state the metric and verify the Urdu-specific value against the primary source.","section":"Table 2; Sec. 4.5"},{"comment":"The methodology does not currently support the claim of a comprehensive systematic review. Section 3 states that the keyword combination 'Urdu Speech Recognition' 'efficiently helped to get most of the Urdu ASR studies across years', but it names no bibliographic databases, no date range, no screening counts, and no inclusion or exclusion criteria, and Figure 2 is a generic four-box diagram. I ask for a reproducible description of the search and screening process, and for an explicit demonstration that no earlier Urdu-specific ASR survey exists (e.g., by discussing how the authors positioned the work relative to Besacier et al. 2014, which already surveys ASR for under-resourced languages including Urdu).","section":"Sec. 3"},{"comment":"The comparison tables mix accuracy percentages with WER values without a consistent metric column, which makes higher-is-better versus lower-is-better ambiguous. For example, Table 3 lists 'SVM, CNN 0.97' next to 'HMM 74%' in the same column, and Table 4 row 5 lists FLEURS CER values of 82.9/83.1 that Table 5 later calls '%CER reductions'. The authors should standardize the tables by adding separate columns for metric and direction, and verifying each entry against the original publication.","section":"Table 3; Table 4"}],"minor_comments":[{"comment":"The sentence beginning 'Traditionally researchers of ASR have explored several techniques, including Gaussian Mixture Models - Hidden Markov Models' is duplicated verbatim at the start of two consecutive paragraphs; one copy should be removed.","section":"Sec. 2.2"},{"comment":"The paragraph beginning 'The insights derived and summarized in this review are poised to captivate...' appears twice, almost verbatim, in the discussion; remove the duplicate.","section":"Sec. 6"},{"comment":"Subsection numbering is duplicated: '4.1.1' and '4.1.3' each appear twice. The subsections for Common Voice and FLEURS should be renumbered so the dataset discussion has a clean sequence.","section":"Sec. 4.1"},{"comment":"The S.No column skips entries 11-14, jumping from 10 to 15; renumber the rows sequentially.","section":"Table 1"},{"comment":"Several references are duplicated with identical author-year labels, including two Ashraf (2010) entries, two Bhogale (2023) entries, and two Khan (2021) entries; disambiguate these with letter suffixes and update the in-text citations accordingly.","section":"References"},{"comment":"The typeset WER formula appears garbled in the manuscript text; please check that it reads (S+D+I)/N × 100.","section":"Eq. (1)"},{"comment":"Whisper is described as a 'massive dataset' in the text; Whisper is a model trained on a large weakly supervised dataset, so the phrasing should be corrected.","section":"Sec. 4.1.3"},{"comment":"There are small terminology slips: 'Discrete Cosine Transform DTS' should likely be DCT, and 'Short Fourier Transforms SFT' is usually called the short-time Fourier transform (STFT).","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"Apart from the inconsistencies described to the authors, I would ask the editor to treat the 'first comprehensive' claim as a publication requirement rather than a rhetorical phrase: the authors need to verify, in writing, that no prior Urdu-specific ASR survey exists and that Section 3's search covers the relevant venues. The paper's fit with a computing/language-resources venue is fine; the main question is whether the authors will do the verification work during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a survey, not an experimental paper. The qualitative storyline is sound: Urdu ASR moved from HMM-GMM baselines to DNNs and then to fine-tuned multilingual models, with data scarcity and mostly private corpora as the binding constraints. If you are new to Urdu ASR, the paper gives you a single place to start: it lists the datasets, toolkits, and reported results scattered across a literature that has no other dedicated survey. That is real value, and the reference list is a good entry point.\n\nThe trouble is the central deliverable. The comparison tables are supposed to let a reader trust the numbers, and they don't. Whisper's Urdu WER appears as 17.5% in Section 4.4 and as 22.6% or 24.2% in Tables 2 and 4. FLEURS Urdu is 1.4 hours in Section 4.1.3, 8.5 hours in Table 1, and 12 hours in Table 5. XLSR-Wav2Vec2 gets a bare '0.49' with no metric. These are not cosmetic slips; a survey built to give comparable figures cannot do so when the headline numbers contradict each other. The WER/PER conflation ('WER is also called phoneme error rate') is a factual error, and the duplicated paragraphs and duplicate citation keys suggest the manuscript was not carefully checked before posting.\n\nThe methodology section is also thinner than the 'systematic review' label requires. Section 3 names a single keyword combination, no databases, no date range, no screening counts, no inclusion criteria. The authors themselves later call the work 'this tutorial' and disclaim 'exhaustive knowledge' in Section 7.1, which sits awkwardly beside the 'first comprehensive systematic review' claims in Sections 1.2 and 7.2. The first-ever claim may be true, but the paper doesn't demonstrate it.\n\nThat said, the flaws are fixable. The narrative holds up, the topic matters, and the compilation is useful enough that a careful revision would be worth having. Whoever takes it on needs to verify every number against the primary source, reconcile the contradictions, document the search protocol honestly, and soften the 'systematic' language. As it stands, I would not cite its numbers.\n\nMy recommendation: send it to peer review, but only with the expectation of a major revision. The reviewer's job should be to force the reconciliation, not to reject the project.","headline":"Useful first map of Urdu ASR, but the headline numbers disagree inside the paper, so the 'comprehensive systematic' claim is not yet supportable.","tokens_in":24401,"tokens_out":1438,"would_cite":false,"duration_ms":16752,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The survey aims to be the first comprehensive review of Urdu automatic speech recognition, mapping datasets, toolkits, and reported word error rates across monolingual and multilingual systems.","keywords":["Urdu ASR","automatic speech recognition","resource-scarce language","speech datasets","multilingual ASR","transfer learning","pre-trained models","systematic review"],"falsifier":"A bibliographic check for an earlier comprehensive Urdu ASR survey published before 2024, or the discovery of a large public Urdu corpus beyond 100 hours that is absent from the paper's dataset tables, would falsify the claim to be the first comprehensive review.","tokens_in":1688,"feed_emoji":"🎙️","tokens_out":2503,"duration_ms":75222,"temperature":0.7,"pith_summary":"This paper sets out to establish that Urdu automatic speech recognition has a recordable research landscape, and that this survey is the first comprehensive systematic review of it. It claims that Urdu, spoken by over 100 million people but short on annotated speech data, has been studied mainly through small datasets, statistical HMM-GMM models, and more recently pre-trained multilingual models. The paper collects the available Urdu speech corpora, the toolkits used to build systems, and the word-error rates reported, and organizes them into monolingual and multilingual categories. A sympathetic reader would value this as a single entry point into a scattered literature: it says where the data is, which tools work, and where the field has stalled.","feed_headline":"Survey maps two decades of Urdu speech recognition","feed_subtitle":"One review collects the datasets, toolkits, and word-error rates that define Urdu ASR, from HMM digits to fine-tuned Whisper.","key_machinery":"The organizing device is a two-way taxonomy: every Urdu ASR study is classified by data type (isolated digits, read speech, spontaneous speech, broadcast, telephone) and by algorithm family (statistical HMM-GMM/SGMM, DNN variants, end-to-end, multilingual pre-trained). The accompanying tables align each study with its toolkit (Sphinx, Kaldi, SRILM, KenLM, ESPnet, Wav2Vec2/XLSR, Whisper) and its reported accuracy, and the discussion reads those tables as evidence that data scarcity, not algorithm choice, is the main constraint on Urdu ASR.","core_discovery":"The paper's central claim is that no earlier work systematically reviewed Urdu ASR as a whole, and that this survey fills that gap by summarizing the datasets, tools, algorithms, and results reported across two decades. It divides monolingual Urdu ASR into traditional statistical approaches (HMM-GMM, SGMM, SVM), deep neural network approaches (TDNN, LSTM, BLSTM, CNN, DNN-HMM), and end-to-end approaches, and treats multilingual systems built on Common Voice, Shrutilipi, IndicSUPERB, Vistaar, FLEURS, Whisper, and XLSR as the route for transfer learning. The best Urdu-specific result it documents comes from a self-supervised Wav2Vec2 model with KenLM at 12.6% word error rate, while Whisper and XLSR fine-tuned on Urdu reach 17.5% and comparable accuracies. On the paper's own account, end-to-end techniques remain mostly untried for Urdu because the available datasets are too small, and the most promising direction is fine-tuning large multilingual pre-trained models.","pith_inferences":["If the survey's coverage holds, the most urgent next step is not another model but a standardized public evaluation set: the WER figures across tables were produced on different vocabularies and test splits, so they are indicative rather than directly comparable.","The same comparison suggests a concrete testable prediction: a Whisper or XLSR model fine-tuned on a combined 300+ hour Urdu corpus with a KenLM language model should beat the reported 12.6% WER monolingual baseline.","The paper's own remarks about dialect and orthography variation imply that a single national Urdu benchmark will likely mislead; a dialect-stratified evaluation (Pakistani, Indian, and diaspora Urdu) would be a more informative yardstick."],"forward_implications":["A new researcher can use the survey's dataset and toolkit tables to reproduce the field's baselines without repeating the search.","The comparison shows that self-supervised multilingual models currently give Urdu's lowest word error rates, suggesting transfer learning rather than new statistical modeling is the productive path.","Because most large Urdu corpora are private, public data like Common Voice and FLEURS will remain the practical foundation for reproducible Urdu ASR unless new open corpora appear.","End-to-end ASR for Urdu stays unexplored, so the survey implies that building a sufficiently large open Urdu corpus would unlock a class of methods not yet evaluated on this language."],"supporting_citations":[{"why":"Surveys ASR for under-resourced languages generally, the baseline the paper extends toward the specific case of Urdu.","marker":"Besacier, 2014"},{"why":"Provides the AUDD Urdu audio digits dataset and CNN baselines used as a reference point for small-data Urdu ASR.","marker":"Chandio, 2021"},{"why":"Reports the 300-hour large-vocabulary Urdu corpus and the TDNN/BLSTM systems that set strong DNN baselines.","marker":"Farooq, 2019"},{"why":"Supplies the Shrutilipi and Vistaar multilingual Indic datasets that include Urdu hours used in the survey's multilingual comparison.","marker":"Bhogale K. S., 2023"},{"why":"Introduces the 1684-hour IndicSUPERB dataset with 86.7 clean Urdu hours and reports the 12.6% Wav2Vec2 WER the survey highlights as best.","marker":"Javed, 2022"},{"why":"Releases Whisper, the pre-trained multilingual model whose Urdu results the survey uses to show transfer-learning potential.","marker":"Radford A, 2023"},{"why":"Fine-tunes XLSR for Urdu on Common Voice, the study the survey treats as a novel transfer-learning demonstration for Urdu.","marker":"Mohiuddin, 2023"}],"fun_headline_variants":["Urdu ASR survey spans two decades of methods","From HMM to Whisper: a review of Urdu speech recognition","First systematic survey of Urdu speech recognition","Wav2Vec2 leads Urdu ASR with 12.6% word error rate"],"cache_read_input_tokens":26240,"weakest_assumption_plain":"The load-bearing premise is that the paper's literature search actually found all or nearly all relevant Urdu ASR work: Section 3 names no databases, date range, screening counts, or inclusion criteria, so the 'first comprehensive systematic review' claim rests on the keyword combination 'Urdu Speech Recognition' being sufficient.","fun_headline_variants_meta":{"raw":{"variants":["Urdu ASR survey spans two decades of methods","From HMM to Whisper: a review of Urdu speech recognition","First systematic survey of Urdu speech recognition","Wav2Vec2 leads Urdu ASR with 12.6% word error rate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000605,"raw_usage":{"total_tokens":2800,"prompt_tokens":903,"completion_tokens":1897,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":1826}},"tokens_in":519,"tokens_out":1897,"duration_ms":13129,"temperature":1.0,"reasoning_tokens":1826,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:21:35.310630+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A bibliographic check for an earlier comprehensive Urdu ASR survey published before 2024, or the discovery of a large public Urdu corpus beyond 100 hours that is absent from the paper's dataset tables, would falsify the claim to be the first comprehensive review.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports the 300-hour large-vocabulary Urdu corpus and the TDNN/BLSTM systems that set strong DNN baselines."}],"review_version":1}