{"id":"96179271-768b-4851-87dc-00454603cea0","arxiv_id":"2505.04531","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 54 studies finds that generative language modelling for low-resource languages relies mostly on transformer models, covers only a small set of languages, and lacks consistent evaluation.","lead":"This paper reviews 54 studies on techniques for building generative language models for low-resource languages, such as data augmentation, back-translation, and prompt engineering. It maps which languages, model architectures, and evaluation metrics have been used, and finds that the field is uneven and hard to compare across studies.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Frequency counts mix primary studies with survey papers, so the method-level map may be an artifact of how the corpus was counted.","rationale":"The review is well-structured, PRISMA-guided, and the qualitative conclusions (transformer dominance, language concentration, evaluation inconsistency) are credible. I considered whether the 'first systematic review' claim is an overclaim; that is hard to settle from the text alone and is not the weakest link. The weakest link is the quantitative synthesis: the corpus intentionally includes survey papers under IC3, yet RQ2/RQ3 frequencies treat them as first-order evidence. This can inflate or double-count methods that are only 'discussed.' The same issue also weakens the representativeness of the sample in a way that is testable from the paper's own tables. The full search strings are missing, but the survey-mixing issue is internal and can be checked without new search infrastructure. Since the needed fix is re-tabulation and fuller reporting, the CONDITIONAL verdict stands rather than a rejection.","tokens_in":41509,"tokens_out":11408,"duration_ms":102289,"concrete_test":"Re-tabulate Figure 7 and Table 3 using only empirical primary studies, excluding review/survey papers such as [29, 86, 108, 109, 123] and any record with QA2=No, and recompute the percentage of records per RQ2/RQ3 category. If the rank order or the top-category percentages shift materially (e.g., back-translation or multilingual modelling becomes the most frequent), the claim that monolingual augmentation is the dominant strategy is an artifact of counting reviews. Additionally, request the complete per-database search strings and search dates; run those strings and quantify how many additional relevant records are returned. If the new records change any headline frequency by more than a few percentage points, the 'first systematic review' synthesis needs revision before acceptance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's central quantitative map depends on treating all 54 included records as comparable units, but the corpus mixes primary studies with at least five survey/review papers ([29, 86, 108, 109, 123]) explicitly included under IC3. In RQ2/RQ3 the authors say papers were categorized 'based on the techniques it discussed or implemented'; as a result, surveys that merely mention a method are counted in the same frequency bars as studies that actually used it. For example, Table 3 lists surveys [29, 86, 109, 123] under 'Augment Monolingual Data' and [86, 108] under 'Back-Translation', so the claim that monolingual augmentation (26%) is the most common strategy may reflect review articles describing the literature rather than empirical deployments. This unit-of-analysis problem is compounded by Section 2.2/Table 1, which reports only four search components rather than the full per-database Boolean strings and no search dates, and Section 5 acknowledges single-reviewer title/abstract screening. If the same underlying primary studies are counted once through survey papers and again directly, or if the unstated search syntax missed relevant records, the reported percentages and research-gap conclusions are not robust.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a PRISMA-style systematic review of 54 studies on strategies to overcome data scarcity in generative language modelling for low-resource languages (LRLs). The review identifies and categorizes technical approaches (monolingual augmentation, back-translation, multilingual modelling, prompt engineering, etc.), analyzes language-family representation, model architectures, and evaluation practices, and reports research gaps and recommendations. The central claims are that transformer-based models dominate, that a small subset of LRLs accounts for most research, and that evaluation is inconsistent across studies.","tokens_in":41691,"tokens_out":6489,"duration_ms":56013,"significance":"If the quantitative map is robust, the review is a useful contribution: it consolidates scattered evidence, provides a structured taxonomy of data-scarcity methods, and gives concrete recommendations for future research, including language-family-centric modelling and more nuanced definitions of 'low-resource'. The paper is transparent about many limitations, ships detailed extraction tables (Tables 2-6), and uses a PICO framework and quality assessment, which are strengths. The falsifiable field-level claims, such as the dominance of transformers and the concentration of studies on a few languages, are important for researchers and funding bodies, and the discussion of geopolitical factors in language selection is a valuable framing.","major_comments":[{"comment":"The central frequency map for RQ2/RQ3 is built on a unit-of-analysis problem: the corpus mixes primary empirical studies with survey/review papers, yet each paper is assigned to technique categories 'based on the techniques it discussed or implemented'. Table 3 explicitly lists surveys [29, 86, 109, 123] under 'Augment Monolingual Data' and [86, 108] under 'Back-Translation', even though these papers review the literature rather than deploy the methods. Consequently, the reported 26% for monolingual augmentation and 24% for back-translation (Figure 7) conflate description with implementation. Because these percentages underpin the review's main empirical conclusions, the authors should re-run the analysis on primary studies only (e.g., QA2=Yes) and report both raw and sensitivity-adjusted frequencies, or clearly separate 'discussed' from 'implemented' categories in the figures and tables.","section":"Section 3.3, Table 3; Section 2.4; Section 5"},{"comment":"The search strategy is not fully reproducible. The text states that 'Table 1 details the search strings that were used for the various digital repositories', but Table 1 shows only four string components ('generative', 'language model OR text model', 'low resource language OR minority language OR endangered language', 'exclusion: classification') with no per-database Boolean syntax, no field restrictions, and no search dates. The PRISMA flow diagram (Figure 1) reports 642 initial records, but without the exact queries and access dates, readers cannot verify the coverage or repeat the search. The authors should provide the full per-database search strings and search dates in an appendix, and should state which metadata fields (title/abstract/keywords) were matched.","section":"Section 2.2, Table 1"},{"comment":"There are internal inconsistencies between the results text and the extracted-data tables that prevent verification of the reported percentages. Reference [106] is cited in Section 3.9 as using perplexity/chrF/METEOR and in Section 3.10 as a question-answering study, but [106] does not appear in Table 2 (the list of 54 included studies) or in any extraction row of Tables 4-6. Similarly, reference [72] is included in the BLEU distribution list in Section 3.9 but is absent from Table 2 and Table 6. Either the tables omit included studies or the citations are wrong; in both cases, the percentages (e.g., 61% BLEU, 7% for perplexity) cannot be audited. The authors must reconcile the reference list, the included-study tables, and all in-text citation groupings.","section":"Section 3.9, Section 3.10, Appendix A (Tables 2, 6)"},{"comment":"The single-reviewer screening process is a load-bearing limitation for the claim that the review represents the relevant literature. The paper openly acknowledges this in Section 5, but Figure 1 and Section 2.3 do not report any reliability checks (e.g., a second reviewer on a random subset, or a comparison of ASReview's active-learning ranking against a manual gold standard). Since the search strings are also incompletely reported (see above) and the ASReview screening excludes papers based on title/abstract, the authors should either add a sensitivity analysis (e.g., re-screening a random sample by a second reviewer) or explicitly state the absence of any inter-rater validation and discuss how this could bias the frequency estimates.","section":"Section 2.3, Figure 1, Section 5"}],"minor_comments":[{"comment":"Typo: 'Scoupus' should be 'Scopus'.","section":"Section 2.2"},{"comment":"The statement that Turkish appears in 9% of papers cites [1, 3, 13, 25, 100, 116], but Table 4 assigns only [13, 25, 100, 116] (and [1]) to Turkish; [3] is listed for Bengali, Telugu, Khmer, and Malay. Please correct the citation grouping or the language assignment.","section":"Section 3.2"},{"comment":"The caption 'Search string composition' does not match the text's claim that the search strings themselves are detailed; consider renaming the table or providing the full strings as described in the major comment.","section":"Table 1"},{"comment":"The statement 'a count of 390k sentences is the average count' lacks a definition of which set the average is over (the languages with a single data source in Figure 15). Specify the denominator and whether the average is per language or per study.","section":"Section 3.8, RQ7"},{"comment":"The PRISMA flow diagram (Figure 1) would be clearer if the 'Excluded: 52' and 'Excluded: 486' boxes also stated the reasons (deduplication; EC1, EC3, EC4) on the diagram itself, as recommended in PRISMA 2020 templates.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful scope and an honest limitations section, but the quantitative results section needs substantial re-analysis before publication. The unit-of-analysis issue (mixing surveys and primary studies) is the most serious technical concern; the missing search strings and the data-table inconsistencies (references [106] and [72]) further undermine verifiability. The authors' emphasis on being 'the first' systematic review in this niche is plausible but should be supported by a brief comparison with the most recent related surveys (e.g., [86], [29]) to demonstrate what is genuinely new. I recommend major revision rather than rejection, as the core roadmap and discussion are valuable and the issues are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First systematic review on data-scarcity methods for generative LMs in low-resource languages. It is a solid piece of survey work: PRISMA-style screening, clearly defined RQs, quality scoring, and useful extraction tables. The headline findings—transformer dominance, concentration on a handful of LRLs, BLEU-heavy evaluation, translation as the main task—are broadly supported by the extracted data. The discussion is thoughtful, especially around why certain language families are overrepresented and the need for standardized reporting of corpus sizes.\n\nThe main soft spot is the unit of analysis in the method frequencies. The review includes survey/review papers under IC3, and Table 3 lists at least five of them ([29,86,108,109,123]) alongside primary studies in the same technique categories. Section 3.3 says each paper was categorized based on techniques it 'discussed or implemented'. That means a survey that mentions back-translation contributes to the 24% figure just like a paper that actually deployed it. The same issue runs through RQ3. The percentages should either be recomputed for empirical deployments only, or clearly labeled as 'mentioned' vs 'implemented'. This does not invalidate the review's map, but it does mean the frequency bars are a mix of description and deployment, and the claim that monolingual augmentation is the most common strategy should be softened.\n\nOther soft spots are more conventional: full search strings and search dates are not reported, screening was done by a single reviewer with ASReview, and non-English papers were excluded. The authors acknowledge the latter two in Section 5, which is honest, but the missing search details make exact replication harder. Minor internal inconsistencies—like classifying Turkish as LRL despite citing a source that calls it high-resource—are handled with an explicit rationale.\n\nBottom line: the review is a useful map, especially for someone entering low-resource NLG. The core findings survive the unit-of-analysis problem even if the exact percentages do not. It deserves a serious referee and should be published after revision; the fixes are straightforward, not structural.","headline":"A useful first systematic map of data-scarcity techniques for low-resource generative NLG, but the headline method frequencies are muddied by counting surveys and primary studies as equivalent units.","tokens_in":42208,"tokens_out":2799,"would_cite":false,"duration_ms":25233,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the first systematic review devoted specifically to data-scarcity strategies for generative language modelling in low-resource languages, synthesising 54 studies and finding transformer dominance, heavy…","keywords":["systematic review","low-resource languages","generative language modelling","data scarcity","data augmentation","back-translation","multilingual models","evaluation metrics"],"falsifier":"One concrete check would be to re-run the search while adding non-English papers and screening full text rather than title and abstract only; if the recovered set shifts the distribution of language families, the frequency of methods, or the share using BLEU, then the review's reported gaps and recommendations are artefacts of its inclusion criteria.","tokens_in":41293,"feed_emoji":"🌐","tokens_out":7293,"duration_ms":69229,"temperature":0.7,"pith_summary":"Generative language models increasingly serve high-resource languages, leaving speakers of low-resource languages behind. This paper argues that the research community has lacked a shared map of how to build such models when text data is scarce, and offers one: a systematic review of 54 studies. It claims this is the first review devoted specifically to data-scarcity strategies for generative language modelling in low-resource languages. The main findings are that transformer-based models dominate, that a small set of languages such as Bengali and Hindi receive most attention, and that evaluation is too inconsistent across studies to compare progress. If this map is accurate, it gives researchers a concrete basis for choosing augmentation methods, reporting data sizes, and designing human evaluation.","feed_headline":"54 studies: low-resource models stay concentrated, under-evaluated","feed_subtitle":"What a 54-study review found: transformer dominance, a few languages, no consistent metrics.","key_machinery":"The machinery of the review is its screening and extraction pipeline: a search expression built from generative, model-type, and language-type terms, run across multiple digital libraries; deduplication and two screening passes, the first assisted by an active-learning prioritisation tool and the second by manual eligibility checks; reference tracking to add five further studies; and a nine-question coding scheme covering languages, methods, augmentation types, publishers, architectures, effectiveness, data scarcity, evaluation, and task. A five-item quality-assessment rubric scores each retained study and is used to weigh the literature. This pipeline converts 642 raw records into 54 analysed studies and produces the frequency distributions that carry every headline finding.","core_discovery":"The paper's central claim is that the literature on generative modelling for low-resource languages clusters around a few technical fixes and a few languages, and that this cluster can be described systematically for the first time. From 54 retained studies, the review reports that augmenting monolingual text, back-translation, multilingual training, and prompt engineering are the main strategies; that translation is by far the most common generative task; that 76% of studies use transformer-based architectures; and that BLEU is the dominant evaluation metric while human evaluation is rare. It further claims that Indo-European languages make up a disproportionate share of the modelled languages, that data reporting is too varied to compare scarcity across studies, and that consistent evaluation standards are missing. The sympathetic reading of the contribution is the structured aggregation itself: a reproducible inventory of methods, languages, architectures, and evaluation practices for a field that previously lacked one.","pith_inferences":["Beyond the paper, the same data imply that the most commonly recommended fixes—back-translation and multilingual pooling—work best for languages that already have machine-translation systems or related neighbours, so the languages with least data may also be the least able to use the dominant methods.","The authors do not pursue this, but their proposed resource-level definition could be turned into a testable index: score languages on corpus size, tokenisation support, compute access, and number of fluent NLP researchers, then check whether the index predicts which languages appear in the literature.","A further extension would use the review's language-family observation to design a controlled experiment: train family-based models on several under-represented families and compare parameter efficiency and quality against monolingual and broad multilingual baselines.","The review's evaluation critique suggests a practical norm: any study claiming a low-resource language model is faithful to the language should report human evaluation alongside automatic metrics."],"forward_implications":["Translation is the dominant task, so the evidence for data-scarcity methods mostly validates them for machine translation; the same methods should not be assumed to work for summarisation, dialogue, or open-ended generation.","Monolingual augmentation, back-translation, multilingual training, and prompt engineering are the four most common strategies, with multilingual and family-of-languages approaches acting as an implicit form of data augmentation.","The transformer dominance and the prevalence of adapting pre-trained models mean that future low-resource work will likely build on existing transformer tooling rather than revisit earlier architectures.","Because BLEU dominates while human evaluation is rare, reported gains may reflect surface-level fluency rather than language fidelity; better evaluation is a precondition for claims about language preservation.","The uneven distribution of modelled languages implies that 'low-resource' is not a single category; a graded measure of data, compute, and researcher availability is needed to target support where it is most lacking."],"supporting_citations":[{"why":"Defines the systematic-review methodology and reporting flow that the paper follows.","marker":"[67]"},{"why":"Supplies the eligibility formulation used to turn the research objective into inclusion criteria and a search string.","marker":"[92]"},{"why":"Provides the active-learning screening tool used in the first study-selection pass.","marker":"[103]"},{"why":"Defines the low-resource/high-resource distinction that sets the review's scope and population.","marker":"[62]"},{"why":"The transformer architecture paper that grounds the finding that transformer-based models dominate low-resource generation.","marker":"[105]"},{"why":"A key included study whose multilingual modelling underpins the finding that multilingual training acts like extra training data.","marker":"[13]"},{"why":"An included study that anchors the language-family finding and the claim that related-language models can match larger models with fewer parameters.","marker":"[12]"},{"why":"Used to argue that BLEU scores are hard to compare across studies, supporting the lack-of-consistent-evaluation finding.","marker":"[73]"}],"fun_headline_variants":["Low-resource NLP review: transformers, few languages, weak eval","First systematic review: low-resource generative models still skewed","54 studies: low-resource generation stuck on transformers and BLEU","Low-resource LLMs: review finds few languages, no benchmark consistency","First LRL generative review: 76% transformers, BLEU dominates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole frequency map rests on the assumption that the 54 studies that survived the search and screening represent the relevant literature; the paper itself concedes that excluding non-English papers and screening titles and abstracts may have left relevant studies out.","fun_headline_variants_meta":{"raw":{"variants":["Low-resource NLP review: transformers, few languages, weak eval","First systematic review: low-resource generative models still skewed","54 studies: low-resource generation stuck on transformers and BLEU","Low-resource LLMs: review finds few languages, no benchmark consistency","First LRL generative review: 76% transformers, BLEU dominates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000847,"raw_usage":{"total_tokens":3681,"prompt_tokens":939,"completion_tokens":2742,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":2649}},"tokens_in":555,"tokens_out":2742,"duration_ms":20309,"temperature":1.0,"reasoning_tokens":2649,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:26:01.520978+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete check would be to re-run the search while adding non-English papers and screening full text rather than title and abstract only; if the recovered set shifts the distribution of language families, the frequency of methods, or the share using BLEU, then the review's reported gaps and recommendations are artefacts of its inclusion criteria.","supporting_citations":[],"review_version":1}