Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper claims to be the first systematic review devoted specifically to data-scarcity strategies for generative language modelling in low-resource languages, synthesising 54 studies and finding transformer dominance, heavy…

desk verdict A useful first systematic map of data-scarcity techniques for low-resource generative NLG, but the headline method frequencies are muddied by counting surveys and primary studies as equivalent units. read the letter →

arxiv 2505.04531 v2 pith:4VUCNUAB submitted 2025-05-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords systematicreviewlow-resourcelanguagesgenerativelanguagemodellingdatascarcityaugmentationback-translationmultilingualmodelsevaluationmetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Generative language models increasingly serve high-resource languages, leaving speakers of low-resource languages behind. This paper argues that the research community has lacked a shared map of how to build such models when text data is scarce, and offers one: a systematic review of 54 studies. It claims this is the first review devoted specifically to data-scarcity strategies for generative language modelling in low-resource languages. The main findings are that transformer-based models dominate, that a small set of languages such as Bengali and Hindi receive most attention, and that evaluation is too inconsistent across studies to compare progress. If this map is accurate, it gives researchers a concrete basis for choosing augmentation methods, reporting data sizes, and designing human evaluation.

What carries the argument

The machinery of the review is its screening and extraction pipeline: a search expression built from generative, model-type, and language-type terms, run across multiple digital libraries; deduplication and two screening passes, the first assisted by an active-learning prioritisation tool and the second by manual eligibility checks; reference tracking to add five further studies; and a nine-question coding scheme covering languages, methods, augmentation types, publishers, architectures, effectiveness, data scarcity, evaluation, and task. A five-item quality-assessment rubric scores each retained study and is used to weigh the literature. This pipeline converts 642 raw records into 54 analysed studies and produces the frequency distributions that carry every headline finding.

What would settle it

One concrete check would be to re-run the search while adding non-English papers and screening full text rather than title and abstract only; if the recovered set shifts the distribution of language families, the frequency of methods, or the share using BLEU, then the review's reported gaps and recommendations are artefacts of its inclusion criteria.

Watch

Extended reading notes

Core claim

The paper's central claim is that the literature on generative modelling for low-resource languages clusters around a few technical fixes and a few languages, and that this cluster can be described systematically for the first time. From 54 retained studies, the review reports that augmenting monolingual text, back-translation, multilingual training, and prompt engineering are the main strategies; that translation is by far the most common generative task; that 76% of studies use transformer-based architectures; and that BLEU is the dominant evaluation metric while human evaluation is rare. It further claims that Indo-European languages make up a disproportionate share of the modelled languages, that data reporting is too varied to compare scarcity across studies, and that consistent evaluation standards are missing. The sympathetic reading of the contribution is the structured aggregation itself: a reproducible inventory of methods, languages, architectures, and evaluation practices for a field that previously lacked one.

Load-bearing premise

The whole frequency map rests on the assumption that the 54 studies that survived the search and screening represent the relevant literature; the paper itself concedes that excluding non-English papers and screening titles and abstracts may have left relevant studies out.

Editorial extensions

If this is right

  • Translation is the dominant task, so the evidence for data-scarcity methods mostly validates them for machine translation; the same methods should not be assumed to work for summarisation, dialogue, or open-ended generation.
  • Monolingual augmentation, back-translation, multilingual training, and prompt engineering are the four most common strategies, with multilingual and family-of-languages approaches acting as an implicit form of data augmentation.
  • The transformer dominance and the prevalence of adapting pre-trained models mean that future low-resource work will likely build on existing transformer tooling rather than revisit earlier architectures.
  • Because BLEU dominates while human evaluation is rare, reported gains may reflect surface-level fluency rather than language fidelity; better evaluation is a precondition for claims about language preservation.
  • The uneven distribution of modelled languages implies that 'low-resource' is not a single category; a graded measure of data, compute, and researcher availability is needed to target support where it is most lacking.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the same data imply that the most commonly recommended fixes—back-translation and multilingual pooling—work best for languages that already have machine-translation systems or related neighbours, so the languages with least data may also be the least able to use the dominant methods.
  • The authors do not pursue this, but their proposed resource-level definition could be turned into a testable index: score languages on corpus size, tokenisation support, compute access, and number of fluent NLP researchers, then check whether the index predicts which languages appear in the literature.
  • A further extension would use the review's language-family observation to design a controlled experiment: train family-based models on several under-represented families and compare parameter efficiency and quality against monolingual and broad multilingual baselines.
  • The review's evaluation critique suggests a practical norm: any study claiming a low-resource language model is faithful to the language should report human evaluation alongside automatic metrics.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper presents a PRISMA-style systematic review of 54 studies on strategies to overcome data scarcity in generative language modelling for low-resource languages (LRLs). The review identifies and categorizes technical approaches (monolingual augmentation, back-translation, multilingual modelling, prompt engineering, etc.), analyzes language-family representation, model architectures, and evaluation practices, and reports research gaps and recommendations. The central claims are that transformer-based models dominate, that a small subset of LRLs accounts for most research, and that evaluation is inconsistent across studies.

Significance. If the quantitative map is robust, the review is a useful contribution: it consolidates scattered evidence, provides a structured taxonomy of data-scarcity methods, and gives concrete recommendations for future research, including language-family-centric modelling and more nuanced definitions of 'low-resource'. The paper is transparent about many limitations, ships detailed extraction tables (Tables 2-6), and uses a PICO framework and quality assessment, which are strengths. The falsifiable field-level claims, such as the dominance of transformers and the concentration of studies on a few languages, are important for researchers and funding bodies, and the discussion of geopolitical factors in language selection is a valuable framing.

major comments (4)
  1. [Section 3.3, Table 3; Section 2.4; Section 5] The central frequency map for RQ2/RQ3 is built on a unit-of-analysis problem: the corpus mixes primary empirical studies with survey/review papers, yet each paper is assigned to technique categories 'based on the techniques it discussed or implemented'. Table 3 explicitly lists surveys [29, 86, 109, 123] under 'Augment Monolingual Data' and [86, 108] under 'Back-Translation', even though these papers review the literature rather than deploy the methods. Consequently, the reported 26% for monolingual augmentation and 24% for back-translation (Figure 7) conflate description with implementation. Because these percentages underpin the review's main empirical conclusions, the authors should re-run the analysis on primary studies only (e.g., QA2=Yes) and report both raw and sensitivity-adjusted frequencies, or clearly separate 'discussed' from 'implemented' categories in the figures and tables.
  2. [Section 2.2, Table 1] The search strategy is not fully reproducible. The text states that 'Table 1 details the search strings that were used for the various digital repositories', but Table 1 shows only four string components ('generative', 'language model OR text model', 'low resource language OR minority language OR endangered language', 'exclusion: classification') with no per-database Boolean syntax, no field restrictions, and no search dates. The PRISMA flow diagram (Figure 1) reports 642 initial records, but without the exact queries and access dates, readers cannot verify the coverage or repeat the search. The authors should provide the full per-database search strings and search dates in an appendix, and should state which metadata fields (title/abstract/keywords) were matched.
  3. [Section 3.9, Section 3.10, Appendix A (Tables 2, 6)] There are internal inconsistencies between the results text and the extracted-data tables that prevent verification of the reported percentages. Reference [106] is cited in Section 3.9 as using perplexity/chrF/METEOR and in Section 3.10 as a question-answering study, but [106] does not appear in Table 2 (the list of 54 included studies) or in any extraction row of Tables 4-6. Similarly, reference [72] is included in the BLEU distribution list in Section 3.9 but is absent from Table 2 and Table 6. Either the tables omit included studies or the citations are wrong; in both cases, the percentages (e.g., 61% BLEU, 7% for perplexity) cannot be audited. The authors must reconcile the reference list, the included-study tables, and all in-text citation groupings.
  4. [Section 2.3, Figure 1, Section 5] The single-reviewer screening process is a load-bearing limitation for the claim that the review represents the relevant literature. The paper openly acknowledges this in Section 5, but Figure 1 and Section 2.3 do not report any reliability checks (e.g., a second reviewer on a random subset, or a comparison of ASReview's active-learning ranking against a manual gold standard). Since the search strings are also incompletely reported (see above) and the ASReview screening excludes papers based on title/abstract, the authors should either add a sensitivity analysis (e.g., re-screening a random sample by a second reviewer) or explicitly state the absence of any inter-rater validation and discuss how this could bias the frequency estimates.
minor comments (5)
  1. [Section 2.2] Typo: 'Scoupus' should be 'Scopus'.
  2. [Section 3.2] The statement that Turkish appears in 9% of papers cites [1, 3, 13, 25, 100, 116], but Table 4 assigns only [13, 25, 100, 116] (and [1]) to Turkish; [3] is listed for Bengali, Telugu, Khmer, and Malay. Please correct the citation grouping or the language assignment.
  3. [Table 1] The caption 'Search string composition' does not match the text's claim that the search strings themselves are detailed; consider renaming the table or providing the full strings as described in the major comment.
  4. [Section 3.8, RQ7] The statement 'a count of 390k sentences is the average count' lacks a definition of which set the average is over (the languages with a single data source in Figure 15). Specify the denominator and whether the average is per language or per study.
  5. [Section 3.1] The PRISMA flow diagram (Figure 1) would be clearer if the 'Excluded: 52' and 'Excluded: 486' boxes also stated the reasons (deduplication; EC1, EC3, EC4) on the diagram itself, as recommended in PRISMA 2020 templates.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review synthesizes 54 external studies; its only self-citation is incidental and none of its findings are forced by construction.

full rationale

This paper is a systematic review, not a derivation or prediction exercise. Its central claims (transformer dominance, concentration on a few LRLs, inconsistent evaluation, most-used data-scarcity methods) are frequency summaries extracted from 54 included studies. There are no fitted parameters, no equations whose outputs coincide with their inputs, and no uniqueness theorem imported from the authors' prior work. The only self-citation, [53], appears in the introduction as an example of a transformer-based classification application ('with the transformer architecture enhancing NLP tasks such as machine translation [77], classification [53] and named entity recognition [6]'). It is not load-bearing for the review's inclusion criteria, search strategy, research questions, or conclusions. The skeptical concern that RQ2/RQ3 counts mix primary studies with survey papers (e.g., [29], [86], [108], [109], [123] are categorized under technical methods) is a methodological unit-of-analysis and validity issue, not circularity: the counts describe what the included corpus says, and the corpus is not defined in terms of the counts. Likewise, the acknowledged limitations (non-English papers excluded, single-reviewer title/abstract screening, reliance on ASReview) affect recall and representativeness but do not make any result equivalent to its input by definition. The novelty claim of being 'the first systematic review focused specifically' on this topic rests on the reported search and screening process rather than on a self-referential construction. Under the stated rules, this warrants a non-finding: score 0, no circular steps identified.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims of this systematic review rest on the representativeness of the literature search, the accuracy of manual data extraction, and the validity of the self-defined quality assessment. These are domain assumptions about review methodology, not mathematical axioms or fitted parameters.

assumptions (3)
  • domain assumption The search and screening process (PRISMA plus ASReview) captured a representative subset of the relevant literature, and the 54 included studies are sufficient to support the reported trends.
    Section 2.2-2.3 and Section 5; the limitations section acknowledges single-reviewer screening and non-English exclusion, which could bias the frequencies and conclusions.
  • domain assumption The manual categorization of technical methods, languages, and language families accurately reflects the content of the 54 studies.
    Section 2.4; minor inconsistencies in the results (e.g., Turkish frequency, incomplete search string table) indicate possible extraction or reporting errors.
  • domain assumption The quality assessment scoring (QA1-QA5) is a valid proxy for study relevance and quality in this field.
    Section 2.4; the scoring rubric is self-defined and not validated against external quality measures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review." pith.science (2026). https://pith.science/paper/4VUCNUAB

@misc{pith2026250504531,
  author       = {Pith},
  title        = {Pith review of: Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4VUCNUAB}},
  note         = {Machine review of arXiv:2505.04531}
}
read the original abstract

Generative language modelling has surged in popularity with the emergence of services such as ChatGPT and Google Gemini. While these models have demonstrated transformative potential in productivity and communication, they overwhelmingly cater to high-resource languages like English. This has amplified concerns over linguistic inequality in natural language processing (NLP). This paper presents the first systematic review focused specifically on strategies to address data scarcity in generative language modelling for low-resource languages (LRL). Drawing from 54 studies, we identify, categorise and evaluate technical approaches, including monolingual data augmentation, back-translation, multilingual training, and prompt engineering, across generative tasks. We also analyse trends in architecture choices, language family representation, and evaluation methods. Our findings highlight a strong reliance on transformer-based models, a concentration on a small subset of LRLs, and a lack of consistent evaluation across studies. We conclude with recommendations for extending these methods to a wider range of LRLs and outline open challenges in building equitable generative language systems. Ultimately, this review aims to support researchers and developers in building inclusive AI tools for underrepresented languages, a necessary step toward empowering LRL speakers and the preservation of linguistic diversity in a world increasingly shaped by large-scale language technologies.

Figures

Figures reproduced from arXiv: 2505.04531 by the authors.

Figure 1
Figure 1. This PRISMA flow diagram describes the process of using the PICO framework-based search string to find and filter relevant [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Distribution of papers based on whether they tackle building a language model for a low-resource language (QA1), and [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Distribution of papers based on whether they discuss methods for addressing data scarcity (QA3), and whether they present [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: CORE ranking distribution for conference papers and Journal Impact Factor 2024 distribution for journal papers (QA5). [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Visualisation of the most frequently modelled low-resource languages (RQ1), shown as a word cloud and a bar chart of the top [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Distribution of the most frequently modelled low-resource language families by papers in this systematic review (RQ1). [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Distribution of the technical methods used to overcome data scarcity in building generative language models by frequency [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Distribution of the monolingual text data augmentation techniques identified in this systematic review by frequency (RQ3). [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Distribution of publishers by frequency (RQ4). [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Distribution of sources by frequency (RQ4). [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]
Figure 11
Figure 11. Figure 11: Trend of survey vs empirical papers by year, and distribution of paper types by publication venue (e.g., conference, journal, [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]
Figure 12
Figure 12. Figure 12: Distribution of publications by country. [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 13
Figure 13. Figure 13: Distribution of the model architectures used by papers in this systematic review by frequency (RQ5). [PITH_FULL_IMAGE:figures/full_fig_p018_13.png]
Figure 14
Figure 14. Figure 14: Model distribution by frequency within the systematic review (RQ5). Note: some studies report general architectures rather [PITH_FULL_IMAGE:figures/full_fig_p018_14.png]
Figure 15
Figure 15. Figure 15: Distribution of sentence counts by language (RQ7). [PITH_FULL_IMAGE:figures/full_fig_p019_15.png]
Figure 16
Figure 16. Figure 16: Distribution of evaluation method by frequency (RQ8). [PITH_FULL_IMAGE:figures/full_fig_p020_16.png]
Figure 17
Figure 17. Figure 17: Distribution of language generation task by method (RQ9). [PITH_FULL_IMAGE:figures/full_fig_p020_17.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Representational Equality in Cross-country Value Simulation: A Systematic Analysis of Large Language Models

    cs.CY 2026-08 conditional novelty 6.0 of 10

    LLM-based value simulation is systematically more accurate for wealthy, high-governance, individualist countries, and common interventions rarely fix the imbalance.

  2. Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Synthetic data for low-resource spoken language models creates a Stability-Expressivity Gap that DGSA and TDSC self-alignment close, enabling SOTA Thai TTS and first Lao zero-shot voice cloning.

Reference graph

Works this paper leans on

269 extracted references · 47 canonical work pages · cited by 2 Pith papers

  1. [106]

    Tam Minh Vo and Khiem Vinh Tran. 2023. Generative Pre-trained Transformer for Vietnamese Community-based COVID-19 Question Answering. doi:10.48550/arXiv.2310.14602 arXiv:2310.14602 [cs]

  2. [72]

    Hong-Hai Phan-Vu, Van-Nam Nguyen, Viet-Trung Tran, and Phan-Thuan Do. 2017. Towards State-of-the-art English-Vietnamese Neural Machine Translation. In Proceedings of the 8th International Symposium on Information and Communication Technology (New York, NY, USA, 2017-12-07) (SoICT ’17). Association for Computing Machinery, 120–126. doi:10.1145/3155133.3155205

  3. [1]

    Emre Can Acikgoz, Mete Erdogan, and Deniz Yuret. 2024. Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking. doi:10.48550/arXiv.2405.04685 arXiv:2405.04685 [cs]

  4. [2]

    Parul Agarwal, Aisha Asif, Shantipriya Parida, Sambit Sekhar, Satya Ranjan Dash, and Subhadarshi Panda. 2023. Generative Chatbot Adaptation for Odia Language: A Critical Evaluation . doi:10.1109/CCPIS59145.2023.10291329 34 Josh McGiff and Nikola S. Nikolov

  5. [3]

    Sumit Agarwal, Suraj Tripathi, Teruko Mitamura, and Carolyn Penstein Rose. 2022. Zero-shot cross-lingual open domain question answering. In Proceedings of the Workshop on Multilingual Information Access (MIA) (Seattle, USA, 2022-07), Akari Asai, Eunsol Choi, Jonathan H. Clark, Junjie Hu, Chia-Hsuan Lee, Jungo Kasai, Shayne Longpre, Ikuya Yamada, and Rui Z...

  6. [4]

    Benyamin Ahmadnia and Bonnie J. Dorr. 2019. Augmenting Neural Machine Translation through Round-Trip Training Approach. 9, 1 (2019), 268–278. doi:10.1515/comp-2019-0019 Publisher: De Gruyter Open Access

  7. [5]

    Talal Almutiri and Farrukh Nadeem. 2022. Markov models applications in natural language processing: a survey. Int. J. Inf. Technol. Comput. Sci 2 (2022), 1–16

  8. [6]

    Mikhail Arkhipov, Maria Trofimova, Yurii Kuratov, and Alexey Sorokin. 2019. Tuning multilingual transformers for language-specific named entity recognition. In Proceedings of the 7th Workshop on Balto-Slavic Natural Language Processing . 89–93

Show all 269 references
  1. [7]

    Baligh Babaali, Mohammed Salem, and Nawaf R. Alharbe. 2024. Breaking language barriers with ChatGPT: enhancing low-resource machine translation between algerian arabic and MSA. (2024). doi:10.1007/s41870-024-01926-7

  2. [8]

    Oliver Bendel and Dalil Jabou. 2024. @llegra: a chatbot for Vallader. 16, 4 (2024), 2035–2045. doi:10.1007/s41870-024-01779-0

  3. [9]

    Tucker Berckmann and Berkan Hiziroglu. 2020. Low-Resource Translation as Language Modeling. InProceedings of the Fifth Conference on Machine Translation (Online, 2020-11), Loïc Barrault, Ondřej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussà, Christian Federmann,...

  4. [10]

    Rajat Subhra Bhowmick, Isha Ganguli, Ananya Paul, Jayanta Paul, and Jaya Sil. 2023. Improving Indic code-mixed to monolingual translation using Mixed Script Augmentation, Generation & Transfer Learning. (2023). doi:10.1145/3606695 Just Accepted

  5. [11]

    Steven Bird. 2022. Local Languages, Third Spaces, and other High-Resource Scenarios. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association ...

  6. [12]

    Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Langua...

  7. [13]

    Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K

    Tyler A. Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K. Bergen. 2023. When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages. doi:10.48550/arXiv.2311.09205 arXiv:2311.09205 [cs]

  8. [14]

    Sayan Chatterjee, Ching Louis Liu, Gareth Rowland, and Tim Hogarth. 2024. The Impact of AI Tool on Engineering at ANZ Bank An Empirical Study on GitHub Copilot within Corporate Environment. arXiv preprint arXiv:2402.05636 (2024)

  9. [15]

    Sandipan Dandapat and Christian Federmann. 2018. Iterative Data Augmentation for Neural Machine Translation: a Low Resource Case Study for English-Telugu. In Proceedings of the 21st Annual Conference of the European Association for Machine Translation (Alicante, Spain, 2018-05...

  10. [16]

    Lorenzo De Mattei, Michele Cafagna, Felice Dell’Orletta, Malvina Nissim, Marco Guerini, Aptus AI, and Fondazione Bruno Kessler. 2020. GePpeTto Carves Italian into a Language Model. Computational Linguistics CLiC-it 2020 (2020), 136

  11. [17]

    Department of Public Expenditure, Infrastructure, Public Service Reform and Digitalisation. 2025. Guidelines for the Responsible Use of AI in the Public Service. Available at: https://www.gov.ie/en/department-of-public-expenditure-infrastructure-public-service-reform-and- digi...

  12. [18]

    Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  13. [19]

    Chenhe Dong, Yinghui Li, Haifan Gong, Miaoxin Chen, Junxin Li, Ying Shen, and Min Yang. 2022. A survey of natural language generation. Comput. Surveys 55, 8 (2022), 1–38

  14. [20]

    C. M. Downey, Terra Blevins, Dhwani Serai, Dwija Parikh, and Shane Steinert-Threlkeld. 2024. Targeted Multilingual Adaptation for Low-resource Language Families. doi:10.48550/arXiv.2405.12413 arXiv:2405.12413 [cs]

  15. [21]

    Matthew S Dryer. 2007. Word order. Language typology and syntactic description 1, 61-131 (2007), 1–1

  16. [22]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)

  17. [23]

    Fei Gao, Jinhua Zhu, Lijun Wu, Yingce Xia, Tao Qin, Xueqi Cheng, Wengang Zhou, and Tie-Yan Liu. 2019. Soft Contextual Data Augmentation for Neural Machine Translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Florence, Italy, ...

  18. [24]

    Silin Gao, Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020. Paraphrase Augmented Task-Oriented Dialog Generation . Online. doi:10.18653/v1/2020.acl- main.60 Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review 35

  19. [25]

    Daniela Gerz, Ivan Vulić, Edoardo Ponti, Jason Naradowsky, Roi Reichart, and Anna Korhonen. 2018. Language Modeling for Morphologically Rich Languages: Character-Aware Modeling for Word-Level Prediction. 6 (2018), 451–465. doi:10.1162/tacl_a_00032

  20. [26]

    The Guardian. 2024. Brat summer: is the long era of clean living finally over? https://www.theguardian.com/lifeandstyle/article/2024/jul/16/brat- summer-is-the-long-era-of-clean-living-finally-over. Accessed: 2025-06-19

  21. [27]

    Ping Guo, Yubing Ren, Yue Hu, Yunpeng Li, Jiarui Zhang, Xingsheng Zhang, and Heyan Huang. 2024. Teaching Large Language Models to Translate on Low-resource Languages with Textbook Prompting. In Proceedings of the 2024 Joint International Conference on Computational Linguistics...

  22. [28]

    Zhu, Guandan Chen, Kaibo Zhang, Boxing Chen, and Fei Huang

    Zilu Guo, Zhongqiang Huang, Kenny Q. Zhu, Guandan Chen, Kaibo Zhang, Boxing Chen, and Fei Huang. 2021. Automatically Paraphrasing via Sentence Reconstruction and Round-trip Translation. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (...

  23. [29]

    Rejwanul Haque, Chao-Hong Liu, and Andy Way. 2021. Recent advances of low-resource neural machine translation. 35, 4 (2021), 451–474. doi:10.1007/s10590-021-09281-1

  24. [30]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. [n. d.]. Training Compute-Optimal Large Language Models. ([n. d.])

  25. [31]

    Heidi Ahmed Holiel, Nancy Mohamed, Arwa Ahmed, and Walaa Medhat. 2023. English-Arabic Text Translation and Abstractive Summarization Using Transformers. doi:10.1109/AICCSA59173.2023.10479257

  26. [32]

    Kung Yin Hong, Lifeng Han, Riza Batista-Navarro, and Goran Nenadic. 2024. CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation. doi:10.48550/arXiv.2405.08172 arXiv:2405.08172 [cs]

  27. [33]

    James Hutson, Pace Ellsworth, and Matt Ellsworth. 2024. Preserving linguistic diversity in the digital age: a scalable model for cultural heritage continuity. Journal of Contemporary Language Research 3, 1 (2024)

  28. [34]

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce C...

  29. [35]

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020)

  30. [36]

    Cherry Khosla and Baljit Singh Saini. 2020. Enhancing performance of deep learning models with different data augmentation techniques: A survey. In 2020 International Conference on Intelligent Engineering and Management (ICIEM) . IEEE, 79–85

  31. [37]

    Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale. 2023. Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback. arXiv preprint arXiv:2303.05453 (2023)

  32. [38]

    Piotr Kłosowski. 2018. Deep learning for natural language processing and language modelling. In 2018 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA). IEEE, 223–228

  33. [39]

    Viet Dac Lai, Nghia Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning. InFindings of the Association for Computational L...

  34. [40]

    Siobhán Ní Laoire. 2016. Irish-English Code-switching: a Sociolinguistic Perspective. In Sociolinguistics in Ireland, Raymond Hickey (Ed.). Palgrave Macmillan UK, 81–106. doi:10.1057/9781137453471_4

  35. [41]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2023. Bloom: A 176b-parameter open-access multilingual language model. (2023)

  36. [42]

    Julie Anne Legate. 1999. The morphosyntax of Irish agreement. MIT working papers in linguistics 33 (1999), 219–240

  37. [43]

    Bin Li, Yixuan Weng, Bin Sun, and Shutao Li. 2022. A Multi-tasking and Multi-stage Chinese Minority Pre-trained Language Model. In Machine Translation (Singapore, 2022), Tong Xiao and Juan Pino (Eds.). Springer Nature, 93–105. doi:10.1007/978-981-19-7960-6_10

  38. [44]

    Zhuang Li, Levon Haroutunian, Raj Tumuluri, Philip Cohen, and Reza Haf. 2024. Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach. In Findings of the Association for Computational Linguistics: EACL 2024 (St. Julian’s,...

  39. [45]

    Ying Lian, Huiting Tang, Mengting Xiang, and Xuefan Dong. 2024. Public attitudes and sentiments toward ChatGPT in China: A text mining analysis based on social media. Technology in Society 76 (2024), 102442

  40. [46]

    Godfrey Lienhardt, John Middleton, and David Tait. 1958. The Western Dinka. Tribes without rulers: Studies in african segmentary systems (1958), 97–135

  41. [47]

    Evelyn Kai-Yan Liu. 2022. Low-Resource Neural Machine Translation: A Case Study of Cantonese. In Proceedings of the Ninth Workshop on NLP for Similar Languages, Varieties and Dialects (Gyeongju, Republic of Korea, 2022-10), Yves Scherrer, Tommi Jauhiainen, Nikola Ljubešić, Pre...

  42. [48]

    Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics 8 (2020), 726–742. 36 Jos...

  43. [49]

    Alexandre Magueresse, Vincent Carles, and Evan Heetderks. 2020. Low-resource languages: A review of past work and future challenges. arXiv preprint arXiv:2006.07264 (2020)

  44. [50]

    Mieradilijiang Maimaiti, Yang Liu, Huanbo Luan, Zegao Pan, and Maosong Sun. 2021. Improving Data Augmentation for Low-Resource NMT Guided by POS-Tagging and Paraphrase Embedding. 20, 6 (2021), 107:1–107:21. doi:10.1145/3464427

  45. [51]

    Zhuoyuan Mao and Yen Yu. 2024. Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages. In Proceedings of the Seventh Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2024) (Bangkok, Thail...

  46. [52]

    Benjamin Marie and Atsushi Fujita. 2020. Synthesizing Parallel Data of User-Generated Texts with Zero-Shot Neural Machine Translation. 8 (2020), 710–725. doi:10.1162/tacl_a_00341

  47. [53]

    Josh McGiff and Nikola S Nikolov. 2024. Bridging the gap in online hate speech detection: A comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter. Applied and Computational Engineering 64 (2024), 63–68

  48. [54]

    Chenggang Mi, Shaoliang Xie, and Yi Fan. 2024. Multi-granularity Knowledge Sharing in Low-resource Neural Machine Translation. 23, 2 (2024), 31:1–31:19. doi:10.1145/3639930

  49. [55]

    Ibomoiye Domor Mienye, Theo G Swart, and George Obaido. 2024. Recurrent neural networks: A comprehensive review of architectures, variants, and applications. Information 15, 9 (2024), 517

  50. [56]

    Agnieszka Mikołajczyk and Michał Grochowski. 2018. Data augmentation for improving deep learning in image classification problem. In 2018 international interdisciplinary PhD workshop (IIPhDW) . IEEE, 117–122

  51. [57]

    Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel. 2023. Scaling Data-Constrained Language Models. In Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globers...

  52. [58]

    Martin Müller and Florian Laurent. 2022. Cedille: A large autoregressive french language model. arXiv preprint arXiv:2202.03371 (2022)

  53. [59]

    Alhassan Mumuni and Fuseini Mumuni. 2022. Data augmentation: A comprehensive survey of modern approaches. Array 16 (2022), 100258

  54. [60]

    Dinesh Kumar Nanduri and Elizabeth M Bonsignore. 2023. Revitalizing Endangered Languages: AI-powered language learning as a catalyst for language appreciation. arXiv preprint arXiv:2304.09394 (2023)

  55. [61]

    Reuben Ng and Ting Yu Joanne Chow. 2024. Powerful tool or too powerful? Early public discourse about ChatGPT across 4 million tweets. Plos one 19, 3 (2024), e0296882

  56. [62]

    Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, and Monojit Choudhury. 2024. The Zeno’s Paradox of ‘Low- Resource’ Languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal...

  57. [63]

    L. N. A. S. H. Nissanka, B. H. R. Pushpananda, and A. R. Weerasinghe. 2020. Exploring Neural Machine Translation for Sinhala-Tamil Languages Pair. In 2020 20th International Conference on Advances in ICT for Emerging Regions (ICTer) (2020-11). 202–207. doi:10.1109/ICTer51097.2...

  58. [64]

    Mitodru Niyogi and Arnab Bhattacharya. 2024. Paramanu: A Family of Novel Efficient Indic Generative Foundation Language Models. doi:10. 48550/arXiv.2401.18034 arXiv:2401.18034 [cs]

  59. [65]

    Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages. In Proceedings of the 1st Workshop on Multilingual Representation Learning . 116–126

  60. [66]

    Arantxa Otegi, Aitor Agirre, Jon Ander Campos, Aitor Soroa, and Eneko Agirre. 2020. Conversational Question Answering in Low Resource Scenarios: A Dataset and Case Study for Basque. In Proceedings of the Twelfth Language Resources and Evaluation Conference (Marseille, France, ...

  61. [67]

    Matthew J Page, David Moher, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al. 2021. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic...

  62. [68]

    ChaeHun Park, Koanho Lee, Hyesu Lim, Jaeseok Kim, Junmo Park, Yu-Jung Heo, Du-Seong Chang, and Jaegul Choo. 2024. Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering. In Findings of the Association for Computational Linguisti...

  63. [69]

    Bruno Peixoto, Rafael Pinto, Miguel Melo, Luciana Cabral, and Maximino Bessa. 2021. Immersive virtual reality for foreign language education: A PRISMA systematic review. IEEE Access 9 (2021), 48952–48962

  64. [70]

    Hieu Pham, Xinyi Wang, Yiming Yang, and Graham Neubig. 2021. Meta Back-translation. doi:10.48550/arXiv.2102.07847 arXiv:2102.07847 [cs]

  65. [71]

    Hanh Pham Van and Huong Le Thanh. 2022. Improving Khmer-Vietnamese Machine Translation with Data Augmentation methods. InProceedings of the 11th International Symposium on Information and Communication Technology (New York, NY, USA, 2022-12-01) (SoICT ’22). Association for Ove...

  66. [73]

    Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. In Proceedings of the Third Conference on Machine Translation: Research Papers (Brussels, Belgium, 2018-10), Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Hu...

  67. [74]

    QuantumBlack, AI by McKinsey. 2025. The state of AI: How organizations are rewiring to capture value. Online report, McKinsey & Company. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai Published March 12, 2025; accessed June 18, 2025

  68. [75]

    Alec Radford. 2018. Improving language understanding by generative pre-training. (2018)

  69. [76]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  70. [77]

    Alessandro Raganato and Jörg Tiedemann. 2018. An analysis of encoder representations in transformer-based machine translation. In Proceedings of the 2018 EMNLP workshop BlackboxNLP: analyzing and interpreting neural networks for NLP . 287–297

  71. [78]

    Georg Rehm and Andy Way. 2023. European language equality: A strategic agenda for digital language equality . Springer Nature

  72. [79]

    Muhammad Razif Rizqullah, Ayu Purwarianti, and Alham Fikri Aji. 2023. QASiNa: Religious Domain Question Answering Using Sirah Nabawiyah . doi:10.1109/ICAICTA59291.2023.10390123

  73. [80]

    Brian Roark, Murat Saraclar, and Michael Collins. 2007. Discriminative n-gram language modeling. Computer Speech & Language 21, 2 (2007), 373–392

  74. [81]

    Lekhraj Saini and Deepti Vidhyarthi. 2023. Bidirectional English-Marathi Translation using Pretrained Models: A Comparative Study of Different Pre-Trained Models. doi:10.1109/INCOFT60753.2023.10425770

  75. [82]

    Shahidul Salim, Hasan Murad, Dola Das, and Faisal Ahmed

    Md. Shahidul Salim, Hasan Murad, Dola Das, and Faisal Ahmed. 2023. BanglaGPT: A Generative Pretrained Transformer-Based Model for Bangla Language. doi:10.1109/ICICT4SD59951.2023.10303383

  76. [83]

    Lamyae Sardi, Ali Idri, and José Luis Fernández-Alemán. 2017. A systematic review of gamification in e-Health. Journal of biomedical informatics 71 (2017), 31–48

  77. [84]

    Barbara Scalvini and Iben Nyholm Debess. 2024. Evaluating the Potential of Language-family-specific Generative Models for Low-resource Data Augmentation: A Faroese Case Study. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Reso...

  78. [85]

    Sheikh Shafayat, H. M. Quamran Hasan, Minhajur Rahman Chowdhury Mahim, Rifki Afina Putri, James Thorne, and Alice Oh. 2024. BEnQA: A Question Answering and Reasoning Benchmark for Bengali and English. doi:10.48550/arXiv.2403.10900 arXiv:2403.10900 [cs]

  79. [86]

    Shumin Shi, Xing Wu, Rihai Su, and Heyan Huang. 2022. Low-resource Neural Machine Translation: Methods and Trends. 21, 5 (2022), 103:1–103:22. doi:10.1145/3524300

  80. [87]

    Connor Shorten and Taghi M Khoshgoftaar. 2019. A survey on image data augmentation for deep learning. Journal of big data 6, 1 (2019), 1–48

  81. [88]

    Connor Shorten, Taghi M Khoshgoftaar, and Borko Furht. 2021. Text data augmentation for deep learning. Journal of big Data 8, 1 (2021), 101

  82. [89]

    Antoine Simoulin and Benoit Crabbé. 2021. Un modèle Transformer Génératif Pré-entrainé pour le _ français. In Traitement Automatique des Langues Naturelles. ATALA, 246–255

  83. [90]

    Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O’Mahony, et al. 2024. Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning. In Proceedings of the 62n...

  84. [91]

    Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, and Valentin Malykh. 2022. Ask Me Anything in Your Native Language. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Seattle...

  85. [92]

    Patricia W Stone. 2002. Popping the (PICO) question in research and evidence-based practice. Applied Nursing Research 15, 3 (2002), 197–198

  86. [93]

    Shang-Yu Su, Chao-Wei Huang, and Yun-Nung Chen. 2019. Dual Supervised Learning for Natural Language Understanding and Generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). ...

  87. [94]

    Yuan Sun, Chaofan Chen, Tianci Xia, and Xiaobing Zhao. 2019. QuGAN: Quasi Generative Adversarial Network for Tibetan Question Answering Corpus Generation. doi:10.1109/ACCESS.2019.2934581

  88. [95]

    Ashwani Tanwar and Prasenjit Majumder. 2020. Translating Morphologically Rich Indian Languages under Zero-Resource Conditions. 19, 6 (2020), 85:1–85:15. doi:10.1145/3407912

  89. [96]

    Luke Taylor and Geoff Nitschke. 2018. Improving deep learning with generic data augmentation. In 2018 IEEE symposium series on computational intelligence (SSCI). IEEE, 1542–1547

  90. [97]

    NLLB Team et al. 2024. Scaling neural machine translation to 200 languages. Nature 630, 8018 (2024), 841. 38 Josh McGiff and Nikola S. Nikolov

  91. [98]

    The World Bank. 2024. Research and development expenditure (% of GDP) - South Sudan. https://data.worldbank.org/indicator/GB.XPD.RSDV. GD.ZS?locations=ZS Accessed: 2025-06-09

  92. [99]

    Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023. Large language models in medicine. Nature medicine 29, 8 (2023), 1930–1940

  93. [100]

    Cagri Toraman. 2024. LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language. doi:10.48550/arXiv. 2405.07745 arXiv:2405.07745 [cs]

  94. [101]

    Matej Ulčar and Marko Robnik-Šikonja. 2023. Sequence-to-sequence pretraining for a less-resourced Slovenian language . doi:10.3389/frai.2023.932519

  95. [102]

    Hisao Usui and Kanako Komiya. 2023. Translation from Historical to Contemporary Japanese Using Japanese T5. In Proceedings of the Joint 3rd International Conference on Natural Language Processing for Digital Humanities and 8th International Workshop on Computational Linguistic...

  96. [103]

    Rens Van De Schoot, Jonathan De Bruin, Raoul Schram, Parisa Zahedi, Jan De Boer, Felix Weijdema, Bianca Kramer, Martijn Huijts, Maarten Hoogerwerf, Gerbrich Ferdinands, et al. 2021. An open source machine learning framework for efficient and transparent systematic reviews. Nat...

  97. [104]

    Greg Van Houdt, Carlos Mosquera, and Gonzalo Nápoles. 2020. A review on the long short-term memory model. Artificial Intelligence Review 53, 8 (2020), 5929–5955

  98. [105]

    A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017)

  99. [107]

    Haifeng Wang, Jiwei Li, Hua Wu, Eduard Hovy, and Yu Sun. 2023. Pre-trained language models and their applications. Engineering 25 (2023), 51–65

  100. [108]

    Jiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, and Jie Zhou. 2022. A Survey on Cross-Lingual Summarization. 10 (2022), 1304–1323. doi:10.1162/tacl_a_00520

  101. [109]

    Rui Wang, Xu Tan, Renqian Luo, Tao Qin, and Tie-Yan Liu. 2021. A survey on low-resource neural machine translation. arXiv preprint arXiv:2107.04239

  102. [110]

    Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018. SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (Brussels, Belgium, 2018-10), Ellen Riloff,...

  103. [111]

    Wilson Wongso, Ananto Joyoadikusumo, Brandon Scott Buana, and Derwin Suhartono. 2023. Many-to-Many Multilingual Translation Model for Languages of Indonesia. 11 (2023), 91385–91397. doi:10.1109/ACCESS.2023.3308818 Conference Name: IEEE Access

  104. [112]

    Zhen Wu and Guo Wang. 2023. A study on the corpus expansion method of neural machine translation based on reverse transcription grammar. In Proceedings of the 2023 5th Asia Pacific Information Technology Conference (New York, NY, USA, 2023-06-12) (APIT ’23). Association for Co...

  105. [113]

    Benfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang, Hongtao Xie, and Yongdong Zhang. 2020. Curriculum learning for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 6095–6104

  106. [114]

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association...

  107. [115]

    Chenyang Yang, Yike Shi, Qianou Ma, Michael Xieyang Liu, Christian Kästner, and Tongshuang Wu. 2025. What Prompts Don’t Say: Understanding and Managing Underspecification in LLM Prompts. arXiv preprint arXiv:2505.13360 (2025)

  108. [116]

    Zeynep Yirmibeşoğlu and Tunga Güngör. 2023. Morphologically Motivated Input Variations and Data Augmentation in Turkish-English Neural Machine Translation. 22, 3 (2023), 92:1–92:31. doi:10.1145/3571073

  109. [117]

    Xinyan Yu, Trina Chatterjee, Akari Asai, Junjie Hu, and Eunsol Choi. 2022. Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources. In Findings of the Association for Computational Linguistics: EMNLP 2022 , Yoav Goldberg, Zornitsa Kozare...

  110. [118]

    Zhiqiang Yu, Zhengtao Yu, Junjun Guo, Yuxin Huang, and Yonghua Wen. 2020. Efficient Low-Resource Neural Machine Translation with Reread and Feedback Mechanism. New York, NY, USA. doi:10.1145/3365244

  111. [119]

    Belén Cruz Zapata, José Luis Fernández-Alemán, Ali Idri, and Ambrosio Toval. 2015. Empirical studies on usability of mHealth apps: a systematic literature review. Journal of medical systems 39 (2015), 1–19

  112. [120]

    Isabelle A Zaugg, Anushah Hossain, and Brendan Molloy. 2022. Digitally-disadvantaged languages. Internet Policy Review 11, 2 (2022), 1–11

  113. [121]

    Chen Zhang, Xiao Liu, Jiuheng Lin, and Yansong Feng. 2024. Teaching Large Language Models an Unseen Language on the Fly. doi:10.48550/arXiv. 2402.19167 arXiv:2402.19167 [cs]

  114. [122]

    Jianyi Zhang, Xu Ji, Zhangchi Zhao, Xiali Hei, and Kim-Kwang Raymond Choo. 2023. Ethical considerations and policy implications for large language models: Guiding responsible development and deployment. arXiv preprint arXiv:2308.02678 (2023)

  115. [123]

    Jinyi Zhang, Ke Su, Haowei Li, Jiannan Mao, Ye Tian, Feng Wen, Chong Guo, and Tadahiro Matsumoto. 2024. Neural Machine Translation for Low-Resource Languages from a Chinese-centric Perspective: A Survey. 23, 6 (2024), 80:1–80:60. doi:10.1145/3665244 Overcoming Data Scarcity in...

  116. [124]

    Shiyue Zhang, Benjamin Frey, and Mohit Bansal. 2020. ChrEn: Cherokee-English Machine Translation for Endangered Language Revitalization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (Online, 2020-11), Bonnie Webber, Trevor C...

  117. [125]

    Kashmiri Prompt engineering -

  118. [126]

    Tibetan Augment monolingual data, monolingual model Paraphrase

  119. [127]

    Cherokee Back-translation -

  120. [128]

    Historical Japanese Augment monolingual data Data enrichment

  121. [129]

    Arabian, Bengali, Telugu Prompt engineering -

  122. [130]

    - Augment monolingual data Soft Contextual Data Augmentation, para- phrase

  123. [131]

    Hindi, Bengali Code mixing, multilingual models -

  124. [132]

    Basque Adaptive learning -

  125. [133]

    - Back-translation, augment monolingual data Re-ordering monolingual sentences, para- phrase

  126. [134]

    Persian Back-translation -

  127. [135]

    Turkish Back-translation -

  128. [136]

    Azerbaijani, Hindi Augment monolingual data Replace main POS tags, paraphrase

  129. [137]

    Sudanese, Javanese Family of LRLs, multilingual models -

  130. [139]

    Balinese, TobaBatak, BatakSimalungun, BatakKaro, Buginese, Javanese, Madurese, Makassarese, TorajaSadan, Sundanese, Ambonese, Acehnese, AralleTabulahan, Berik, Balantak, PakpakDairi, Bauzi, Galela, Gorontalo, Hawu, Iban, Abun, DaaKaili, LampungApi, Meyah, Minangkabau, Kupang, ...

  131. [140]

    Telugu Back-translation -

  132. [141]

    - Back-translation, augment monolingual data -

  133. [142]

    - Augment monolingual data Paraphrase

  134. [143]

    Azerbaijani, Belarusian, Glacian, Slovak Back-translation -

  135. [144]

    Assamese, Bengali, Hindi, Konkani, Maithili, Marathi, Odia, Sanskrit, Tamil, Telugu Multilingual models, family of LRLs -

  136. [145]

    Estonian, Komi, Mari, Erzya, Veps, Udmurt, Sámi, Karelian, Moksha, Livonian, Votic, Ingrian Multilingual models, family of LRLs -

  137. [146]

    Cantonese Back-translation -

  138. [147]

    Turkish, Lithuanian, Hindi, Catalan, Slovak, NorwegianBokmal, Estonian, Bengali, Latvian, Serbian, Slovenian, Tamil, Albanian, Azerbaijani, Urdu, Nepali, Macedonian, Kazakh, Georgian, Armenian, Belarusian, Esperanto, Croatian, Malayalam, Icelandic, Welsh, Telugu, Galician, Hau...

  139. [148]

    Bengali Mass translation, augment monolingual data Mass translation

  140. [149]

    Telugu, Tamil, Gujarati, Punjabi, Hindi Multilingual models -

  141. [150]

    Amharic, Catalan, Greek, Estonian, Basque, Farsi, Hindi, Croatian, Javanese, Georgian, Khmer, Kannada, Lithuanian, Latvian, Malay, Mongolian, Burmese, MinNan, Norwegian, Slovak, Slovene, Serbian, Tamil, Tagalog, Turkish Multilingual models -

  142. [151]

    UpperSorbian Back-translation -

  143. [152]

    Cantonese Multilingual models -

  144. [153]

    Khmer Back-translation -

  145. [154]

    Tamil, Sinhala Back-translation -

  146. [155]

    Bengali, Telegu, Khmer, Malay Prompt engineering -

  147. [156]

    Faroese Prompt engineering -

  148. [157]

    - Prompt engineering -

  149. [158]

    Slovenian Monolingual model -

  150. [159]

    Vallader Prompt engineering -

  151. [160]

    - Augment monolingual data Re-ordering monolingual sentences

  152. [161]

    Algerian Multilingual models -

  153. [162]

    Tibetan, Mongolian, Uyghur Multilingual models, family of LRLs -

  154. [163]

    - Back-translation -

  155. [164]

    Uyghur Monolingual model -

  156. [166]

    - Augment monolingual data Fadaee, re-ordering monolingual sentences

  157. [167]

    - Adaptive learning -

  158. [168]

    Afrikaans, Amharic, Belarusian, Welsh, Irish, Scottish, Galician, Hausa, Georgian, Kazakh, Khmer, Kyrgyz, Limburgish, Myanmar, Bokmål, Nynorsk, Occitan, Sinhala, Tajik, Turkmen, Tatar, Uighur, Uzbek, Yiddish Prompt engineering -

  159. [169]

    Turkish Adaptive learning, prompt engineering, vo- cab extension -

  160. [170]

    Zhuang Prompt engineering -

  161. [171]

    Turkish Adaptive learning, monolingual model -

  162. [172]

    - Mass translation, augment monolingual data Mass translation, Fadaee, Soft Contextual Data Augmentation, SwitchOut

  163. [173]

    - Back-translation, augment monolingual data Fadaee, SwitchOut, SMRT Simulated mul- tiple reference training, replace main POS tags

  164. [174]

    - Augment monolingual data Re-ordering monolingual sentences, para- phrase

  165. [175]

    Bengali Monolingual model -

  166. [176]

    Marathi Prompt engineering -

  167. [177]

    Indonesian Monolingual model, adaptive learning -

  168. [178]

    Data extracted for RQ4 and RQ5

    Odia Prompt engineering - Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review 41 Table 5. Data extracted for RQ4 and RQ5. The RQ4 column identifies the publisher of each study. The publication/venue column indicates the sou...

  169. [179]

    ACL ACL EACL - The European Chapter of the ACL USA 2024 Transformer mt5, flant5base

  170. [180]

    IEEE IEEE Access China 2019 RNN, Transformer, GAN QuGAN, BERT

  171. [181]

    ACL ACL EMNLP - Empirical Methods in Natural Language Processing USA 2020 RNN, Transformer BERT

  172. [182]

    ACL ACL NLP4DH - International Conference on Natural Language Processing for Digital Humanities Japan 2023 Transformer t5

  173. [183]

    ACL ACL NAACL - North American Chapter of the Association for Computational Lin- guistics Russia 2022 Transformer XLM, RoBERTa

  174. [184]

    ACL ACL - Annual Meeting of the Association for Computational Linguistics China 2019 Transformer -

  175. [185]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing India 2023 Transformer mT5

  176. [186]

    ACL ACL LREC - International Conference on Language Resources and Evaluation Spain 2020 Transformer mBERT, BERTeus

  177. [187]

    ACL ACL - Transactions of the Association for Computational Linguistics Japan 2020 Transformer XLM

  178. [188]

    De Gruyter De Gruyter - Open Computer Science USA 2019 LSTM, RNN -

  179. [189]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing Turkey 2023 Transformer, BiLSTM, RNN, LSTM -

  180. [190]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2021 Transformer -

  181. [191]

    ACL ACL EMNLP - Empirical Methods in Natural Language Processing Hong Kong 2021 Transformer mBART, GPT2

  182. [192]

    IJCAI IJCAI - International Joint Conference on Artificial Intelligence China 2021 set2sequence -

  183. [193]

    IEEE IEEE Access Indonesia 2023 Transformer mt5

  184. [194]

    ACL ACL - Proceedings of the 21st Annual Conference of the European Association for Machine Translation USA 2018 SMT, Transformer mBERT

  185. [195]

    ACL ACL EMNLP - Empirical Methods in Natural Language Processing USA 2018 Transformer -

  186. [196]

    ACL ACL - Annual Meeting of the Association for Computational Linguistics China 2020 Seq2Seq, Transformer -

  187. [197]

    ARXIV ARXIV USA 2021 Transformer -

  188. [198]

    ARXIV ARXIV India 2024 Transformer -

  189. [199]

    ARXIV ARXIV USA 2024 Transformer XLM

  190. [200]

    ARXIV ARXIV UK 2024 Transformer mBART

  191. [201]

    ARXIV ARXIV USA 2023 Transformer GPT2

  192. [202]

    ARXIV ARXIV Korea 2024 Transformer CLaude2.1, LLama27b, LLama213b, Mis- tral7b

  193. [203]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing India 2020 GAN -

  194. [204]

    ACL ACL - Transactions of the Association for Computational Linguistics UK 2018 LSTM, RNN -

  195. [205]

    ACL ACL WMT - Conference on Machine Translation USA 2020 Transformer GPT2

  196. [206]

    ACL ACL - Workshop on NLP for Similar Languages, Varieties and Dialects Sweden 2022 Transformer, BiLSTM, RNN, LSTM -

  197. [207]

    ACM ACM SoICT - Symposium on Information and Communication Technology Vietnam 2022 Transformer mBART50

  198. [208]

    IEEE IEEE - International Conference on Advances in ICT for Emerging Regions Sri Lanka 2020 LSTM, RNN -

  199. [209]

    ACL ACL - Proceedings of the Workshop on Multilingual Information Access (MIA) USA 2022 Transformer mt5, mBERT, CORA

  200. [210]

    ACL ACL LREC - International Conference on Language Resources and Evaluation Faroe Islands 2024 Transformer -

  201. [211]

    ACL ACL LREC - Joint International Conference on Computational Linguistics, Language Resources and Evaluation China 2024 Transformer -

  202. [212]

    Frontiers Frontiers in Artificial Intelligence Slovenia 2023 Transformer t5

  203. [213]

    Springer Springer - International Journal of Information Technology Switzerland 2024 Transformer -

  204. [214]

    ACM ACM APIT - Asia Pacific Information Technology Conference China 2023 - -

  205. [215]

    Springer Springer - International Journal of Information Technology Algeria 2024 Transformer -

  206. [216]

    Springer Springer - Machine Translation China 2022 Transformer -

  207. [217]

    ACL ACL - Transactions of the Association for Computational Linguistics China 2022 - -

  208. [218]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2024 Transformer -

  209. [219]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2020 Transformer -

  210. [220]

    ARXIV ARXIV China 2021 - -

  211. [221]

    IEEE ACS/IEEE - International Conference on Computer Systems and Applications Egypt 2023 Transformer, LSTM, RNN Helskini

  212. [222]

    ACL ACL LoResMT - Workshop on Technologies for Machine Translation of Low-Resource Languages USA 2024 Transformer Bloomz

  213. [223]

    ARXIV ARXIV Turkey 2024 Transformer LLaMa7b

  214. [224]

    ARXIV ARXIV China 2024 - -

  215. [225]

    ARXIV ARXIV Turkey 2024 Transformer Mistral7B, GPT2

  216. [226]

    Springer Springer - Machine Translation Ireland 2021 - -

  217. [227]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2022 - -

  218. [228]

    ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing Japan 2024 - -

  219. [229]

    IEEE IEEE ICICT4SD - International Conference on Information and Communication Tech- nology for Sustainable Development Bangladesh 2023 Transformer GPT2

  220. [230]

    IEEE IEEE INCOFT - International Conference on Futuristic Technologies India 2023 - -

  221. [231]

    IEEE IEEE ICAICTA - International Conference of Advanced Informatics: Concept, Theory and Application Indonesia 2023 Transformer mBERT, XLM, IndoBERT

  222. [232]

    Nikolov Table 6

    IEEE IEEE CCPIS - International Conference on Circuits, Power and Intelligent Systems India 2023 Transformer - 42 Josh McGiff and Nikola S. Nikolov Table 6. Data extracted for RQ7, RQ8, and RQ9. The RQ7 column describes the scarcity of data available for low-resource languages...

  223. [233]

    Kashmiri: 26kS BLEU, BERTScore, chrf++ Translation

  224. [234]

    Tibetan: 22kS BLEU Question Answering

  225. [235]

    Cherokee: 14kS BLEU Translation

  226. [236]

    - Recall, F1-score, exact match Question Answering

  227. [237]

    Hindi: 5kS, Bengali: 5kS BLEU, ROUGE, METEOR Translation

  228. [238]

    Basque: 2kS F1-score Question Answering

  229. [239]

    Turkish: 207kS BLEU Translation

  230. [240]

    Azerbaijan: 3.1mT, Hindi: 7.3mT, Uzbek: 2.2mT, Turkish: 0.8mT BLEU, METEOR Translation

  231. [241]

    - BLEU, Xtreme Indosum, TyDiQA, XPer- sona Language model, translation, summarisa- tion, conversation, question answering

  232. [242]

    - BLEU, ROUGE Translation

  233. [243]

    - SacreBLEU Translation

  234. [244]

    Telugu: 750kS BLEU, human evaluation Translation

  235. [245]

    - BLEU, Entity match rate, Success F1 Conversation

  236. [246]

    Azerbaijani: 5946S, Glacian: 10kS BLEU Translation

  237. [247]

    - Human evaluation Language model

  238. [248]

    Estonian: 6.1GB Accuracy, unlabeled attachment score Language model

  239. [249]

    Cantonese: 1.1mS SacreBLEU, hLEPOR, COMET, BERTscore, Human evaluation Translation

  240. [250]

    Iranian Persian: 1bT, Modern Greek: 1bT, Standard Arabic: 1bT, Turkish: 1bT, Lithuanian: 100mT, Hindi: 100mT, Catalan: 100mT, Slovak: 100mT, Norwegian Bokmal: 100mT, Estonian: 100mT, Bengali: 100mT, Latvian: 100mT, Serbian: 100mT, Slovenian: 100mT, Tamil: 100mT, Albanian: 100m...

  241. [251]

    Bengali: 5kS COPA Question Answering

  242. [252]

    Telugu: 70kS,Tamil: 70kS,Gujarati: 70kS,Punjabi: 70kS,Hindi: 70kS BLEU, TER Translation

  243. [253]

    Amharic: 511kS, Catalan: 788kS, Greek: 744kS, Estonian: 556kS, Basque: 647kS, Farsi: 738kS, Hindi: 666kS, Croatian: 620kS, Javanese: 622kS, Georgian: 580kS, Khmer: 579kS, Kannada: 434kS, Lithuanian: 554kS, Latvian: 587kS, Malay: 702kS, Mongolian: 629kS, Burmese: 576kS, Min-Nan...

  244. [254]

    UpperSorbian: 600kS BLEU Language model, translation

  245. [255]

    Cantonese: 35kS SacreBLEU Translation

  246. [256]

    Khmer: 70kS BLEU Translation

  247. [257]

    Tamil: 26kS, Sinhala: 26kS BLEU Translation

  248. [258]

    - Recall Question Answering

  249. [259]

    - Sentence-BERT, BLEU, chrF -

  250. [260]

    - COMET, BLEURT, chrF++ Translation

  251. [261]

    - BoolQ, Commitment Bank, COPA, Recog- nizing Textual Entailment, The Winograd Schema Challenge, classification report Language model

  252. [262]

    - BLEU, METEOR, ROUGE, CIDEr Language model

  253. [263]

    Uyghur: 20kS BLEU Translation

  254. [264]

    - BLEU, chrF, ROUGE Language model, summarisation, transla- tion

  255. [265]

    Afrikaans: 275kS, Amharic: 89kS, Belarusian: 67kS, Welsh: 289kS, Irish: 289kS, Scottish Gaelic: 16kS, Galician: 515kS, Hausa: 97kS, Georgian: 377kS, Kazakh: 79kS, Khmer: 111kS, Kyrgyz: 27kS, Limburgish: 25kS, Burmese: 24kS, Norwegian Bokmål: 142kS, Norwegian Nynorsk: 486kS, Oc...

  256. [266]

    Turkish: 273.9mT Perplexity, classification report Language model

  257. [267]

    Zhuang: 5kS Bleu, chrF Translation

  258. [268]

    Turkish: 180GB ARC-TR, TruthfulQA-TR Language model, question answering

  259. [269]

    Bengali: 26GB Perplexity Language model

  260. [270]

    - BLEU, precision, chrF, TER, Rouge Translation

  261. [271]

    - Exact match, F1-score, substring match Question Answering

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.