REVIEW 4 major objections 5 minor 2 cited by
Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims to be the first systematic review devoted specifically to data-scarcity strategies for generative language modelling in low-resource languages, synthesising 54 studies and finding transformer dominance, heavy…
desk verdict A useful first systematic map of data-scarcity techniques for low-resource generative NLG, but the headline method frequencies are muddied by counting surveys and primary studies as equivalent units. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery of the review is its screening and extraction pipeline: a search expression built from generative, model-type, and language-type terms, run across multiple digital libraries; deduplication and two screening passes, the first assisted by an active-learning prioritisation tool and the second by manual eligibility checks; reference tracking to add five further studies; and a nine-question coding scheme covering languages, methods, augmentation types, publishers, architectures, effectiveness, data scarcity, evaluation, and task. A five-item quality-assessment rubric scores each retained study and is used to weigh the literature. This pipeline converts 642 raw records into 54 analysed studies and produces the frequency distributions that carry every headline finding.
What would settle it
One concrete check would be to re-run the search while adding non-English papers and screening full text rather than title and abstract only; if the recovered set shifts the distribution of language families, the frequency of methods, or the share using BLEU, then the review's reported gaps and recommendations are artefacts of its inclusion criteria.
Extended reading notes
Core claim
The paper's central claim is that the literature on generative modelling for low-resource languages clusters around a few technical fixes and a few languages, and that this cluster can be described systematically for the first time. From 54 retained studies, the review reports that augmenting monolingual text, back-translation, multilingual training, and prompt engineering are the main strategies; that translation is by far the most common generative task; that 76% of studies use transformer-based architectures; and that BLEU is the dominant evaluation metric while human evaluation is rare. It further claims that Indo-European languages make up a disproportionate share of the modelled languages, that data reporting is too varied to compare scarcity across studies, and that consistent evaluation standards are missing. The sympathetic reading of the contribution is the structured aggregation itself: a reproducible inventory of methods, languages, architectures, and evaluation practices for a field that previously lacked one.
Load-bearing premise
The whole frequency map rests on the assumption that the 54 studies that survived the search and screening represent the relevant literature; the paper itself concedes that excluding non-English papers and screening titles and abstracts may have left relevant studies out.
Editorial extensions
If this is right
- Translation is the dominant task, so the evidence for data-scarcity methods mostly validates them for machine translation; the same methods should not be assumed to work for summarisation, dialogue, or open-ended generation.
- Monolingual augmentation, back-translation, multilingual training, and prompt engineering are the four most common strategies, with multilingual and family-of-languages approaches acting as an implicit form of data augmentation.
- The transformer dominance and the prevalence of adapting pre-trained models mean that future low-resource work will likely build on existing transformer tooling rather than revisit earlier architectures.
- Because BLEU dominates while human evaluation is rare, reported gains may reflect surface-level fluency rather than language fidelity; better evaluation is a precondition for claims about language preservation.
- The uneven distribution of modelled languages implies that 'low-resource' is not a single category; a graded measure of data, compute, and researcher availability is needed to target support where it is most lacking.
Reading between the lines
- Beyond the paper, the same data imply that the most commonly recommended fixes—back-translation and multilingual pooling—work best for languages that already have machine-translation systems or related neighbours, so the languages with least data may also be the least able to use the dominant methods.
- The authors do not pursue this, but their proposed resource-level definition could be turned into a testable index: score languages on corpus size, tokenisation support, compute access, and number of fluent NLP researchers, then check whether the index predicts which languages appear in the literature.
- A further extension would use the review's language-family observation to design a controlled experiment: train family-based models on several under-represented families and compare parameter efficiency and quality against monolingual and broad multilingual baselines.
- The review's evaluation critique suggests a practical norm: any study claiming a low-resource language model is faithful to the language should report human evaluation alongside automatic metrics.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a PRISMA-style systematic review of 54 studies on strategies to overcome data scarcity in generative language modelling for low-resource languages (LRLs). The review identifies and categorizes technical approaches (monolingual augmentation, back-translation, multilingual modelling, prompt engineering, etc.), analyzes language-family representation, model architectures, and evaluation practices, and reports research gaps and recommendations. The central claims are that transformer-based models dominate, that a small subset of LRLs accounts for most research, and that evaluation is inconsistent across studies.
Significance. If the quantitative map is robust, the review is a useful contribution: it consolidates scattered evidence, provides a structured taxonomy of data-scarcity methods, and gives concrete recommendations for future research, including language-family-centric modelling and more nuanced definitions of 'low-resource'. The paper is transparent about many limitations, ships detailed extraction tables (Tables 2-6), and uses a PICO framework and quality assessment, which are strengths. The falsifiable field-level claims, such as the dominance of transformers and the concentration of studies on a few languages, are important for researchers and funding bodies, and the discussion of geopolitical factors in language selection is a valuable framing.
major comments (4)
- [Section 3.3, Table 3; Section 2.4; Section 5] The central frequency map for RQ2/RQ3 is built on a unit-of-analysis problem: the corpus mixes primary empirical studies with survey/review papers, yet each paper is assigned to technique categories 'based on the techniques it discussed or implemented'. Table 3 explicitly lists surveys [29, 86, 109, 123] under 'Augment Monolingual Data' and [86, 108] under 'Back-Translation', even though these papers review the literature rather than deploy the methods. Consequently, the reported 26% for monolingual augmentation and 24% for back-translation (Figure 7) conflate description with implementation. Because these percentages underpin the review's main empirical conclusions, the authors should re-run the analysis on primary studies only (e.g., QA2=Yes) and report both raw and sensitivity-adjusted frequencies, or clearly separate 'discussed' from 'implemented' categories in the figures and tables.
- [Section 2.2, Table 1] The search strategy is not fully reproducible. The text states that 'Table 1 details the search strings that were used for the various digital repositories', but Table 1 shows only four string components ('generative', 'language model OR text model', 'low resource language OR minority language OR endangered language', 'exclusion: classification') with no per-database Boolean syntax, no field restrictions, and no search dates. The PRISMA flow diagram (Figure 1) reports 642 initial records, but without the exact queries and access dates, readers cannot verify the coverage or repeat the search. The authors should provide the full per-database search strings and search dates in an appendix, and should state which metadata fields (title/abstract/keywords) were matched.
- [Section 3.9, Section 3.10, Appendix A (Tables 2, 6)] There are internal inconsistencies between the results text and the extracted-data tables that prevent verification of the reported percentages. Reference [106] is cited in Section 3.9 as using perplexity/chrF/METEOR and in Section 3.10 as a question-answering study, but [106] does not appear in Table 2 (the list of 54 included studies) or in any extraction row of Tables 4-6. Similarly, reference [72] is included in the BLEU distribution list in Section 3.9 but is absent from Table 2 and Table 6. Either the tables omit included studies or the citations are wrong; in both cases, the percentages (e.g., 61% BLEU, 7% for perplexity) cannot be audited. The authors must reconcile the reference list, the included-study tables, and all in-text citation groupings.
- [Section 2.3, Figure 1, Section 5] The single-reviewer screening process is a load-bearing limitation for the claim that the review represents the relevant literature. The paper openly acknowledges this in Section 5, but Figure 1 and Section 2.3 do not report any reliability checks (e.g., a second reviewer on a random subset, or a comparison of ASReview's active-learning ranking against a manual gold standard). Since the search strings are also incompletely reported (see above) and the ASReview screening excludes papers based on title/abstract, the authors should either add a sensitivity analysis (e.g., re-screening a random sample by a second reviewer) or explicitly state the absence of any inter-rater validation and discuss how this could bias the frequency estimates.
minor comments (5)
- [Section 2.2] Typo: 'Scoupus' should be 'Scopus'.
- [Section 3.2] The statement that Turkish appears in 9% of papers cites [1, 3, 13, 25, 100, 116], but Table 4 assigns only [13, 25, 100, 116] (and [1]) to Turkish; [3] is listed for Bengali, Telugu, Khmer, and Malay. Please correct the citation grouping or the language assignment.
- [Table 1] The caption 'Search string composition' does not match the text's claim that the search strings themselves are detailed; consider renaming the table or providing the full strings as described in the major comment.
- [Section 3.8, RQ7] The statement 'a count of 390k sentences is the average count' lacks a definition of which set the average is over (the languages with a single data source in Figure 15). Specify the denominator and whether the average is per language or per study.
- [Section 3.1] The PRISMA flow diagram (Figure 1) would be clearer if the 'Excluded: 52' and 'Excluded: 486' boxes also stated the reasons (deduplication; EC1, EC3, EC4) on the diagram itself, as recommended in PRISMA 2020 templates.
Circularity Check
No significant circularity: the review synthesizes 54 external studies; its only self-citation is incidental and none of its findings are forced by construction.
full rationale
This paper is a systematic review, not a derivation or prediction exercise. Its central claims (transformer dominance, concentration on a few LRLs, inconsistent evaluation, most-used data-scarcity methods) are frequency summaries extracted from 54 included studies. There are no fitted parameters, no equations whose outputs coincide with their inputs, and no uniqueness theorem imported from the authors' prior work. The only self-citation, [53], appears in the introduction as an example of a transformer-based classification application ('with the transformer architecture enhancing NLP tasks such as machine translation [77], classification [53] and named entity recognition [6]'). It is not load-bearing for the review's inclusion criteria, search strategy, research questions, or conclusions. The skeptical concern that RQ2/RQ3 counts mix primary studies with survey papers (e.g., [29], [86], [108], [109], [123] are categorized under technical methods) is a methodological unit-of-analysis and validity issue, not circularity: the counts describe what the included corpus says, and the corpus is not defined in terms of the counts. Likewise, the acknowledged limitations (non-English papers excluded, single-reviewer title/abstract screening, reliance on ASReview) affect recall and representativeness but do not make any result equivalent to its input by definition. The novelty claim of being 'the first systematic review focused specifically' on this topic rests on the reported search and screening process rather than on a self-referential construction. Under the stated rules, this warrants a non-finding: score 0, no circular steps identified.
Assumptions & free parameters
assumptions (3)
- domain assumption The search and screening process (PRISMA plus ASReview) captured a representative subset of the relevant literature, and the 54 included studies are sufficient to support the reported trends.
- domain assumption The manual categorization of technical methods, languages, and language families accurately reflects the content of the 54 studies.
- domain assumption The quality assessment scoring (QA1-QA5) is a valid proxy for study relevance and quality in this field.
Cite this review
Pith. "Pith review of Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review." pith.science (2026). https://pith.science/paper/4VUCNUAB
@misc{pith2026250504531,
author = {Pith},
title = {Pith review of: Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review},
year = {2026},
howpublished = {\url{https://pith.science/paper/4VUCNUAB}},
note = {Machine review of arXiv:2505.04531}
}
read the original abstract
Generative language modelling has surged in popularity with the emergence of services such as ChatGPT and Google Gemini. While these models have demonstrated transformative potential in productivity and communication, they overwhelmingly cater to high-resource languages like English. This has amplified concerns over linguistic inequality in natural language processing (NLP). This paper presents the first systematic review focused specifically on strategies to address data scarcity in generative language modelling for low-resource languages (LRL). Drawing from 54 studies, we identify, categorise and evaluate technical approaches, including monolingual data augmentation, back-translation, multilingual training, and prompt engineering, across generative tasks. We also analyse trends in architecture choices, language family representation, and evaluation methods. Our findings highlight a strong reliance on transformer-based models, a concentration on a small subset of LRLs, and a lack of consistent evaluation across studies. We conclude with recommendations for extending these methods to a wider range of LRLs and outline open challenges in building equitable generative language systems. Ultimately, this review aims to support researchers and developers in building inclusive AI tools for underrepresented languages, a necessary step toward empowering LRL speakers and the preservation of linguistic diversity in a world increasingly shaped by large-scale language technologies.
Figures
Figures from the paper (14 more)
Forward citations
Cited by 2 Pith papers
-
Representational Equality in Cross-country Value Simulation: A Systematic Analysis of Large Language Models
LLM-based value simulation is systematically more accurate for wealthy, high-governance, individualist countries, and common interventions rarely fix the imbalance.
-
Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models
Synthetic data for low-resource spoken language models creates a Stability-Expressivity Gap that DGSA and TDSC self-alignment close, enabling SOTA Thai TTS and first Lao zero-shot voice cloning.
Reference graph
Works this paper leans on
-
[106]
Tam Minh Vo and Khiem Vinh Tran. 2023. Generative Pre-trained Transformer for Vietnamese Community-based COVID-19 Question Answering. doi:10.48550/arXiv.2310.14602 arXiv:2310.14602 [cs]
work page Pith review arXiv doi:10.48550/arxiv.2310.14602 2023
-
[72]
Hong-Hai Phan-Vu, Van-Nam Nguyen, Viet-Trung Tran, and Phan-Thuan Do. 2017. Towards State-of-the-art English-Vietnamese Neural Machine Translation. In Proceedings of the 8th International Symposium on Information and Communication Technology (New York, NY, USA, 2017-12-07) (SoICT ’17). Association for Computing Machinery, 120–126. doi:10.1145/3155133.3155205
-
[1]
Emre Can Acikgoz, Mete Erdogan, and Deniz Yuret. 2024. Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking. doi:10.48550/arXiv.2405.04685 arXiv:2405.04685 [cs]
-
[2]
Parul Agarwal, Aisha Asif, Shantipriya Parida, Sambit Sekhar, Satya Ranjan Dash, and Subhadarshi Panda. 2023. Generative Chatbot Adaptation for Odia Language: A Critical Evaluation . doi:10.1109/CCPIS59145.2023.10291329 34 Josh McGiff and Nikola S. Nikolov
arXiv 2023
-
[3]
Sumit Agarwal, Suraj Tripathi, Teruko Mitamura, and Carolyn Penstein Rose. 2022. Zero-shot cross-lingual open domain question answering. In Proceedings of the Workshop on Multilingual Information Access (MIA) (Seattle, USA, 2022-07), Akari Asai, Eunsol Choi, Jonathan H. Clark, Junjie Hu, Chia-Hsuan Lee, Jungo Kasai, Shayne Longpre, Ikuya Yamada, and Rui Z...
-
[4]
Benyamin Ahmadnia and Bonnie J. Dorr. 2019. Augmenting Neural Machine Translation through Round-Trip Training Approach. 9, 1 (2019), 268–278. doi:10.1515/comp-2019-0019 Publisher: De Gruyter Open Access
-
[5]
Talal Almutiri and Farrukh Nadeem. 2022. Markov models applications in natural language processing: a survey. Int. J. Inf. Technol. Comput. Sci 2 (2022), 1–16
2022
-
[6]
Mikhail Arkhipov, Maria Trofimova, Yurii Kuratov, and Alexey Sorokin. 2019. Tuning multilingual transformers for language-specific named entity recognition. In Proceedings of the 7th Workshop on Balto-Slavic Natural Language Processing . 89–93
2019
Show all 269 references
-
[7]
Baligh Babaali, Mohammed Salem, and Nawaf R. Alharbe. 2024. Breaking language barriers with ChatGPT: enhancing low-resource machine translation between algerian arabic and MSA. (2024). doi:10.1007/s41870-024-01926-7
2024 doi
-
[8]
Oliver Bendel and Dalil Jabou. 2024. @llegra: a chatbot for Vallader. 16, 4 (2024), 2035–2045. doi:10.1007/s41870-024-01779-0
2024 doi
-
[9]
Tucker Berckmann and Berkan Hiziroglu. 2020. Low-Resource Translation as Language Modeling. InProceedings of the Fifth Conference on Machine Translation (Online, 2020-11), Loïc Barrault, Ondřej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussà, Christian Federmann,...
2020
-
[10]
Rajat Subhra Bhowmick, Isha Ganguli, Ananya Paul, Jayanta Paul, and Jaya Sil. 2023. Improving Indic code-mixed to monolingual translation using Mixed Script Augmentation, Generation & Transfer Learning. (2023). doi:10.1145/3606695 Just Accepted
2023 doi
-
[11]
Steven Bird. 2022. Local Languages, Third Spaces, and other High-Resource Scenarios. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association ...
2022 doi
-
[12]
Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Langua...
2021
- [13]
-
[14]
Sayan Chatterjee, Ching Louis Liu, Gareth Rowland, and Tim Hogarth. 2024. The Impact of AI Tool on Engineering at ANZ Bank An Empirical Study on GitHub Copilot within Corporate Environment. arXiv preprint arXiv:2402.05636 (2024)
2024 arXiv
-
[15]
Sandipan Dandapat and Christian Federmann. 2018. Iterative Data Augmentation for Neural Machine Translation: a Low Resource Case Study for English-Telugu. In Proceedings of the 21st Annual Conference of the European Association for Machine Translation (Alicante, Spain, 2018-05...
2018
-
[16]
Lorenzo De Mattei, Michele Cafagna, Felice Dell’Orletta, Malvina Nissim, Marco Guerini, Aptus AI, and Fondazione Bruno Kessler. 2020. GePpeTto Carves Italian into a Language Model. Computational Linguistics CLiC-it 2020 (2020), 136
2020
-
[17]
Department of Public Expenditure, Infrastructure, Public Service Reform and Digitalisation. 2025. Guidelines for the Responsible Use of AI in the Public Service. Available at: https://www.gov.ie/en/department-of-public-expenditure-infrastructure-public-service-reform-and- digi...
2025
-
[18]
Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[19]
Chenhe Dong, Yinghui Li, Haifan Gong, Miaoxin Chen, Junxin Li, Ying Shen, and Min Yang. 2022. A survey of natural language generation. Comput. Surveys 55, 8 (2022), 1–38
2022
- [20]
-
[21]
Matthew S Dryer. 2007. Word order. Language typology and syntactic description 1, 61-131 (2007), 1–1
2007
-
[22]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[23]
Fei Gao, Jinhua Zhu, Lijun Wu, Yingce Xia, Tao Qin, Xueqi Cheng, Wengang Zhou, and Tie-Yan Liu. 2019. Soft Contextual Data Augmentation for Neural Machine Translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Florence, Italy, ...
2019 doi
-
[24]
Silin Gao, Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020. Paraphrase Augmented Task-Oriented Dialog Generation . Online. doi:10.18653/v1/2020.acl- main.60 Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review 35
2020 doi
-
[25]
Daniela Gerz, Ivan Vulić, Edoardo Ponti, Jason Naradowsky, Roi Reichart, and Anna Korhonen. 2018. Language Modeling for Morphologically Rich Languages: Character-Aware Modeling for Word-Level Prediction. 6 (2018), 451–465. doi:10.1162/tacl_a_00032
2018 doi
-
[26]
The Guardian. 2024. Brat summer: is the long era of clean living finally over? https://www.theguardian.com/lifeandstyle/article/2024/jul/16/brat- summer-is-the-long-era-of-clean-living-finally-over. Accessed: 2025-06-19
2024
-
[27]
Ping Guo, Yubing Ren, Yue Hu, Yunpeng Li, Jiarui Zhang, Xingsheng Zhang, and Heyan Huang. 2024. Teaching Large Language Models to Translate on Low-resource Languages with Textbook Prompting. In Proceedings of the 2024 Joint International Conference on Computational Linguistics...
2024
-
[28]
Zhu, Guandan Chen, Kaibo Zhang, Boxing Chen, and Fei Huang
Zilu Guo, Zhongqiang Huang, Kenny Q. Zhu, Guandan Chen, Kaibo Zhang, Boxing Chen, and Fei Huang. 2021. Automatically Paraphrasing via Sentence Reconstruction and Round-trip Translation. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (...
2021 doi
-
[29]
Rejwanul Haque, Chao-Hong Liu, and Andy Way. 2021. Recent advances of low-resource neural machine translation. 35, 4 (2021), 451–474. doi:10.1007/s10590-021-09281-1
2021 doi
-
[30]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. [n. d.]. Training Compute-Optimal Large Language Models. ([n. d.])
-
[31]
Heidi Ahmed Holiel, Nancy Mohamed, Arwa Ahmed, and Walaa Medhat. 2023. English-Arabic Text Translation and Abstractive Summarization Using Transformers. doi:10.1109/AICCSA59173.2023.10479257
2023
- [32]
-
[33]
James Hutson, Pace Ellsworth, and Matt Ellsworth. 2024. Preserving linguistic diversity in the digital age: a scalable model for cultural heritage continuity. Journal of Contemporary Language Research 3, 1 (2024)
2024
-
[34]
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce C...
2020 doi
-
[35]
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020)
2020 arXiv
-
[36]
Cherry Khosla and Baljit Singh Saini. 2020. Enhancing performance of deep learning models with different data augmentation techniques: A survey. In 2020 International Conference on Intelligent Engineering and Management (ICIEM) . IEEE, 79–85
2020
-
[37]
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale. 2023. Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback. arXiv preprint arXiv:2303.05453 (2023)
2023 arXiv
-
[38]
Piotr Kłosowski. 2018. Deep learning for natural language processing and language modelling. In 2018 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA). IEEE, 223–228
2018
-
[39]
Viet Dac Lai, Nghia Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning. InFindings of the Association for Computational L...
2023 doi
-
[40]
Siobhán Ní Laoire. 2016. Irish-English Code-switching: a Sociolinguistic Perspective. In Sociolinguistics in Ireland, Raymond Hickey (Ed.). Palgrave Macmillan UK, 81–106. doi:10.1057/9781137453471_4
2016 doi
-
[41]
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2023. Bloom: A 176b-parameter open-access multilingual language model. (2023)
2023
-
[42]
Julie Anne Legate. 1999. The morphosyntax of Irish agreement. MIT working papers in linguistics 33 (1999), 219–240
1999
-
[43]
Bin Li, Yixuan Weng, Bin Sun, and Shutao Li. 2022. A Multi-tasking and Multi-stage Chinese Minority Pre-trained Language Model. In Machine Translation (Singapore, 2022), Tong Xiao and Juan Pino (Eds.). Springer Nature, 93–105. doi:10.1007/978-981-19-7960-6_10
2022 doi
-
[44]
Zhuang Li, Levon Haroutunian, Raj Tumuluri, Philip Cohen, and Reza Haf. 2024. Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach. In Findings of the Association for Computational Linguistics: EACL 2024 (St. Julian’s,...
2024
-
[45]
Ying Lian, Huiting Tang, Mengting Xiang, and Xuefan Dong. 2024. Public attitudes and sentiments toward ChatGPT in China: A text mining analysis based on social media. Technology in Society 76 (2024), 102442
2024
-
[46]
Godfrey Lienhardt, John Middleton, and David Tait. 1958. The Western Dinka. Tribes without rulers: Studies in african segmentary systems (1958), 97–135
1958
-
[47]
Evelyn Kai-Yan Liu. 2022. Low-Resource Neural Machine Translation: A Case Study of Cantonese. In Proceedings of the Ninth Workshop on NLP for Similar Languages, Varieties and Dialects (Gyeongju, Republic of Korea, 2022-10), Yves Scherrer, Tommi Jauhiainen, Nikola Ljubešić, Pre...
2022
-
[48]
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics 8 (2020), 726–742. 36 Jos...
2020
-
[49]
Alexandre Magueresse, Vincent Carles, and Evan Heetderks. 2020. Low-resource languages: A review of past work and future challenges. arXiv preprint arXiv:2006.07264 (2020)
2020 arXiv
-
[50]
Mieradilijiang Maimaiti, Yang Liu, Huanbo Luan, Zegao Pan, and Maosong Sun. 2021. Improving Data Augmentation for Low-Resource NMT Guided by POS-Tagging and Paraphrase Embedding. 20, 6 (2021), 107:1–107:21. doi:10.1145/3464427
2021 doi
-
[51]
Zhuoyuan Mao and Yen Yu. 2024. Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages. In Proceedings of the Seventh Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2024) (Bangkok, Thail...
2024
-
[52]
Benjamin Marie and Atsushi Fujita. 2020. Synthesizing Parallel Data of User-Generated Texts with Zero-Shot Neural Machine Translation. 8 (2020), 710–725. doi:10.1162/tacl_a_00341
2020 doi
-
[53]
Josh McGiff and Nikola S Nikolov. 2024. Bridging the gap in online hate speech detection: A comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter. Applied and Computational Engineering 64 (2024), 63–68
2024
-
[54]
Chenggang Mi, Shaoliang Xie, and Yi Fan. 2024. Multi-granularity Knowledge Sharing in Low-resource Neural Machine Translation. 23, 2 (2024), 31:1–31:19. doi:10.1145/3639930
2024 doi
-
[55]
Ibomoiye Domor Mienye, Theo G Swart, and George Obaido. 2024. Recurrent neural networks: A comprehensive review of architectures, variants, and applications. Information 15, 9 (2024), 517
2024
-
[56]
Agnieszka Mikołajczyk and Michał Grochowski. 2018. Data augmentation for improving deep learning in image classification problem. In 2018 international interdisciplinary PhD workshop (IIPhDW) . IEEE, 117–122
2018
-
[57]
Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel. 2023. Scaling Data-Constrained Language Models. In Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globers...
2023
-
[58]
Martin Müller and Florian Laurent. 2022. Cedille: A large autoregressive french language model. arXiv preprint arXiv:2202.03371 (2022)
2022 arXiv
-
[59]
Alhassan Mumuni and Fuseini Mumuni. 2022. Data augmentation: A comprehensive survey of modern approaches. Array 16 (2022), 100258
2022
-
[60]
Dinesh Kumar Nanduri and Elizabeth M Bonsignore. 2023. Revitalizing Endangered Languages: AI-powered language learning as a catalyst for language appreciation. arXiv preprint arXiv:2304.09394 (2023)
2023
-
[61]
Reuben Ng and Ting Yu Joanne Chow. 2024. Powerful tool or too powerful? Early public discourse about ChatGPT across 4 million tweets. Plos one 19, 3 (2024), e0296882
2024
-
[62]
Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, and Monojit Choudhury. 2024. The Zeno’s Paradox of ‘Low- Resource’ Languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal...
2024
-
[63]
L. N. A. S. H. Nissanka, B. H. R. Pushpananda, and A. R. Weerasinghe. 2020. Exploring Neural Machine Translation for Sinhala-Tamil Languages Pair. In 2020 20th International Conference on Advances in ICT for Emerging Regions (ICTer) (2020-11). 202–207. doi:10.1109/ICTer51097.2...
2020
-
[64]
Mitodru Niyogi and Arnab Bhattacharya. 2024. Paramanu: A Family of Novel Efficient Indic Generative Foundation Language Models. doi:10. 48550/arXiv.2401.18034 arXiv:2401.18034 [cs]
2024 doi
-
[65]
Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages. In Proceedings of the 1st Workshop on Multilingual Representation Learning . 116–126
2021
-
[66]
Arantxa Otegi, Aitor Agirre, Jon Ander Campos, Aitor Soroa, and Eneko Agirre. 2020. Conversational Question Answering in Low Resource Scenarios: A Dataset and Case Study for Basque. In Proceedings of the Twelfth Language Resources and Evaluation Conference (Marseille, France, ...
2020
-
[67]
Matthew J Page, David Moher, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al. 2021. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic...
2021
-
[68]
ChaeHun Park, Koanho Lee, Hyesu Lim, Jaeseok Kim, Junmo Park, Yu-Jung Heo, Du-Seong Chang, and Jaegul Choo. 2024. Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering. In Findings of the Association for Computational Linguisti...
2024 doi
-
[69]
Bruno Peixoto, Rafael Pinto, Miguel Melo, Luciana Cabral, and Maximino Bessa. 2021. Immersive virtual reality for foreign language education: A PRISMA systematic review. IEEE Access 9 (2021), 48952–48962
2021
- [70]
-
[71]
Hanh Pham Van and Huong Le Thanh. 2022. Improving Khmer-Vietnamese Machine Translation with Data Augmentation methods. InProceedings of the 11th International Symposium on Information and Communication Technology (New York, NY, USA, 2022-12-01) (SoICT ’22). Association for Ove...
2022
-
[73]
Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. In Proceedings of the Third Conference on Machine Translation: Research Papers (Brussels, Belgium, 2018-10), Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Hu...
2018
-
[74]
QuantumBlack, AI by McKinsey. 2025. The state of AI: How organizations are rewiring to capture value. Online report, McKinsey & Company. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai Published March 12, 2025; accessed June 18, 2025
2025
-
[75]
Alec Radford. 2018. Improving language understanding by generative pre-training. (2018)
2018
-
[76]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[77]
Alessandro Raganato and Jörg Tiedemann. 2018. An analysis of encoder representations in transformer-based machine translation. In Proceedings of the 2018 EMNLP workshop BlackboxNLP: analyzing and interpreting neural networks for NLP . 287–297
2018
-
[78]
Georg Rehm and Andy Way. 2023. European language equality: A strategic agenda for digital language equality . Springer Nature
2023
-
[79]
Muhammad Razif Rizqullah, Ayu Purwarianti, and Alham Fikri Aji. 2023. QASiNa: Religious Domain Question Answering Using Sirah Nabawiyah . doi:10.1109/ICAICTA59291.2023.10390123
2023
-
[80]
Brian Roark, Murat Saraclar, and Michael Collins. 2007. Discriminative n-gram language modeling. Computer Speech & Language 21, 2 (2007), 373–392
2007
-
[81]
Lekhraj Saini and Deepti Vidhyarthi. 2023. Bidirectional English-Marathi Translation using Pretrained Models: A Comparative Study of Different Pre-Trained Models. doi:10.1109/INCOFT60753.2023.10425770
2023
-
[82]
Shahidul Salim, Hasan Murad, Dola Das, and Faisal Ahmed
Md. Shahidul Salim, Hasan Murad, Dola Das, and Faisal Ahmed. 2023. BanglaGPT: A Generative Pretrained Transformer-Based Model for Bangla Language. doi:10.1109/ICICT4SD59951.2023.10303383
2023
-
[83]
Lamyae Sardi, Ali Idri, and José Luis Fernández-Alemán. 2017. A systematic review of gamification in e-Health. Journal of biomedical informatics 71 (2017), 31–48
2017
-
[84]
Barbara Scalvini and Iben Nyholm Debess. 2024. Evaluating the Potential of Language-family-specific Generative Models for Low-resource Data Augmentation: A Faroese Case Study. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Reso...
2024
- [85]
-
[86]
Shumin Shi, Xing Wu, Rihai Su, and Heyan Huang. 2022. Low-resource Neural Machine Translation: Methods and Trends. 21, 5 (2022), 103:1–103:22. doi:10.1145/3524300
2022 doi
-
[87]
Connor Shorten and Taghi M Khoshgoftaar. 2019. A survey on image data augmentation for deep learning. Journal of big data 6, 1 (2019), 1–48
2019
-
[88]
Connor Shorten, Taghi M Khoshgoftaar, and Borko Furht. 2021. Text data augmentation for deep learning. Journal of big Data 8, 1 (2021), 101
2021
-
[89]
Antoine Simoulin and Benoit Crabbé. 2021. Un modèle Transformer Génératif Pré-entrainé pour le _ français. In Traitement Automatique des Langues Naturelles. ATALA, 246–255
2021
-
[90]
Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O’Mahony, et al. 2024. Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning. In Proceedings of the 62n...
2024
-
[91]
Nikita Sorokin, Dmitry Abulkhanov, Irina Piontkovskaya, and Valentin Malykh. 2022. Ask Me Anything in Your Native Language. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Seattle...
2022 doi
-
[92]
Patricia W Stone. 2002. Popping the (PICO) question in research and evidence-based practice. Applied Nursing Research 15, 3 (2002), 197–198
2002
-
[93]
Shang-Yu Su, Chao-Wei Huang, and Yun-Nung Chen. 2019. Dual Supervised Learning for Natural Language Understanding and Generation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). ...
2019 doi
-
[94]
Yuan Sun, Chaofan Chen, Tianci Xia, and Xiaobing Zhao. 2019. QuGAN: Quasi Generative Adversarial Network for Tibetan Question Answering Corpus Generation. doi:10.1109/ACCESS.2019.2934581
2019
-
[95]
Ashwani Tanwar and Prasenjit Majumder. 2020. Translating Morphologically Rich Indian Languages under Zero-Resource Conditions. 19, 6 (2020), 85:1–85:15. doi:10.1145/3407912
2020 doi
-
[96]
Luke Taylor and Geoff Nitschke. 2018. Improving deep learning with generic data augmentation. In 2018 IEEE symposium series on computational intelligence (SSCI). IEEE, 1542–1547
2018
-
[97]
NLLB Team et al. 2024. Scaling neural machine translation to 200 languages. Nature 630, 8018 (2024), 841. 38 Josh McGiff and Nikola S. Nikolov
2024
-
[98]
The World Bank. 2024. Research and development expenditure (% of GDP) - South Sudan. https://data.worldbank.org/indicator/GB.XPD.RSDV. GD.ZS?locations=ZS Accessed: 2025-06-09
2024
-
[99]
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023. Large language models in medicine. Nature medicine 29, 8 (2023), 1930–1940
2023
- [100]
-
[101]
Matej Ulčar and Marko Robnik-Šikonja. 2023. Sequence-to-sequence pretraining for a less-resourced Slovenian language . doi:10.3389/frai.2023.932519
2023
-
[102]
Hisao Usui and Kanako Komiya. 2023. Translation from Historical to Contemporary Japanese Using Japanese T5. In Proceedings of the Joint 3rd International Conference on Natural Language Processing for Digital Humanities and 8th International Workshop on Computational Linguistic...
2023
-
[103]
Rens Van De Schoot, Jonathan De Bruin, Raoul Schram, Parisa Zahedi, Jan De Boer, Felix Weijdema, Bianca Kramer, Martijn Huijts, Maarten Hoogerwerf, Gerbrich Ferdinands, et al. 2021. An open source machine learning framework for efficient and transparent systematic reviews. Nat...
2021
-
[104]
Greg Van Houdt, Carlos Mosquera, and Gonzalo Nápoles. 2020. A review on the long short-term memory model. Artificial Intelligence Review 53, 8 (2020), 5929–5955
2020
-
[105]
A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017)
2017
-
[107]
Haifeng Wang, Jiwei Li, Hua Wu, Eduard Hovy, and Yu Sun. 2023. Pre-trained language models and their applications. Engineering 25 (2023), 51–65
2023
-
[108]
Jiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, and Jie Zhou. 2022. A Survey on Cross-Lingual Summarization. 10 (2022), 1304–1323. doi:10.1162/tacl_a_00520
2022 doi
-
[109]
Rui Wang, Xu Tan, Renqian Luo, Tao Qin, and Tie-Yan Liu. 2021. A survey on low-resource neural machine translation. arXiv preprint arXiv:2107.04239
2021 arXiv
-
[110]
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018. SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (Brussels, Belgium, 2018-10), Ellen Riloff,...
2018 doi
-
[111]
Wilson Wongso, Ananto Joyoadikusumo, Brandon Scott Buana, and Derwin Suhartono. 2023. Many-to-Many Multilingual Translation Model for Languages of Indonesia. 11 (2023), 91385–91397. doi:10.1109/ACCESS.2023.3308818 Conference Name: IEEE Access
2023
-
[112]
Zhen Wu and Guo Wang. 2023. A study on the corpus expansion method of neural machine translation based on reverse transcription grammar. In Proceedings of the 2023 5th Asia Pacific Information Technology Conference (New York, NY, USA, 2023-06-12) (APIT ’23). Association for Co...
2023
-
[113]
Benfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang, Hongtao Xie, and Yongdong Zhang. 2020. Curriculum learning for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 6095–6104
2020
-
[114]
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association...
2021
-
[115]
Chenyang Yang, Yike Shi, Qianou Ma, Michael Xieyang Liu, Christian Kästner, and Tongshuang Wu. 2025. What Prompts Don’t Say: Understanding and Managing Underspecification in LLM Prompts. arXiv preprint arXiv:2505.13360 (2025)
2025 arXiv
-
[116]
Zeynep Yirmibeşoğlu and Tunga Güngör. 2023. Morphologically Motivated Input Variations and Data Augmentation in Turkish-English Neural Machine Translation. 22, 3 (2023), 92:1–92:31. doi:10.1145/3571073
2023 doi
-
[117]
Xinyan Yu, Trina Chatterjee, Akari Asai, Junjie Hu, and Eunsol Choi. 2022. Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources. In Findings of the Association for Computational Linguistics: EMNLP 2022 , Yoav Goldberg, Zornitsa Kozare...
2022 doi
-
[118]
Zhiqiang Yu, Zhengtao Yu, Junjun Guo, Yuxin Huang, and Yonghua Wen. 2020. Efficient Low-Resource Neural Machine Translation with Reread and Feedback Mechanism. New York, NY, USA. doi:10.1145/3365244
2020 doi
-
[119]
Belén Cruz Zapata, José Luis Fernández-Alemán, Ali Idri, and Ambrosio Toval. 2015. Empirical studies on usability of mHealth apps: a systematic literature review. Journal of medical systems 39 (2015), 1–19
2015
-
[120]
Isabelle A Zaugg, Anushah Hossain, and Brendan Molloy. 2022. Digitally-disadvantaged languages. Internet Policy Review 11, 2 (2022), 1–11
2022
- [121]
-
[122]
Jianyi Zhang, Xu Ji, Zhangchi Zhao, Xiali Hei, and Kim-Kwang Raymond Choo. 2023. Ethical considerations and policy implications for large language models: Guiding responsible development and deployment. arXiv preprint arXiv:2308.02678 (2023)
2023 arXiv
-
[123]
Jinyi Zhang, Ke Su, Haowei Li, Jiannan Mao, Ye Tian, Feng Wen, Chong Guo, and Tadahiro Matsumoto. 2024. Neural Machine Translation for Low-Resource Languages from a Chinese-centric Perspective: A Survey. 23, 6 (2024), 80:1–80:60. doi:10.1145/3665244 Overcoming Data Scarcity in...
2024 doi
-
[124]
Shiyue Zhang, Benjamin Frey, and Mohit Bansal. 2020. ChrEn: Cherokee-English Machine Translation for Endangered Language Revitalization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (Online, 2020-11), Bonnie Webber, Trevor C...
2020 doi
-
[125]
Kashmiri Prompt engineering -
-
[126]
Tibetan Augment monolingual data, monolingual model Paraphrase
-
[127]
Cherokee Back-translation -
-
[128]
Historical Japanese Augment monolingual data Data enrichment
-
[129]
Arabian, Bengali, Telugu Prompt engineering -
-
[130]
- Augment monolingual data Soft Contextual Data Augmentation, para- phrase
-
[131]
Hindi, Bengali Code mixing, multilingual models -
-
[132]
Basque Adaptive learning -
-
[133]
- Back-translation, augment monolingual data Re-ordering monolingual sentences, para- phrase
-
[134]
Persian Back-translation -
-
[135]
Turkish Back-translation -
-
[136]
Azerbaijani, Hindi Augment monolingual data Replace main POS tags, paraphrase
-
[137]
Sudanese, Javanese Family of LRLs, multilingual models -
-
[139]
Balinese, TobaBatak, BatakSimalungun, BatakKaro, Buginese, Javanese, Madurese, Makassarese, TorajaSadan, Sundanese, Ambonese, Acehnese, AralleTabulahan, Berik, Balantak, PakpakDairi, Bauzi, Galela, Gorontalo, Hawu, Iban, Abun, DaaKaili, LampungApi, Meyah, Minangkabau, Kupang, ...
-
[140]
Telugu Back-translation -
-
[141]
- Back-translation, augment monolingual data -
-
[142]
- Augment monolingual data Paraphrase
-
[143]
Azerbaijani, Belarusian, Glacian, Slovak Back-translation -
-
[144]
Assamese, Bengali, Hindi, Konkani, Maithili, Marathi, Odia, Sanskrit, Tamil, Telugu Multilingual models, family of LRLs -
-
[145]
Estonian, Komi, Mari, Erzya, Veps, Udmurt, Sámi, Karelian, Moksha, Livonian, Votic, Ingrian Multilingual models, family of LRLs -
-
[146]
Cantonese Back-translation -
-
[147]
Turkish, Lithuanian, Hindi, Catalan, Slovak, NorwegianBokmal, Estonian, Bengali, Latvian, Serbian, Slovenian, Tamil, Albanian, Azerbaijani, Urdu, Nepali, Macedonian, Kazakh, Georgian, Armenian, Belarusian, Esperanto, Croatian, Malayalam, Icelandic, Welsh, Telugu, Galician, Hau...
-
[148]
Bengali Mass translation, augment monolingual data Mass translation
-
[149]
Telugu, Tamil, Gujarati, Punjabi, Hindi Multilingual models -
-
[150]
Amharic, Catalan, Greek, Estonian, Basque, Farsi, Hindi, Croatian, Javanese, Georgian, Khmer, Kannada, Lithuanian, Latvian, Malay, Mongolian, Burmese, MinNan, Norwegian, Slovak, Slovene, Serbian, Tamil, Tagalog, Turkish Multilingual models -
-
[151]
UpperSorbian Back-translation -
-
[152]
Cantonese Multilingual models -
-
[153]
Khmer Back-translation -
-
[154]
Tamil, Sinhala Back-translation -
-
[155]
Bengali, Telegu, Khmer, Malay Prompt engineering -
-
[156]
Faroese Prompt engineering -
-
[157]
- Prompt engineering -
-
[158]
Slovenian Monolingual model -
-
[159]
Vallader Prompt engineering -
-
[160]
- Augment monolingual data Re-ordering monolingual sentences
-
[161]
Algerian Multilingual models -
-
[162]
Tibetan, Mongolian, Uyghur Multilingual models, family of LRLs -
-
[163]
- Back-translation -
-
[164]
Uyghur Monolingual model -
-
[166]
- Augment monolingual data Fadaee, re-ordering monolingual sentences
-
[167]
- Adaptive learning -
-
[168]
Afrikaans, Amharic, Belarusian, Welsh, Irish, Scottish, Galician, Hausa, Georgian, Kazakh, Khmer, Kyrgyz, Limburgish, Myanmar, Bokmål, Nynorsk, Occitan, Sinhala, Tajik, Turkmen, Tatar, Uighur, Uzbek, Yiddish Prompt engineering -
-
[169]
Turkish Adaptive learning, prompt engineering, vo- cab extension -
-
[170]
Zhuang Prompt engineering -
-
[171]
Turkish Adaptive learning, monolingual model -
-
[172]
- Mass translation, augment monolingual data Mass translation, Fadaee, Soft Contextual Data Augmentation, SwitchOut
-
[173]
- Back-translation, augment monolingual data Fadaee, SwitchOut, SMRT Simulated mul- tiple reference training, replace main POS tags
-
[174]
- Augment monolingual data Re-ordering monolingual sentences, para- phrase
-
[175]
Bengali Monolingual model -
-
[176]
Marathi Prompt engineering -
-
[177]
Indonesian Monolingual model, adaptive learning -
-
[178]
Data extracted for RQ4 and RQ5
Odia Prompt engineering - Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review 41 Table 5. Data extracted for RQ4 and RQ5. The RQ4 column identifies the publisher of each study. The publication/venue column indicates the sou...
-
[179]
ACL ACL EACL - The European Chapter of the ACL USA 2024 Transformer mt5, flant5base
2024
-
[180]
IEEE IEEE Access China 2019 RNN, Transformer, GAN QuGAN, BERT
2019
-
[181]
ACL ACL EMNLP - Empirical Methods in Natural Language Processing USA 2020 RNN, Transformer BERT
2020
-
[182]
ACL ACL NLP4DH - International Conference on Natural Language Processing for Digital Humanities Japan 2023 Transformer t5
2023
-
[183]
ACL ACL NAACL - North American Chapter of the Association for Computational Lin- guistics Russia 2022 Transformer XLM, RoBERTa
2022
-
[184]
ACL ACL - Annual Meeting of the Association for Computational Linguistics China 2019 Transformer -
2019
-
[185]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing India 2023 Transformer mT5
2023
-
[186]
ACL ACL LREC - International Conference on Language Resources and Evaluation Spain 2020 Transformer mBERT, BERTeus
2020
-
[187]
ACL ACL - Transactions of the Association for Computational Linguistics Japan 2020 Transformer XLM
2020
-
[188]
De Gruyter De Gruyter - Open Computer Science USA 2019 LSTM, RNN -
2019
-
[189]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing Turkey 2023 Transformer, BiLSTM, RNN, LSTM -
2023
-
[190]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2021 Transformer -
2021
-
[191]
ACL ACL EMNLP - Empirical Methods in Natural Language Processing Hong Kong 2021 Transformer mBART, GPT2
2021
-
[192]
IJCAI IJCAI - International Joint Conference on Artificial Intelligence China 2021 set2sequence -
2021
-
[193]
IEEE IEEE Access Indonesia 2023 Transformer mt5
2023
-
[194]
ACL ACL - Proceedings of the 21st Annual Conference of the European Association for Machine Translation USA 2018 SMT, Transformer mBERT
2018
-
[195]
ACL ACL EMNLP - Empirical Methods in Natural Language Processing USA 2018 Transformer -
2018
-
[196]
ACL ACL - Annual Meeting of the Association for Computational Linguistics China 2020 Seq2Seq, Transformer -
2020
-
[197]
ARXIV ARXIV USA 2021 Transformer -
2021
-
[198]
ARXIV ARXIV India 2024 Transformer -
2024
-
[199]
ARXIV ARXIV USA 2024 Transformer XLM
2024
-
[200]
ARXIV ARXIV UK 2024 Transformer mBART
2024
-
[201]
ARXIV ARXIV USA 2023 Transformer GPT2
2023
-
[202]
ARXIV ARXIV Korea 2024 Transformer CLaude2.1, LLama27b, LLama213b, Mis- tral7b
2024
-
[203]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing India 2020 GAN -
2020
-
[204]
ACL ACL - Transactions of the Association for Computational Linguistics UK 2018 LSTM, RNN -
2018
-
[205]
ACL ACL WMT - Conference on Machine Translation USA 2020 Transformer GPT2
2020
-
[206]
ACL ACL - Workshop on NLP for Similar Languages, Varieties and Dialects Sweden 2022 Transformer, BiLSTM, RNN, LSTM -
2022
-
[207]
ACM ACM SoICT - Symposium on Information and Communication Technology Vietnam 2022 Transformer mBART50
2022
-
[208]
IEEE IEEE - International Conference on Advances in ICT for Emerging Regions Sri Lanka 2020 LSTM, RNN -
2020
-
[209]
ACL ACL - Proceedings of the Workshop on Multilingual Information Access (MIA) USA 2022 Transformer mt5, mBERT, CORA
2022
-
[210]
ACL ACL LREC - International Conference on Language Resources and Evaluation Faroe Islands 2024 Transformer -
2024
-
[211]
ACL ACL LREC - Joint International Conference on Computational Linguistics, Language Resources and Evaluation China 2024 Transformer -
2024
-
[212]
Frontiers Frontiers in Artificial Intelligence Slovenia 2023 Transformer t5
2023
-
[213]
Springer Springer - International Journal of Information Technology Switzerland 2024 Transformer -
2024
-
[214]
ACM ACM APIT - Asia Pacific Information Technology Conference China 2023 - -
2023
-
[215]
Springer Springer - International Journal of Information Technology Algeria 2024 Transformer -
2024
-
[216]
Springer Springer - Machine Translation China 2022 Transformer -
2022
-
[217]
ACL ACL - Transactions of the Association for Computational Linguistics China 2022 - -
2022
-
[218]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2024 Transformer -
2024
-
[219]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2020 Transformer -
2020
-
[220]
ARXIV ARXIV China 2021 - -
2021
-
[221]
IEEE ACS/IEEE - International Conference on Computer Systems and Applications Egypt 2023 Transformer, LSTM, RNN Helskini
2023
-
[222]
ACL ACL LoResMT - Workshop on Technologies for Machine Translation of Low-Resource Languages USA 2024 Transformer Bloomz
2024
-
[223]
ARXIV ARXIV Turkey 2024 Transformer LLaMa7b
2024
-
[224]
ARXIV ARXIV China 2024 - -
2024
-
[225]
ARXIV ARXIV Turkey 2024 Transformer Mistral7B, GPT2
2024
-
[226]
Springer Springer - Machine Translation Ireland 2021 - -
2021
-
[227]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing China 2022 - -
2022
-
[228]
ACM ACM TALLIP - Transactions on Asian and Low-Resource Language Information Processing Japan 2024 - -
2024
-
[229]
IEEE IEEE ICICT4SD - International Conference on Information and Communication Tech- nology for Sustainable Development Bangladesh 2023 Transformer GPT2
2023
-
[230]
IEEE IEEE INCOFT - International Conference on Futuristic Technologies India 2023 - -
2023
-
[231]
IEEE IEEE ICAICTA - International Conference of Advanced Informatics: Concept, Theory and Application Indonesia 2023 Transformer mBERT, XLM, IndoBERT
2023
-
[232]
Nikolov Table 6
IEEE IEEE CCPIS - International Conference on Circuits, Power and Intelligent Systems India 2023 Transformer - 42 Josh McGiff and Nikola S. Nikolov Table 6. Data extracted for RQ7, RQ8, and RQ9. The RQ7 column describes the scarcity of data available for low-resource languages...
2023
-
[233]
Kashmiri: 26kS BLEU, BERTScore, chrf++ Translation
-
[234]
Tibetan: 22kS BLEU Question Answering
-
[235]
Cherokee: 14kS BLEU Translation
-
[236]
- Recall, F1-score, exact match Question Answering
-
[237]
Hindi: 5kS, Bengali: 5kS BLEU, ROUGE, METEOR Translation
-
[238]
Basque: 2kS F1-score Question Answering
-
[239]
Turkish: 207kS BLEU Translation
-
[240]
Azerbaijan: 3.1mT, Hindi: 7.3mT, Uzbek: 2.2mT, Turkish: 0.8mT BLEU, METEOR Translation
-
[241]
- BLEU, Xtreme Indosum, TyDiQA, XPer- sona Language model, translation, summarisa- tion, conversation, question answering
-
[242]
- BLEU, ROUGE Translation
-
[243]
- SacreBLEU Translation
-
[244]
Telugu: 750kS BLEU, human evaluation Translation
-
[245]
- BLEU, Entity match rate, Success F1 Conversation
-
[246]
Azerbaijani: 5946S, Glacian: 10kS BLEU Translation
-
[247]
- Human evaluation Language model
-
[248]
Estonian: 6.1GB Accuracy, unlabeled attachment score Language model
-
[249]
Cantonese: 1.1mS SacreBLEU, hLEPOR, COMET, BERTscore, Human evaluation Translation
-
[250]
Iranian Persian: 1bT, Modern Greek: 1bT, Standard Arabic: 1bT, Turkish: 1bT, Lithuanian: 100mT, Hindi: 100mT, Catalan: 100mT, Slovak: 100mT, Norwegian Bokmal: 100mT, Estonian: 100mT, Bengali: 100mT, Latvian: 100mT, Serbian: 100mT, Slovenian: 100mT, Tamil: 100mT, Albanian: 100m...
-
[251]
Bengali: 5kS COPA Question Answering
-
[252]
Telugu: 70kS,Tamil: 70kS,Gujarati: 70kS,Punjabi: 70kS,Hindi: 70kS BLEU, TER Translation
-
[253]
Amharic: 511kS, Catalan: 788kS, Greek: 744kS, Estonian: 556kS, Basque: 647kS, Farsi: 738kS, Hindi: 666kS, Croatian: 620kS, Javanese: 622kS, Georgian: 580kS, Khmer: 579kS, Kannada: 434kS, Lithuanian: 554kS, Latvian: 587kS, Malay: 702kS, Mongolian: 629kS, Burmese: 576kS, Min-Nan...
-
[254]
UpperSorbian: 600kS BLEU Language model, translation
-
[255]
Cantonese: 35kS SacreBLEU Translation
-
[256]
Khmer: 70kS BLEU Translation
-
[257]
Tamil: 26kS, Sinhala: 26kS BLEU Translation
-
[258]
- Recall Question Answering
-
[259]
- Sentence-BERT, BLEU, chrF -
-
[260]
- COMET, BLEURT, chrF++ Translation
-
[261]
- BoolQ, Commitment Bank, COPA, Recog- nizing Textual Entailment, The Winograd Schema Challenge, classification report Language model
-
[262]
- BLEU, METEOR, ROUGE, CIDEr Language model
-
[263]
Uyghur: 20kS BLEU Translation
-
[264]
- BLEU, chrF, ROUGE Language model, summarisation, transla- tion
-
[265]
Afrikaans: 275kS, Amharic: 89kS, Belarusian: 67kS, Welsh: 289kS, Irish: 289kS, Scottish Gaelic: 16kS, Galician: 515kS, Hausa: 97kS, Georgian: 377kS, Kazakh: 79kS, Khmer: 111kS, Kyrgyz: 27kS, Limburgish: 25kS, Burmese: 24kS, Norwegian Bokmål: 142kS, Norwegian Nynorsk: 486kS, Oc...
-
[266]
Turkish: 273.9mT Perplexity, classification report Language model
-
[267]
Zhuang: 5kS Bleu, chrF Translation
-
[268]
Turkish: 180GB ARC-TR, TruthfulQA-TR Language model, question answering
-
[269]
Bengali: 26GB Perplexity Language model
-
[270]
- BLEU, precision, chrF, TER, Rouge Translation
-
[271]
- Exact match, F1-score, substring match Question Answering
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.