REVIEW 4 major objections 5 minor 1 cited by
Small Languages, Big Models: A Study of Continual Training on Languages of Norway
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A three-stage continual training recipe adapts a 12B English model to Norwegian and Sámi, beating open baselines on most native tasks while running 30% faster.
desk verdict Worth reading for the open model and the efficiency result, but the paper's central quality claim for the three-stage recipe is not supported by the experiments as run. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The three-stage continual pretraining pipeline: Stage 1 trains a new byte-level BPE tokenizer on the target corpus; Stage 2 copies and averages original embeddings into the new embedding matrix and trains only embeddings for 1,000 steps to avoid loss spikes and catastrophic forgetting; Stage 3 unfreezes all parameters and trains the full model on 250B tokens. A secondary mechanism is the hybrid masked-causal objective, which combines standard causal language modeling with masked next-token prediction so the model can be used in multiple inference modes.
What would settle it
Train an 11.4B model from scratch on the same 250B-token mixture, or train an 11.4B model with tokenizer replacement but without the embedding-update stage, and compare on NorQuAD, NoReC, and Tatoeba; if either matches NorMistral-11B, the three-stage claim is not load-bearing.
Extended reading notes
Core claim
The central claim is that replacing the tokenizer, realigning embeddings, and then training the whole model efficiently transfers knowledge from an English-centric 12B model to Bokmål, Nynorsk, and Northern Sámi. The paper reports that NorMistral-11B reaches state-of-the-art results on most tasks in the NorEval suite, improves English-to-Northern Sámi translation from near zero to about 50 BLEU at 16 shots, and runs more than 30% faster than Mistral-Nemo-12B because its 51,200-token vocabulary produces shorter sequences. The authors also show that the same model can be used as a causal generator, a masked language model, or a prefix language model by mixing causal language modeling with masked next-token prediction at a 90/10 ratio.
Load-bearing premise
The evidence that the three-stage recipe beats simpler continual training comes from 7-billion-parameter models trained on one corpus, while the released model is 11.4 billion parameters trained on a different, larger mixture, so the paper assumes the 7B ranking carries over to the final model and to the tiny Northern Sámi portion.
Editorial extensions
If this is right
- A language with only tens of millions of words can be added to an existing large model by continued training on upsampled data, with no architecture change.
- Replacing the tokenizer yields a substantial inference speedup: 30% shorter sequences translate to more than 30% faster per-sample processing.
- The same checkpoint can serve as a generative model and as a bidirectional encoder, so a single training run covers both use cases.
- Continual training can push performance past the base model on native-language benchmarks while risking some regression on multilingual datasets like Belebele.
Reading between the lines
- The recipe is likely to transfer to other small languages with closely related large languages, because the corpus already mixes Bokmål with Danish, Swedish, Icelandic, Faroese, English, and code.
- The 7B ablations that justify the three-stage recipe may not fully determine the 11.4B result; the released model uses a different corpus mixture, so the relative value of the embedding-update stage at scale remains unproven.
- A fairer test of the method would compare against simple vocabulary extension plus full training under the same inference budget, since the paper excludes such baselines on efficiency grounds.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-stage continual pretraining recipe for adapting a 12B English-centric model (Mistral-Nemo-12B) to Norwegian Bokmål, Nynorsk, and Northern Sámi: (1) train a new domain-optimized BPE tokenizer, (2) realign input/output embeddings via 1,000 steps of frozen-backbone training, and (3) train the full model. The authors release NorMistral-11B, three 7B checkpoints, training/evaluation code, and a new Northern Sámi web corpus. They report that the new tokenizer cuts sequence lengths by roughly 30% and gives a corresponding inference speedup (Table 1, Appendix A), and that NorMistral-11B outperforms several existing Norwegian models on most of the NorEval benchmarks (Table 2). The paper also describes a hybrid masked-causal training objective and demonstrates flexible use of the model as a causal, prefix, or bidirectional LM (Section 3.2, Table 3). Section 5 presents ablations at 7B scale, including architecture choice, from-scratch vs. warm-starting, the hybrid objective, and number of training steps. The central claim is that this three-stage continual training substantially improves both downstream performance and inference efficiency for low-resource languages.
Significance. If the central claim were fully supported, this would be a practically valuable recipe for adapting large English-centric models to lower-resource languages, because it reduces inference cost while reusing existing pretrained knowledge. The paper has clear strengths: the models, code, evaluation prompts, and the new Sámi corpus are openly released; the tokenizer efficiency evidence in Table 1 and Appendix A is concrete and machine-checkable; and the hybrid masked-causal objective is an interesting design choice whose downstream effects are partially documented. The main weakness is that the performance advantage of the three-stage recipe over simple continual training is not established, because no same-scale baseline of that kind is run. The 7B ablations are on a different corpus and scale than the released model, and the evaluation protocol reports maximum scores across prompts, which can inflate apparent gains. These issues are load-bearing for the paper's claimed contribution.
major comments (4)
- [Section 5, Table 4] The only continual-training comparison is 'init. from scratch' versus 'three-stage continual' at 7B scale on the Norwegian Colossal Corpus. The paper explicitly does not consider simple continual training or adapter tuning because they 'necessarily lead to inefficient inference' (Section 5). That rationale addresses inference cost, not model quality. The abstract and introduction claim that the three-stage approach 'substantially improves the downstream performance'; to support that claim, the paper needs a baseline that continually trains the same base model on the same corpus while keeping the original tokenizer (or extending the vocabulary) and evaluates at the same scale. Without such a baseline, the evidence does not show that the three-stage recipe is better than simpler alternatives in quality, only that it is faster. This is a central, load-bearing omission.
- [Section 3.1, Stage 2] The embedding-update stage is asserted to prevent loss spikes and catastrophic forgetting, but no ablation removes or alters this stage. Since the 'three-stage' method is the paper's main contribution, the contribution of Stage 2 should be measured directly, for example by comparing full training from randomly initialized new embeddings, from averaged sub-token embeddings without realignment, and from the proposed 1,000-step realignment. As written, the necessity of the second stage is an unverified component of the central recipe.
- [Section 4, Table 2] NorMistral-11B regresses relative to Mistral-Nemo-12B on Belebele (56.7 vs. 62.8) and NorOpenBookQA (77.9 vs. 86.9 on Bokmål, 77.8 vs. 86.7 on Nynorsk). The text acknowledges this and says it 'requires a further study,' but the abstract and introduction still claim that the approach 'substantially improves the downstream performance' without qualifying these exceptions. Given that the base model outperforms the adapted model on two of the benchmarks, the unconditional performance claim should be revised to state which tasks improve and which regress, and ideally the paper should provide some analysis of why these regressions occur rather than leaving them unexplained.
- [Section 3.3 and Table 2 caption] The evaluation protocol reports the maximum score across five prompt templates and across chosen k-shot settings. The full per-prompt results in Appendix B show very high prompt sensitivity (e.g., Belebele scores for NorMistral-11B range from 22.8 to 56.7 in Table 6). Reporting the maximum rather than a mean or median can systematically inflate performance and makes the headline comparisons sensitive to prompt-search luck. The paper should either justify this protocol, report both the maximum and the mean/median with variance, or provide the full distribution of scores in the main tables so that the SOTA claim is robust to prompt selection.
minor comments (5)
- [Section 4, first paragraph] The text refers to 'NorOpenBookQ' but the benchmark is named NorOpenBookQA; this typo should be corrected.
- [Section 6, 'Norwegian language models'] The phrase 'ranging fron 15M parameters' should read 'ranging from 15M parameters'.
- [Appendix A, first sentence] The phrase 'in-context-lerning' should be 'in-context learning'.
- [Appendix B.8] The text says 'The complete evaluation on ASK-GEC is provided in Table 15,' but the table immediately following is Table 14; the citation is off by one.
- [Appendix B.10.1] The sentence 'We used the following four prompt templates for testing all language models on translation to Northern Sámi' appears under the English-to-Bokmål translation subsection; it should refer to Bokmål translation.
Circularity Check
No significant circularity: the three-stage recipe and its efficiency and quality claims are empirically tested, with only minor non-load-bearing self-citations.
full rationale
The central derivation chain is empirical rather than definitional: Stage 1 trains a new BPE tokenizer on the target corpus, Stage 2 copies and realigns embeddings, and Stage 3 trains the full model. The claimed efficiency gain is not predicted from the tokenizer's fitted parameters; it is measured directly in Appendix A (Tables 1 and 5) as a consequence of replacing the vocabulary, not as a hidden reuse of a fitted quantity. The claimed quality gain is supported by Table 2, which compares against seven external baselines, and Table 4, which compares from-scratch versus warm-started training at 7B scale; neither comparison reduces by construction to the training recipe. The paper explicitly excludes simple continual training and adapter tuning for efficiency reasons, stating in Section 5 that they are not considered 'because they necessarily lead to inefficient inference'; the absence of a same-scale quality baseline against those methods is a methodological gap, not a circularity. The hybrid masked-causal objective is motivated partly by self-citations (Samuel 2024; Charpentier and Samuel 2024), but Table 4 shows the hybrid objective did not produce an overall gain, so those citations are not load-bearing for the main claim. The same-group benchmarks (NorEval, NRK-Quiz-QA, NorOpenBookQA, NorCommonsenseQA) are native-speaker datasets used as outcome measures, not training inputs, and do not make the derivation circular. No equation or fitted parameter is renamed as a prediction, so no specific circular reduction can be exhibited from the paper's own equations or construction.
Assumptions & free parameters
free parameters (6)
- Vocabulary size of new tokenizer =
51,200 tokens
- Upsampling factors for Nynorsk and Northern Sámi =
4x and 16x (Figure 1)
- Causal/MNTP objective ratio =
90% / 10%
- Embedding-only update steps =
1,000 steps
- Peak learning rate and schedule =
1e-4, 1,000 warmup, 10,000 decay steps
- Total training steps =
60,000
assumptions (5)
- domain assumption Chinchilla scaling laws (Hoffmann et al. 2022) apply to the Norwegian data-constrained setting, making 250B tokens compute-optimal for 11.4B parameters.
- domain assumption Muennighoff et al. (2023) repetition results: up to four repeats are harmless, so upsampling Nynorsk and Northern Sámi by 4x and 16x will not hurt.
- ad hoc to paper The conclusions of the 7B ablations on the Norwegian Colossal Corpus transfer to the 11.4B model trained on the broader Mímir Core mixture.
- domain assumption Embedding transfer by copying overlapping token vectors and averaging sub-token vectors, then 1,000 steps of frozen-backbone training, aligns the new tokenizer well enough to avoid catastrophic forgetting.
- ad hoc to paper The evaluation protocol that reports the maximum score across prompts and shot counts is a valid way to measure model quality.
Cite this review
Pith. "Pith review of Small Languages, Big Models: A Study of Continual Training on Languages of Norway." pith.science (2026). https://pith.science/paper/PM442M4E
@misc{pith2026241206484,
author = {Pith},
title = {Pith review of: Small Languages, Big Models: A Study of Continual Training on Languages of Norway},
year = {2026},
howpublished = {\url{https://pith.science/paper/PM442M4E}},
note = {Machine review of arXiv:2412.06484}
}
read the original abstract
Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern S\'ami. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokm\r{a}l, Nynorsk, and Northern S\'ami with 11.4 billion parameters: NorMistral-11B.
Figures
Forward citations
Cited by 1 Pith paper
-
Multi-label Scandinavian Language Identification (SLIDE)
SLIDE introduces a manually labeled multi-label evaluation set and BERT/FastText models for Scandinavian language identification, using machine translation identity as a silver-labeling signal.
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023. https://openreview.net/forum?id=hmOwOZWzYE GQA : Training generalized multi-query transformer models from multi-head checkpoints . In The 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[4]
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. https://aclanthology.org/2024.acl-long.44 The belebele benchmark: a parallel reading comprehension dataset in 122 language variants . In Proceedings of the 62nd Annual Meeting of the ...
2024
-
[5]
Adrien Barbaresi. 2021. https://doi.org/10.18653/v1/2021.acl-demo.15 Trafilatura: A web scraping library and command-line tool for text discovery and extraction . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, page...
-
[6]
Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. https://openreview.net/forum?id=IW1PR7vEBf LLM 2vec: Large language models are secretly powerful text encoders . In First Conference on Language Modeling
2024
-
[7]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[8]
Christopher Bryant, Mariano Felice, and Ted Briscoe. 2017. https://doi.org/10.18653/v1/P17-1074 Automatic annotation and evaluation of error types for grammatical error correction . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 793--805, Vancouver, Canada. Association for Computat...
Show all 65 references
-
[9]
Lucas Georges Gabriel Charpentier and David Samuel. 2024. http://arxiv.org/abs/2410.24159 Gpt or bert: why not both?
2024 arXiv
-
[10]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning ...
2020 doi
-
[11]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. https://openreview.net/forum?id=dXiGWqBoxaD GPT 3.int8(): 8-bit matrix multiplication for transformers at scale . In Advances in Neural Information Processing Systems
2022
-
[12]
Tim Dettmers and Luke Zettlemoyer. 2023. https://proceedings.mlr.press/v202/dettmers23a.html The case for 4-bit precision: k-bit inference scaling laws . In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Rese...
2023
-
[13]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[14]
Ethan Ewer, Daewon Chae, Thomas Zeng, Jinkyu Kim, and Kangwook Lee. 2024. http://arxiv.org/abs/2410.01600 ENTP : Encoder-only next token prediction
2024 arXiv
-
[15]
Philip Gage. 1994. https://api.semanticscholar.org/CorpusID:59804030 A new algorithm for data compression . The C Users Journal archive, 12:23--38
1994
-
[16]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2024 doi
-
[17]
Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Ba \ n \'o n, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ram \' rez-S \'a nchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and J \"o rg Tiedemann. 2024. https://aclanthology.org/2024.lr...
2024
-
[18]
Giellatekno and Divvun. 2016. https://doi.org/10.18710/8AK7KZ SIKOR North Saami corpus
2016 doi
-
[19]
Alexander H \"a gele, Elie Bakouch, Atli Kosson, Loubna Ben allal, Leandro Von Werra, and Martin Jaggi. 2024. https://openreview.net/forum?id=ompl7supoX Scaling laws and compute-optimal training beyond fixed training durations . In Workshop on Efficient Systems for Foundation ...
2024
-
[20]
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Sim...
2022
-
[21]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lo RA : Low-rank adaptation of large language models . In International Conference on Learning Representations
2022
-
[22]
Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, Andr \'e Martins, Fran c ois Yvon, and Hinrich Sch \"u tze. 2023. https://doi.org/10.18653/v1/2023.acl-long.61 Glot500: Scaling multilingual ...
2023 doi
-
[23]
Sardana Ivanova, Fredrik Andreassen, Matias Jentoft, Sondre Wold, and Lilja vrelid. 2023. https://aclanthology.org/2023.nodalida-1.17 N or Q u AD : N orwegian question answering dataset . In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pag...
2023
-
[24]
Matias Jentoft. 2023. https://www.duo.uio.no/handle/10852/103885 Grammatical error correction with byte-level language models . Master's thesis, University of Oslo
2023
-
[25]
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...
2023 arXiv
-
[26]
Amir Hossein Kargaran, Ayyoob Imani, Fran c ois Yvon, and Hinrich Schuetze. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.410 G lot LID : Language identification for low-resource languages . In Findings of the Association for Computational Linguistics: EMNLP 2023, page...
2023 doi
-
[27]
Per Kummervold, Freddy Wetjen, and Javier de la Rosa. 2022. https://aclanthology.org/2022.lrec-1.410 The N orwegian colossal corpus: A text corpus for training large N orwegian language models . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pag...
2022
-
[28]
Per E Kummervold, Javier De la Rosa, Freddy Wetjen, and Svein Arne Brygfjeld. 2021. https://aclanthology.org/2021.nodalida-main.3 Operationalizing a national digital library: The case for a N orwegian transformer model . In Proceedings of the 23rd Nordic Conference on Computat...
2021
-
[29]
Andrey Kutuzov, Jeremy Barnes, Erik Velldal, Lilja vrelid, and Stephan Oepen. 2021. https://aclanthology.org/2021.nodalida-main.4 Large-scale contextualised language modelling for N orwegian . In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa)...
2021
-
[30]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics
2004
-
[31]
Lauvrak, Jon Espen Ingvaldsen, Simen Eide, Jon Atle Gulla, and Zhirong Yang
Peng Liu, Lemei Zhang, Terje Farup, Even W. Lauvrak, Jon Espen Ingvaldsen, Simen Eide, Jon Atle Gulla, and Zhirong Yang. 2024 a . http://arxiv.org/abs/2312.01314 NLEB ench+ N or GLM : A comprehensive empirical analysis and benchmark dataset for generative language models in N orwegian
2024 arXiv
-
[32]
Xiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu, and Dahua Lin. 2024 b . https://openreview.net/forum?id=JO7k0SJ5V6 Scaling laws of R o PE -based extrapolation . In The Twelfth International Conference on Learning Representations
2024
-
[33]
Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations
2019
-
[34]
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Z...
2024 arXiv
-
[35]
Sheng Lu, Hendrik Schuff, and Iryna Gurevych. 2024. https://doi.org/10.18653/v1/2024.naacl-long.325 How are prompts different in terms of sensitivity? In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human La...
2024 doi
-
[36]
Risto Luukkonen, Jonathan Burdge, Elaine Zosa, Aarne Talman, Ville Komulainen, Vaino Hatanpaa, Peter Sarlin, and Sampo Pyysalo. 2024. https://api.semanticscholar.org/CorpusID:268857103 Poro 34b and the blessing of multilinguality . ArXiv, abs/2404.01856
2024 arXiv
-
[37]
Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki ...
2023 doi
- [38]
-
[39]
Michael McCloskey and Neal J. Cohen. 1989. https://doi.org/https://doi.org/10.1016/S0079-7421(08)60536-8 Catastrophic interference in connectionist networks: The sequential learning problem . In Gordon H. Bower, editor, Psychology of Learning and Motivation, volume 24 of Psych...
1989 doi
-
[40]
Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023. https://openreview.net/forum?id=j5BuTrEj35 Scaling data-constrained language models . In Thirty-seventh Conference on Neural I...
2023
-
[41]
Rossi, and Thien Huu Nguyen
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2023. http://arxiv.org/abs/2309.09400 Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages
2023 arXiv
-
[42]
Nguyen and Julian Salazar
Toan Q. Nguyen and Julian Salazar. 2019. https://aclanthology.org/2019.iwslt-1.17 Transformers without tears: Improving the normalization of self-attention . In Proceedings of the 16th International Conference on Spoken Language Translation, Hong Kong. Association for Computat...
2019
-
[43]
Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Haji c , Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016. https://aclanthology.org/L16-1262 U niversal D ependencies v1: A mul...
2016
-
[44]
Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Haji c , Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020. https://aclanthology.org/2020.lrec-1.497 U niversal D ependencies v2: An evergrowing multilingual treebank co...
2020
-
[45]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...
2002
-
[46]
Ronny Paul, Himanshu Buckchash, Shantipriya Parida, and Dilip K. Prasad. 2024. https://api.semanticscholar.org/CorpusID:269635411 Towards a more inclusive ai: Progress and perspectives in large language model training for the s \'a mi language . ArXiv, abs/2405.05777
2024 arXiv
-
[47]
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. http://arxiv.org/abs/2406.17557 The F ine W eb datasets: Decanting the web for the finest text data at scale
2024 arXiv
-
[48]
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. https://openai.com/index/language-unsupervised/ Improving language understanding by generative pre-training . OpenAI Blog
2018
-
[49]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1)
2020
-
[50]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. https://dl.acm.org/doi/10.5555/3433701.3433727 ZeRO : memory optimizations toward training trillion parameter models . In Proceedings of the International Conference for High Performance Computing, Network...
2020
-
[51]
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. https://doi.org/10.1145/3394486.3406703 Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters . In Proceedings of the 26th ACM SIGKDD International Conferenc...
2020
-
[52]
David Samuel. 2024. https://openreview.net/forum?id=BCA9NMZkLS BERT s are generative in-context learners . In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[53]
David Samuel, Andrey Kutuzov, Samia Touileb, Erik Velldal, Lilja vrelid, Egil R nningstad, Elina Sigdel, and Anna Palatkina. 2023. https://aclanthology.org/2023.nodalida-1.61 N or B ench -- a benchmark for N orwegian language models . In Proceedings of the 24th Nordic Conferen...
2023
-
[54]
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili'c, Daniel Hesslow, Roman Castagn'e, Alexandra Sasha Luccioni, François Yvon, Matthias Gall \'e , Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Be...
2022 arXiv
-
[55]
Noam M. Shazeer. 2020. https://api.semanticscholar.org/CorpusID:211096588 GLU variants improve transformer . ArXiv, abs/2002.05202
2020 arXiv
-
[56]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. http://arxiv.org/abs/1909.08053 Megatron- LM : T raining multi-billion parameter language models using model parallelism
2020 arXiv
-
[57]
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021. http://arxiv.org/abs/2104.09864 Roformer: Enhanced transformer with rotary position embedding
2021 arXiv
-
[58]
J \"o rg Tiedemann. 2020. https://aclanthology.org/2020.wmt-1.139 The tatoeba translation challenge -- realistic data sets for low resource and multilingual MT . In Proceedings of the Fifth Conference on Machine Translation, pages 1174--1182, Online. Association for Computatio...
2020
-
[59]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://api.semanticscholar.org...
2023 arXiv
-
[60]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf Attention is all you need . In Advances in Ne...
2017
-
[61]
Erik Velldal, Lilja vrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, and Fredrik J rgensen. 2018. https://aclanthology.org/L18-1661 N o R e C : The N orwegian review corpus . In Proceedings of the Eleventh International Conference on Language Resources and Ev...
2018
-
[62]
Wietse de Vries and Malvina Nissim. 2021. https://doi.org/10.18653/v1/2021.findings-acl.74 As good as new. how to successfully recycle E nglish GPT -2 to make models for other languages . In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 836-...
2021 doi
-
[63]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020 doi
-
[64]
Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, and Vassilina Nikoulina. 2023. https://doi....
2023 doi
-
[65]
Biao Zhang and Rico Sennrich. 2019. https://openreview.net/references/pdf?id=S1qBAf6rr Root Mean Square Layer Normalization . In Advances in Neural Information Processing Systems 32, Vancouver, Canada
2019
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.