Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Small Languages, Big Models: A Study of Continual Training on Languages of Norway

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A three-stage continual training recipe adapts a 12B English model to Norwegian and Sámi, beating open baselines on most native tasks while running 30% faster.

desk verdict Worth reading for the open model and the efficiency result, but the paper's central quality claim for the three-stage recipe is not supported by the experiments as run. read the letter →

arxiv 2412.06484 v2 pith:PM442M4E submitted 2024-12-09 cs.CL

classification cs.CL
keywords continualpretrainingtokenizeradaptationlow-resourcelanguagesNorwegianNorthernSámilanguagemodelembeddingrealignmenthybridmasked-causal
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether a large English-centric model can be made to serve languages with far less data, such as Norwegian and especially Northern Sámi. It proposes a three-stage continual training method: build a byte-level BPE tokenizer for the target corpus, realign the embedding matrix to the new vocabulary while freezing all other parameters, then train the full model. The resulting NorMistral-11B outperforms comparable open models on most Norwegian benchmarks and is more than 30% faster on Norwegian inputs because its tokenizer cuts average sequence length by about 30%. If this recipe holds, it gives low-resource languages a practical path to large models without training from scratch.

What carries the argument

The three-stage continual pretraining pipeline: Stage 1 trains a new byte-level BPE tokenizer on the target corpus; Stage 2 copies and averages original embeddings into the new embedding matrix and trains only embeddings for 1,000 steps to avoid loss spikes and catastrophic forgetting; Stage 3 unfreezes all parameters and trains the full model on 250B tokens. A secondary mechanism is the hybrid masked-causal objective, which combines standard causal language modeling with masked next-token prediction so the model can be used in multiple inference modes.

What would settle it

Train an 11.4B model from scratch on the same 250B-token mixture, or train an 11.4B model with tokenizer replacement but without the embedding-update stage, and compare on NorQuAD, NoReC, and Tatoeba; if either matches NorMistral-11B, the three-stage claim is not load-bearing.

Watch

Extended reading notes

Core claim

The central claim is that replacing the tokenizer, realigning embeddings, and then training the whole model efficiently transfers knowledge from an English-centric 12B model to Bokmål, Nynorsk, and Northern Sámi. The paper reports that NorMistral-11B reaches state-of-the-art results on most tasks in the NorEval suite, improves English-to-Northern Sámi translation from near zero to about 50 BLEU at 16 shots, and runs more than 30% faster than Mistral-Nemo-12B because its 51,200-token vocabulary produces shorter sequences. The authors also show that the same model can be used as a causal generator, a masked language model, or a prefix language model by mixing causal language modeling with masked next-token prediction at a 90/10 ratio.

Load-bearing premise

The evidence that the three-stage recipe beats simpler continual training comes from 7-billion-parameter models trained on one corpus, while the released model is 11.4 billion parameters trained on a different, larger mixture, so the paper assumes the 7B ranking carries over to the final model and to the tiny Northern Sámi portion.

Editorial extensions

If this is right

  • A language with only tens of millions of words can be added to an existing large model by continued training on upsampled data, with no architecture change.
  • Replacing the tokenizer yields a substantial inference speedup: 30% shorter sequences translate to more than 30% faster per-sample processing.
  • The same checkpoint can serve as a generative model and as a bidirectional encoder, so a single training run covers both use cases.
  • Continual training can push performance past the base model on native-language benchmarks while risking some regression on multilingual datasets like Belebele.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The recipe is likely to transfer to other small languages with closely related large languages, because the corpus already mixes Bokmål with Danish, Swedish, Icelandic, Faroese, English, and code.
  • The 7B ablations that justify the three-stage recipe may not fully determine the 11.4B result; the released model uses a different corpus mixture, so the relative value of the embedding-update stage at scale remains unproven.
  • A fairer test of the method would compare against simple vocabulary extension plus full training under the same inference budget, since the paper excludes such baselines on efficiency grounds.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a three-stage continual pretraining recipe for adapting a 12B English-centric model (Mistral-Nemo-12B) to Norwegian Bokmål, Nynorsk, and Northern Sámi: (1) train a new domain-optimized BPE tokenizer, (2) realign input/output embeddings via 1,000 steps of frozen-backbone training, and (3) train the full model. The authors release NorMistral-11B, three 7B checkpoints, training/evaluation code, and a new Northern Sámi web corpus. They report that the new tokenizer cuts sequence lengths by roughly 30% and gives a corresponding inference speedup (Table 1, Appendix A), and that NorMistral-11B outperforms several existing Norwegian models on most of the NorEval benchmarks (Table 2). The paper also describes a hybrid masked-causal training objective and demonstrates flexible use of the model as a causal, prefix, or bidirectional LM (Section 3.2, Table 3). Section 5 presents ablations at 7B scale, including architecture choice, from-scratch vs. warm-starting, the hybrid objective, and number of training steps. The central claim is that this three-stage continual training substantially improves both downstream performance and inference efficiency for low-resource languages.

Significance. If the central claim were fully supported, this would be a practically valuable recipe for adapting large English-centric models to lower-resource languages, because it reduces inference cost while reusing existing pretrained knowledge. The paper has clear strengths: the models, code, evaluation prompts, and the new Sámi corpus are openly released; the tokenizer efficiency evidence in Table 1 and Appendix A is concrete and machine-checkable; and the hybrid masked-causal objective is an interesting design choice whose downstream effects are partially documented. The main weakness is that the performance advantage of the three-stage recipe over simple continual training is not established, because no same-scale baseline of that kind is run. The 7B ablations are on a different corpus and scale than the released model, and the evaluation protocol reports maximum scores across prompts, which can inflate apparent gains. These issues are load-bearing for the paper's claimed contribution.

major comments (4)
  1. [Section 5, Table 4] The only continual-training comparison is 'init. from scratch' versus 'three-stage continual' at 7B scale on the Norwegian Colossal Corpus. The paper explicitly does not consider simple continual training or adapter tuning because they 'necessarily lead to inefficient inference' (Section 5). That rationale addresses inference cost, not model quality. The abstract and introduction claim that the three-stage approach 'substantially improves the downstream performance'; to support that claim, the paper needs a baseline that continually trains the same base model on the same corpus while keeping the original tokenizer (or extending the vocabulary) and evaluates at the same scale. Without such a baseline, the evidence does not show that the three-stage recipe is better than simpler alternatives in quality, only that it is faster. This is a central, load-bearing omission.
  2. [Section 3.1, Stage 2] The embedding-update stage is asserted to prevent loss spikes and catastrophic forgetting, but no ablation removes or alters this stage. Since the 'three-stage' method is the paper's main contribution, the contribution of Stage 2 should be measured directly, for example by comparing full training from randomly initialized new embeddings, from averaged sub-token embeddings without realignment, and from the proposed 1,000-step realignment. As written, the necessity of the second stage is an unverified component of the central recipe.
  3. [Section 4, Table 2] NorMistral-11B regresses relative to Mistral-Nemo-12B on Belebele (56.7 vs. 62.8) and NorOpenBookQA (77.9 vs. 86.9 on Bokmål, 77.8 vs. 86.7 on Nynorsk). The text acknowledges this and says it 'requires a further study,' but the abstract and introduction still claim that the approach 'substantially improves the downstream performance' without qualifying these exceptions. Given that the base model outperforms the adapted model on two of the benchmarks, the unconditional performance claim should be revised to state which tasks improve and which regress, and ideally the paper should provide some analysis of why these regressions occur rather than leaving them unexplained.
  4. [Section 3.3 and Table 2 caption] The evaluation protocol reports the maximum score across five prompt templates and across chosen k-shot settings. The full per-prompt results in Appendix B show very high prompt sensitivity (e.g., Belebele scores for NorMistral-11B range from 22.8 to 56.7 in Table 6). Reporting the maximum rather than a mean or median can systematically inflate performance and makes the headline comparisons sensitive to prompt-search luck. The paper should either justify this protocol, report both the maximum and the mean/median with variance, or provide the full distribution of scores in the main tables so that the SOTA claim is robust to prompt selection.
minor comments (5)
  1. [Section 4, first paragraph] The text refers to 'NorOpenBookQ' but the benchmark is named NorOpenBookQA; this typo should be corrected.
  2. [Section 6, 'Norwegian language models'] The phrase 'ranging fron 15M parameters' should read 'ranging from 15M parameters'.
  3. [Appendix A, first sentence] The phrase 'in-context-lerning' should be 'in-context learning'.
  4. [Appendix B.8] The text says 'The complete evaluation on ASK-GEC is provided in Table 15,' but the table immediately following is Table 14; the citation is off by one.
  5. [Appendix B.10.1] The sentence 'We used the following four prompt templates for testing all language models on translation to Northern Sámi' appears under the English-to-Bokmål translation subsection; it should refer to Bokmål translation.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the three-stage recipe and its efficiency and quality claims are empirically tested, with only minor non-load-bearing self-citations.

full rationale

The central derivation chain is empirical rather than definitional: Stage 1 trains a new BPE tokenizer on the target corpus, Stage 2 copies and realigns embeddings, and Stage 3 trains the full model. The claimed efficiency gain is not predicted from the tokenizer's fitted parameters; it is measured directly in Appendix A (Tables 1 and 5) as a consequence of replacing the vocabulary, not as a hidden reuse of a fitted quantity. The claimed quality gain is supported by Table 2, which compares against seven external baselines, and Table 4, which compares from-scratch versus warm-started training at 7B scale; neither comparison reduces by construction to the training recipe. The paper explicitly excludes simple continual training and adapter tuning for efficiency reasons, stating in Section 5 that they are not considered 'because they necessarily lead to inefficient inference'; the absence of a same-scale quality baseline against those methods is a methodological gap, not a circularity. The hybrid masked-causal objective is motivated partly by self-citations (Samuel 2024; Charpentier and Samuel 2024), but Table 4 shows the hybrid objective did not produce an overall gain, so those citations are not load-bearing for the main claim. The same-group benchmarks (NorEval, NRK-Quiz-QA, NorOpenBookQA, NorCommonsenseQA) are native-speaker datasets used as outcome measures, not training inputs, and do not make the derivation circular. No equation or fitted parameter is renamed as a prediction, so no specific circular reduction can be exhibited from the paper's own equations or construction.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central evaluation and training choices rest on several empirical scaling laws and hand-picked hyperparameters. The free parameters are mostly standard training choices; the upsampling factors and objective ratio are chosen by hand. The axioms are external empirical results (Chinchilla, Muennighoff repetition limits) and the unverified transfer of 7B ablation conclusions to the 11B model. No new theoretical entities are introduced.

free parameters (6)
  • Vocabulary size of new tokenizer = 51,200 tokens
    Chosen to compress Norwegian text (Table 1) and reduce parameters; not optimized against a validation objective.
  • Upsampling factors for Nynorsk and Northern Sámi = 4x and 16x (Figure 1)
    Hand-picked to balance data scarcity while following Muennighoff et al. (2023); not tuned by experiments in this paper.
  • Causal/MNTP objective ratio = 90% / 10%
    Described as 'conservative' in Section 3.2; no sweep is reported.
  • Embedding-only update steps = 1,000 steps
    Chosen in Stage 2 (Section 3.1) without an ablation.
  • Peak learning rate and schedule = 1e-4, 1,000 warmup, 10,000 decay steps
    Standard AdamW values inherited from prior Mistral training; not fitted.
  • Total training steps = 60,000
    Set to match 250B tokens; justified by Chinchilla but not by any experiment here.
assumptions (5)
  • domain assumption Chinchilla scaling laws (Hoffmann et al. 2022) apply to the Norwegian data-constrained setting, making 250B tokens compute-optimal for 11.4B parameters.
    Invoked in Section 2.1 to justify the 250B token budget and 11.4B size.
  • domain assumption Muennighoff et al. (2023) repetition results: up to four repeats are harmless, so upsampling Nynorsk and Northern Sámi by 4x and 16x will not hurt.
    Used in Sections 2.1-2.2 and Section 5 to justify repeating low-resource data; the paper extends the cited result beyond four repeats.
  • ad hoc to paper The conclusions of the 7B ablations on the Norwegian Colossal Corpus transfer to the 11.4B model trained on the broader Mímir Core mixture.
    Table 4 and Section 5 compare 7B models only; the released 11B model is never compared against a from-scratch or simpler continual baseline at its own scale.
  • domain assumption Embedding transfer by copying overlapping token vectors and averaging sub-token vectors, then 1,000 steps of frozen-backbone training, aligns the new tokenizer well enough to avoid catastrophic forgetting.
    Core to Stage 2 (Section 3.1); drawn from de Vries and Nissim (2021), not re-derived or ablated in this paper.
  • ad hoc to paper The evaluation protocol that reports the maximum score across prompts and shot counts is a valid way to measure model quality.
    Section 3.3 and Table 2 caption; this inflates scores and is not corrected for multiple comparisons.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Small Languages, Big Models: A Study of Continual Training on Languages of Norway." pith.science (2026). https://pith.science/paper/PM442M4E

@misc{pith2026241206484,
  author       = {Pith},
  title        = {Pith review of: Small Languages, Big Models: A Study of Continual Training on Languages of Norway},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PM442M4E}},
  note         = {Machine review of arXiv:2412.06484}
}
read the original abstract

Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern S\'ami. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokm\r{a}l, Nynorsk, and Northern S\'ami with 11.4 billion parameters: NorMistral-11B.

Figures

Figures reproduced from arXiv: 2412.06484 by the authors.

Figure 1
Figure 1. Language composition of training corpus The left figure shows the proportions of languages in the final corpus mixture, with the target languages of Norway in blue, related languages in red, and other data sources in gray. The right figure then displays the upsampling factors used to get the aforementioned proportions. In the following sections, we first describe the train￾ing corpus of NorMistral-11B in Section 2; … view at source ↗
Figure 2
Figure 2. Three-stage continual pretraining We propose a novel continual pretraining pipeline consisting of 1 creating a new tokenizer optimized for the training corpus, 2 realigning the embedding weights to the new tokens, and 3 training the full language model. Arrows symbolize changes between stages, while double-lines represent no changes. 2.3 Data sources Existing corpora We source most of the data from existing publicly… view at source ↗
Figure 3
Figure 3. Inference modes of NorMistral-11B The hybrid masked-causal pretraining allows the model to be more flexible during inference. It can not only serve as a unidirectional causal language model (left), but also as a fully bidirectional masked language model (middle), or as a partially bidirectional prefix language model (right). The diagrams illustrate possible attention connections. of causal models (Samuel, 2024). Fol… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-label Scandinavian Language Identification (SLIDE)

    cs.CL 2025-02 conditional novelty 6.0 of 10

    SLIDE introduces a manually labeled multi-label evaluation set and BERT/FastText models for Scandinavian language identification, using machine translation identity as a silver-labeling signal.

Reference graph

Works this paper leans on

65 extracted references · 18 canonical work pages · cited by 1 Pith paper

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023. https://openreview.net/forum?id=hmOwOZWzYE GQA : Training generalized multi-query transformer models from multi-head checkpoints . In The 2023 Conference on Empirical Methods in Natural Language Processing

  4. [4]

    Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. https://aclanthology.org/2024.acl-long.44 The belebele benchmark: a parallel reading comprehension dataset in 122 language variants . In Proceedings of the 62nd Annual Meeting of the ...

  5. [5]

    Adrien Barbaresi. 2021. https://doi.org/10.18653/v1/2021.acl-demo.15 Trafilatura: A web scraping library and command-line tool for text discovery and extraction . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, page...

  6. [6]

    Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. https://openreview.net/forum?id=IW1PR7vEBf LLM 2vec: Large language models are secretly powerful text encoders . In First Conference on Language Modeling

  7. [7]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...

  8. [8]

    Christopher Bryant, Mariano Felice, and Ted Briscoe. 2017. https://doi.org/10.18653/v1/P17-1074 Automatic annotation and evaluation of error types for grammatical error correction . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 793--805, Vancouver, Canada. Association for Computat...

Show all 65 references
  1. [9]

    Lucas Georges Gabriel Charpentier and David Samuel. 2024. http://arxiv.org/abs/2410.24159 Gpt or bert: why not both?

  2. [10]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning ...

  3. [11]

    Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. https://openreview.net/forum?id=dXiGWqBoxaD GPT 3.int8(): 8-bit matrix multiplication for transformers at scale . In Advances in Neural Information Processing Systems

  4. [12]

    Tim Dettmers and Luke Zettlemoyer. 2023. https://proceedings.mlr.press/v202/dettmers23a.html The case for 4-bit precision: k-bit inference scaling laws . In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Rese...

  5. [13]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...

  6. [14]

    Ethan Ewer, Daewon Chae, Thomas Zeng, Jinkyu Kim, and Kangwook Lee. 2024. http://arxiv.org/abs/2410.01600 ENTP : Encoder-only next token prediction

  7. [15]

    Philip Gage. 1994. https://api.semanticscholar.org/CorpusID:59804030 A new algorithm for data compression . The C Users Journal archive, 12:23--38

  8. [16]

    Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...

  9. [17]

    Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Ba \ n \'o n, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ram \' rez-S \'a nchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and J \"o rg Tiedemann. 2024. https://aclanthology.org/2024.lr...

  10. [18]

    Giellatekno and Divvun. 2016. https://doi.org/10.18710/8AK7KZ SIKOR North Saami corpus

  11. [19]

    Alexander H \"a gele, Elie Bakouch, Atli Kosson, Loubna Ben allal, Leandro Von Werra, and Martin Jaggi. 2024. https://openreview.net/forum?id=ompl7supoX Scaling laws and compute-optimal training beyond fixed training durations . In Workshop on Efficient Systems for Foundation ...

  12. [20]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Sim...

  13. [21]

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lo RA : Low-rank adaptation of large language models . In International Conference on Learning Representations

  14. [22]

    Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, Andr \'e Martins, Fran c ois Yvon, and Hinrich Sch \"u tze. 2023. https://doi.org/10.18653/v1/2023.acl-long.61 Glot500: Scaling multilingual ...

  15. [23]

    Sardana Ivanova, Fredrik Andreassen, Matias Jentoft, Sondre Wold, and Lilja vrelid. 2023. https://aclanthology.org/2023.nodalida-1.17 N or Q u AD : N orwegian question answering dataset . In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pag...

  16. [24]

    Matias Jentoft. 2023. https://www.duo.uio.no/handle/10852/103885 Grammatical error correction with byte-level language models . Master's thesis, University of Oslo

  17. [25]

    Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...

  18. [26]

    Amir Hossein Kargaran, Ayyoob Imani, Fran c ois Yvon, and Hinrich Schuetze. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.410 G lot LID : Language identification for low-resource languages . In Findings of the Association for Computational Linguistics: EMNLP 2023, page...

  19. [27]

    Per Kummervold, Freddy Wetjen, and Javier de la Rosa. 2022. https://aclanthology.org/2022.lrec-1.410 The N orwegian colossal corpus: A text corpus for training large N orwegian language models . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pag...

  20. [28]

    Per E Kummervold, Javier De la Rosa, Freddy Wetjen, and Svein Arne Brygfjeld. 2021. https://aclanthology.org/2021.nodalida-main.3 Operationalizing a national digital library: The case for a N orwegian transformer model . In Proceedings of the 23rd Nordic Conference on Computat...

  21. [29]

    Andrey Kutuzov, Jeremy Barnes, Erik Velldal, Lilja vrelid, and Stephan Oepen. 2021. https://aclanthology.org/2021.nodalida-main.4 Large-scale contextualised language modelling for N orwegian . In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa)...

  22. [30]

    Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics

  23. [31]

    Lauvrak, Jon Espen Ingvaldsen, Simen Eide, Jon Atle Gulla, and Zhirong Yang

    Peng Liu, Lemei Zhang, Terje Farup, Even W. Lauvrak, Jon Espen Ingvaldsen, Simen Eide, Jon Atle Gulla, and Zhirong Yang. 2024 a . http://arxiv.org/abs/2312.01314 NLEB ench+ N or GLM : A comprehensive empirical analysis and benchmark dataset for generative language models in N orwegian

  24. [32]

    Xiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu, and Dahua Lin. 2024 b . https://openreview.net/forum?id=JO7k0SJ5V6 Scaling laws of R o PE -based extrapolation . In The Twelfth International Conference on Learning Representations

  25. [33]

    Ilya Loshchilov and Frank Hutter. 2019. https://openreview.net/forum?id=Bkg6RiCqY7 Decoupled weight decay regularization . In International Conference on Learning Representations

  26. [34]

    Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Z...

  27. [35]

    Sheng Lu, Hendrik Schuff, and Iryna Gurevych. 2024. https://doi.org/10.18653/v1/2024.naacl-long.325 How are prompts different in terms of sensitivity? In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human La...

  28. [36]

    Risto Luukkonen, Jonathan Burdge, Elaine Zosa, Aarne Talman, Ville Komulainen, Vaino Hatanpaa, Peter Sarlin, and Sampo Pyysalo. 2024. https://api.semanticscholar.org/CorpusID:268857103 Poro 34b and the blessing of multilinguality . ArXiv, abs/2404.01856

  29. [37]

    Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki ...

  30. [38]

    Ang Lv, Kaiyi Zhang, Shufang Xie, Quan Tu, Yuhan Chen, Ji-Rong Wen, and Rui Yan. 2023. https://doi.org/10.48550/arXiv.2311.07468 Are we falling in a middle-intelligence trap? an analysis and mitigation of the reversal curse . CoRR, abs/2311.07468

  31. [39]

    Michael McCloskey and Neal J. Cohen. 1989. https://doi.org/https://doi.org/10.1016/S0079-7421(08)60536-8 Catastrophic interference in connectionist networks: The sequential learning problem . In Gordon H. Bower, editor, Psychology of Learning and Motivation, volume 24 of Psych...

  32. [40]

    Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023. https://openreview.net/forum?id=j5BuTrEj35 Scaling data-constrained language models . In Thirty-seventh Conference on Neural I...

  33. [41]

    Rossi, and Thien Huu Nguyen

    Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2023. http://arxiv.org/abs/2309.09400 Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages

  34. [42]

    Nguyen and Julian Salazar

    Toan Q. Nguyen and Julian Salazar. 2019. https://aclanthology.org/2019.iwslt-1.17 Transformers without tears: Improving the normalization of self-attention . In Proceedings of the 16th International Conference on Spoken Language Translation, Hong Kong. Association for Computat...

  35. [43]

    Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman

    Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Haji c , Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016. https://aclanthology.org/L16-1262 U niversal D ependencies v1: A mul...

  36. [44]

    Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman

    Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Haji c , Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020. https://aclanthology.org/2020.lrec-1.497 U niversal D ependencies v2: An evergrowing multilingual treebank co...

  37. [45]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...

  38. [46]

    Ronny Paul, Himanshu Buckchash, Shantipriya Parida, and Dilip K. Prasad. 2024. https://api.semanticscholar.org/CorpusID:269635411 Towards a more inclusive ai: Progress and perspectives in large language model training for the s \'a mi language . ArXiv, abs/2405.05777

  39. [47]

    Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. http://arxiv.org/abs/2406.17557 The F ine W eb datasets: Decanting the web for the finest text data at scale

  40. [48]

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. https://openai.com/index/language-unsupervised/ Improving language understanding by generative pre-training . OpenAI Blog

  41. [49]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1)

  42. [50]

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. https://dl.acm.org/doi/10.5555/3433701.3433727 ZeRO : memory optimizations toward training trillion parameter models . In Proceedings of the International Conference for High Performance Computing, Network...

  43. [51]

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. https://doi.org/10.1145/3394486.3406703 Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters . In Proceedings of the 26th ACM SIGKDD International Conferenc...

  44. [52]

    David Samuel. 2024. https://openreview.net/forum?id=BCA9NMZkLS BERT s are generative in-context learners . In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  45. [53]

    David Samuel, Andrey Kutuzov, Samia Touileb, Erik Velldal, Lilja vrelid, Egil R nningstad, Elina Sigdel, and Anna Palatkina. 2023. https://aclanthology.org/2023.nodalida-1.61 N or B ench -- a benchmark for N orwegian language models . In Proceedings of the 24th Nordic Conferen...

  46. [54]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili'c, Daniel Hesslow, Roman Castagn'e, Alexandra Sasha Luccioni, François Yvon, Matthias Gall \'e , Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Be...

  47. [55]

    Noam M. Shazeer. 2020. https://api.semanticscholar.org/CorpusID:211096588 GLU variants improve transformer . ArXiv, abs/2002.05202

  48. [56]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. http://arxiv.org/abs/1909.08053 Megatron- LM : T raining multi-billion parameter language models using model parallelism

  49. [57]

    Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021. http://arxiv.org/abs/2104.09864 Roformer: Enhanced transformer with rotary position embedding

  50. [58]

    J \"o rg Tiedemann. 2020. https://aclanthology.org/2020.wmt-1.139 The tatoeba translation challenge -- realistic data sets for low resource and multilingual MT . In Proceedings of the Fifth Conference on Machine Translation, pages 1174--1182, Online. Association for Computatio...

  51. [59]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://api.semanticscholar.org...

  52. [60]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf Attention is all you need . In Advances in Ne...

  53. [61]

    Erik Velldal, Lilja vrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, and Fredrik J rgensen. 2018. https://aclanthology.org/L18-1661 N o R e C : The N orwegian review corpus . In Proceedings of the Eleventh International Conference on Language Resources and Ev...

  54. [62]

    Wietse de Vries and Malvina Nissim. 2021. https://doi.org/10.18653/v1/2021.findings-acl.74 As good as new. how to successfully recycle E nglish GPT -2 to make models for other languages . In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 836-...

  55. [63]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...

  56. [64]

    Zheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, Genta Winata, Stella Biderman, Edward Raff, Dragomir Radev, and Vassilina Nikoulina. 2023. https://doi....

  57. [65]

    Biao Zhang and Rico Sennrich. 2019. https://openreview.net/references/pdf?id=S1qBAf6rr Root Mean Square Layer Normalization . In Advances in Neural Information Processing Systems 32, Vancouver, Canada

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.