Pith. sign in

REVIEW 3 major objections 6 minor 95 references

SnakModel: Lessons Learned from Training an Open Danish Large Language Model

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read By continuously pre-training a 7-billion-parameter English model on 13.6 billion curated Danish words and tuning it on 3.7 million Danish instructions, this paper produces a Danish LLM that beats every other Llama-2-7B-based model on the…

desk verdict Solid open Danish LLM and corpus with honest leakage disclosure, but the headline average rests on unmeasured overlap with two benchmark tasks; needs a clean-subset re-evaluation. read the letter →

arxiv 2412.12956 v1 pith:OTX56KZQ submitted 2024-12-17 cs.CL

classification cs.CL
keywords Danishlanguagemodelcontinuouspre-traininginstructiontuningScandEvallow-resourceNLPtrainingdynamicsweightdivergence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that a strong Danish large language model can be built by taking an existing English 7B model, continuing its training on a heavily curated, deduplicated corpus of 13.6B Danish words, and then instruction-tuning on 3.7M Danish instruction–answer pairs. It reports that the resulting model, SnakModel-7Binstruct, reaches an average score of 56.63 on the Danish tasks of ScandEval, beating every other Llama-2-7B-based system, including ones given the same Danish instruction data. The authors also claim that most downstream gains appear within the first 2,000–5,000 continued-pretraining steps, that one epoch of instruction tuning is enough, and that parameter change concentrates in the embedding layer, the SwiGLU up-projection, and the language-model head. A sympathetic reader would care because the paper turns these observations into concrete training guidelines for small language communities with limited compute.

What carries the argument

The central mechanism is continued pre-training on a curated Danish corpus, evaluated through two diagnostic lenses: intermediate checkpoints scored on ScandEval to trace when Danish competence emerges, and principal subspace angles (a measure of parameter change) computed before and after adaptation to locate where it happens. The continuation trains Llama 2-7B for 12,500 steps on 13.6B words with a low peak learning rate of $1.5 \times 10^{-5}$ to avoid gradient explosions, followed by one epoch of LoRA instruction tuning with rank 128. The subspace-angle analysis shows that most change concentrates in the embedding, the SwiGLU gate and up-projection, and the language-model head, with self-attention barely moving.

What would settle it

Compare the test splits of DANSK NER and ScaLA against the SnakModel pre-training corpus: if a substantial fraction of test examples appears verbatim in the training data, re-evaluate on strictly disjoint versions; the central claim of superior generalization would collapse if the NER and LA gains disappear on those clean splits.

Watch

Extended reading notes

Core claim

The central claim is that continuous pre-training of Llama 2-7B on a large native Danish corpus, followed by instruction tuning, yields the best Llama-2-7B-based model for Danish, with an average ScandEval score of 56.63 versus 46.48 for the English base and 49.50 for the chat version. The largest gains appear on tasks built from natural Danish text rather than translations: linguistic acceptability rises from 33.43 to 52.91, proverb understanding from 38.69 to 71.05, and citizenship-test accuracy from 57.05 to 71.88. The paper also shows that continued pre-training temporarily degrades named-entity recognition and question answering, but that instruction tuning recovers both, and that close-to-final performance is reached after only a fraction of the pre-training corpus has been seen.

Load-bearing premise

The evaluation is only trustworthy as a measure of generalization if the documented overlap between the training corpus and the ScandEval data—the DANSK NER set was included verbatim and parts of ScaLA were included—does not inflate the reported NER and LA scores through memorization.

Editorial extensions

If this is right

  • Danish language modeling can start from an existing English 7B model plus a ~13.6B-word curated corpus instead of training from scratch.
  • Instruction tuning after 2,000–5,000 steps of continued pre-training may reach near-final performance, so expensive long pre-training can be truncated.
  • A single epoch of instruction tuning on translated instruction data is enough to restore and amplify instruction-following behavior in the target language.
  • Because adaptation concentrates in the embedding, SwiGLU up-projection, and language-model head, future Danish-adaptation runs could train only those parameters for similar efficiency gains.
  • Native-language data, rather than translated data, is what drives gains on culture-specific tasks such as proverb understanding and citizenship questions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the documented leakage is substantial, the reported NER and LA improvements may overstate true generalization; a clean held-out evaluation would put a realistic bound on the adaptation gain.
  • The parameter-localization result suggests that for typologically close languages, parameter-efficient targeting of embeddings and feed-forward layers could approximate full continued pre-training at a fraction of the cost—an extension the paper does not directly test.
  • The 2,000–5,000-step saturation point raises the question of whether even smaller corpora (a few billion words) would suffice for Danish, which could be tested by pre-training on progressively smaller slices.
  • The recipe may transfer to other Germanic mid-resource languages, but for typologically distant targets self-attention may need more updating, so the weight-divergence pattern should be re-measured per language family.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents SnakModel, a Danish LLM created by continuously pre-training Llama2-7B on 13.6B Danish words and then instruction-tuning on 3.7M Danish instructions. It evaluates the model on the Danish portion of the ScandEval benchmark across eight tasks, compares it with contemporary Llama2-7B- and Mistral-7B-based models, analyzes intermediate training dynamics and weight divergence, and distills recommendations for adapting LLMs to lower-resource languages. The central claim is that SnakModel-7Binstruct outperforms all other Llama2-7B-based models, including those trained on the same Danish instruction data, with an average score of 56.63 (Table 3).

Significance. The paper is a genuinely useful resource paper: it releases model weights, intermediate checkpoints, data-collection scripts, and evaluation code, and the +INSTda ablation is a well-designed control that isolates the effect of Danish pre-training from instruction tuning. The training-dynamics and weight-divergence analyses are informative and go beyond a single leaderboard. If the evaluation is confirmed to be clean, the guidance for mid-resource languages (e.g., one epoch of instruction tuning, focus on embedding/feed-forward updates) would be a valuable contribution.

major comments (3)
  1. [Section 4.1 (Leakage) and Section 5, Table 3] The paper acknowledges in Section 4.1 that the DANSK NER dataset was "completely included (without labels)" in the pre-training corpus and that "many parts" of the ScaLA dataset were included in original form, with 6/200 AngryTweets tweets also found. Because the ScandEval test split used in Table 3 is built from these datasets, the reported NER and LA scores, and to a lesser degree SENTI, may reflect memorization of input text rather than generalization. The paper never quantifies the fraction of the actual test splits present in the deduplicated pre-training corpus, and it never recomputes the benchmark averages on overlap-free subsets. Since the largest margin over the strongest +INSTda baseline is on LA (+9.51) and the NER margin is only +0.06, the headline claim that SNAKMODEL outperforms all other Llama2-7B-based models is not yet supported on the contaminated tasks. Please report per-task overlap counts (exact and near-duplicate) and re-run Table 3 on clean subsets, or otherwise show that the 56.63 average and the rank ordering are robust to excluding contaminated instances.
  2. [Section 5, Table 3] All benchmark scores are reported for a single run, with no variance estimates or significance tests. Several task-level margins over the second-best Llama2-7B model are below one percentage point (NER +0.06, QA +0.26, SENTI +0.78), so the claim of "outperforms all other LLAMA 2-7B-based models" rests on differences that could easily be within run-to-run noise. I ask for at least bootstrap confidence intervals on the average, and ideally repeated instruction-tuning runs for the main comparison, to make the central claim statistically grounded.
  3. [Section 4.3, Figure 2b, and Section 6] The training-dynamics conclusion that "instruction tuning after 2,000–5,000 steps ... may already be sufficient to obtain close-to-final performance" is drawn from validation scores on the same contaminated tasks (LA and NER). If the LA and NER gains are partly due to the model having seen the test text during pre-training, the plateau pattern in Figure 2b may be an artifact of memorization rather than a genuine signal about when Danish pre-training becomes useful. Please verify the plateau on clean validation subsets before presenting this as guidance for future work.
minor comments (6)
  1. [Section 4.1] The leakage test samples "200 random 8-grams from each of our datasets" but does not specify whether the search was run against the pre-training corpus after deduplication or before it; please clarify, and also report exact whole-sequence overlaps with the test splits rather than only 8-gram hits.
  2. [Section 4.1] The statement that "many parts" of ScaLA were included in original form would be much more informative with a quantitative count or a proportion; as written, the reader cannot gauge the severity of the overlap.
  3. [Table 3] The "AVG." column appears to be an unweighted average of eight task scores with different metrics; please state this explicitly and consider adding a note on how the average behaves if the contaminated tasks are removed.
  4. [Section 4.2] LoRA rank 128 is described as applied to "all parameters within the model," but it is not stated whether this includes the input embeddings and the language-modeling head; this matters for the weight-divergence interpretation in Figure 4.
  5. [Figures 1 and 2] The captions use "SNAK MODEL" with a space; please use "SNAKMODEL" consistently to match the rest of the paper.
  6. [Figure 2] The dashed line in Figure 2 is not explained in the caption; please identify what it represents.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the central benchmark claim is an external empirical measurement, and the Section 4.1 leakage disclosure is a data-contamination caveat rather than a circular derivation.

full rationale

The paper's main claim, that SNAKMODEL-7Binstruct outperforms other LLAMA 2-7B-based models on the Danish part of ScandEval, is an empirical comparison of trained checkpoints against external benchmark tasks; no performance value in Table 3 is derived by definition from a training objective or from a fitted parameter. The rubric's circularity patterns do not apply: there is no self-definitional relation between the pretraining corpus and the reported metric, no fitted input is renamed as a prediction, and no 'uniqueness theorem' is imported from the authors' prior work. The paper does contain self-citations (Müller-Eberstein et al., 2024, for principal subspace angles; Singh et al., 2024, for the Aya collection, on which one of the present authors is a co-author), but these are measurement-method and data-provenance attributions, not load-bearing support for the benchmark outcome. The only passage that could be mistaken for circularity is the Section 4.1 leakage check, which states: 'The DANSK NER dataset was completely included (without labels), as it was sampled from Gigaword, and many parts of the ScaLA dataset were also included in its original form in GigaWord and CC100.' This is a data-contamination disclosure: it raises a legitimate validity concern about whether those two task scores are inflated by memorization, and a clean-subset re-evaluation would strengthen the paper. However, it is not a circular step in the derivation sense, because the benchmark scores are measured outputs of the trained model, not inputs to its construction or to the paper's reasoning. The final average 56.63 is reported transparently with per-task scores, and the comparison to baselines is external; the paper's own disclosure makes the limitation visible rather than hiding it. Under the stated rubric, which requires exhibiting a specific reduction such as Eq. X = Eq. Y by construction or a fitted parameter renamed as a prediction, no such reduction is present. Score 0 reflects the absence of derivation circularity; the contamination caveat is a separate empirical-validity issue, not a circularity one.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The paper's central result is empirical and rests on hyperparameter choices, benchmark validity, and data curation decisions rather than on new theoretical entities. The main burden is the acknowledged benchmark contamination (DANSK NER and ScaLA overlap), tracked under red flags.

free parameters (6)
  • Pre-training peak learning rate = 1.5e-5
    Chosen after two failed runs at higher learning rates that caused gradient explosions; directly affects convergence and final loss.
  • Pre-training context length = 4096
    Standard for Llama2, chosen without ablation and retained throughout.
  • Global batch size (pre-training) = 512
    Chosen for stability on four A100 GPUs; not ablated.
  • Instruction tuning LoRA rank = 128
    Chosen to approximate full fine-tuning; directly affects instruction tuning capacity.
  • Instruction tuning learning rate = 2e-4
    Constant learning rate for LoRA; not ablated against other values.
  • fastText language identification threshold = 0.6
    Used in data filtering; removes 28% of documents and affects corpus composition.
assumptions (5)
  • domain assumption ScandEval is a valid and sufficient benchmark for Danish LLM evaluation.
    All model comparisons and design decisions are judged on the eight ScandEval tasks; the paper does not test on other Danish benchmarks.
  • domain assumption Llama2's SentencePiece tokenizer with 32K subwords is adequate for Danish without retraining or extension.
    Section 4.1 states the assumption that Danish and English share sufficient vocabulary overlap; no tokenizer analysis is provided.
  • domain assumption The leakage tests are sufficient to determine benchmark contamination.
    Section 4.1 checks 200 random 8-grams per dataset and prompts Llama2, but the acknowledged DANSK NER and ScaLA overlaps show the tests are not fully sufficient.
  • domain assumption Automatically translated instruction data is of sufficient quality for Danish instruction tuning.
    Section 3.2 relies on manual inspection and selection of translated datasets; no human evaluation of translation quality is reported.
  • standard math Standard deep learning optimizer and scaling assumptions hold.
    Training uses AdamW, gradient clipping, and standard transformer practices without proof of optimality; acceptable as background.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SnakModel: Lessons Learned from Training an Open Danish Large Language Model." pith.science (2026). https://pith.science/paper/OTX56KZQ

@misc{pith2026241212956,
  author       = {Pith},
  title        = {Pith review of: SnakModel: Lessons Learned from Training an Open Danish Large Language Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OTX56KZQ}},
  note         = {Machine review of arXiv:2412.12956}
}
read the original abstract

We present SnakModel, a Danish large language model (LLM) based on Llama2-7B, which we continuously pre-train on 13.6B Danish words, and further tune on 3.7M Danish instructions. As best practices for creating LLMs for smaller language communities have yet to be established, we examine the effects of early modeling and training decisions on downstream performance throughout the entire training pipeline, including (1) the creation of a strictly curated corpus of Danish text from diverse sources; (2) the language modeling and instruction-tuning training process itself, including the analysis of intermediate training dynamics, and ablations across different hyperparameters; (3) an evaluation on eight language and culturally-specific tasks. Across these experiments SnakModel achieves the highest overall performance, outperforming multiple contemporary Llama2-7B-based models. By making SnakModel, the majority of our pre-training corpus, and the associated code available under open licenses, we hope to foster further research and development in Danish Natural Language Processing, and establish training guidelines for languages with similar resource constraints.

Figures

Figures reproduced from arXiv: 2412.12956 by the authors.

Figure 1
Figure 1. SNAKMODEL-7Bbase Pre-training Be￾haviour. We report the stable language model loss during training and validation. ter sample (without labels). The DANSK NER dataset was completely included (without labels), as it was sampled from Gigaword, and many parts of the ScaLA dataset were also included in its orig￾inal form in GigaWord and CC100. The code for all leakage tests is included in our code repository. 4.2 Instruc… view at source ↗
Figure 2
Figure 2. SNAKMODEL Training Dynamics of LM pre-training, instruction tuning, and multi-epoch instruction tuning, as measured on the ScandEval (validation) tasks of linguistic acceptability (LA), named entity recognition (NER), sentiment analysis (SENTI), summarization (SUMM), commonsense reasoning (CSR), question answering (QA), proverb meaning (TM), and citizenship tests (CT). delimiters following LLAMA2-7Bchat 12; (3) ALPA… view at source ↗
Figure 4
Figure 4. Parameter-wise Weight Divergence of SNAKMODEL-7Bbase as measured in mean SSA. Darker bars represent EMB and LMH respectively. the model, and subsequent target language special￾ization in later layers (Wendler et al., 2024) [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 34 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024. https://arxiv.org/abs/2404.14219 Phi-3 technical report: A highly capable language model locally on your phone . ArXiv preprint, abs/2404.14219

  4. [4]

    AI-Sweden. 2024. https://huggingface.co/AI-Sweden-Models/Llama-3-8B Ai-sweden-models/llama-3-8b

  5. [5]

    Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019. https://doi.org/10.18653/v1/W19-1909 Publicly available clinical BERT embeddings . In Proceedings of the 2nd Clinical Natural Language Processing Workshop, pages 72--78, Minneapolis, Minnesota, USA. Association for Computational Linguistics

  6. [6]

    Duarte M Alves, Jos \'e Pombal, Nuno M Guerreiro, Pedro H Martins, Jo \ a o Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, et al. 2024. https://arxiv.org/abs/2402.17733 Tower: An open multilingual large language model for translation-related tasks . ArXiv preprint, abs/2402.17733

  7. [7]

    Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, et al. 2024. https://arxiv.org/abs/2405.15032 Aya 23: Open weight releases to further multilingual progress . ArXiv preprint, abs/2405.15032

  8. [8]

    Giuseppe Attardi. 2015. Wikiextractor. https://github.com/attardi/wikiextractor

Show all 95 references
  1. [9]

    Andrea Bacciu, Cesare Campagnano, Giovanni Trappolini, and Fabrizio Silvestri. 2024. https://aclanthology.org/2024.lrec-main.388 D ante LLM : Let ' s push I talian LLM research forward! In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lan...

  2. [10]

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. https://arxiv.org/abs/2309.16609 Qwen technical report . ArXiv preprint, abs/2309.16609

  3. [11]

    Simone Balloccu, Patr \' cia Schmidtov \'a , Mateusz Lango, and Ondrej Dusek. 2024. https://aclanthology.org/2024.eacl-long.5 Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLM s . In Proceedings of the 18th Conference of the European Chap...

  4. [12]

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023. Pythia: A suite for analyzing large language models across training and scaling. In ...

  5. [13]

    Alejandro Hernández Cano, Matteo Pagliardini, Andreas Köpf, Kyle Matoba, Amirkeivan Mohtashami, Xingyao Wang, Olivia Simin Fan, Axel Marmet, Deniz Bayazit, Igor Krawczuk, Zeming Chen, Francesco Salvi, Antoine Bosselut, and Martin Jaggi. 2023. https://github.com/epfLLM/Megatron...

  6. [14]

    Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023. https://doi.org/10.18653/v1/2023.c3nlp-1.7 Assessing cross-cultural alignment between C hat GPT and human societies: An empirical study . In Proceedings of the First Workshop on Cross-Cultur...

  7. [15]

    Chang, Caleb Chiam, Liye Fu, Andrew Wang, Justine Zhang, and Cristian Danescu-Niculescu-Mizil

    Jonathan P. Chang, Caleb Chiam, Liye Fu, Andrew Wang, Justine Zhang, and Cristian Danescu-Niculescu-Mizil. 2020. https://doi.org/10.18653/v1/2020.sigdial-1.8 C onvo K it: A toolkit for the analysis of conversations . In Proceedings of the 21th Annual Meeting of the Special Int...

  8. [16]

    Christiansen, Kristian Tyl \'e n, Riccardo Fusaroli, Dorthe Bleses, Anders H jen, Fabio Trecca, Christina Dideriksen, and Byurakn Ishkhanyan

    Morten H. Christiansen, Kristian Tyl \'e n, Riccardo Fusaroli, Dorthe Bleses, Anders H jen, Fabio Trecca, Christina Dideriksen, and Byurakn Ishkhanyan. 2023. https://projects.au.dk/the-puzzle-of-danish The puzzle of danish

  9. [17]

    Ciosici and Leon Derczynski

    Manuel R. Ciosici and Leon Derczynski. 2022. https://arxiv.org/abs/2208.12097 Training a t5 using lab-sized resources

  10. [18]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning ...

  11. [19]

    Marta R Costa-juss \`a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022. https://arxiv.org/abs/2207.04672 No language left behind: Scaling human-centered machine translation . ArX...

  12. [20]

    John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, et al. 2024. https://arxiv.org/abs/2412.04261 Aya expanse: Combining research breakthroughs for a new multilingual fr...

  13. [21]

    Danish-Foundation-Models-Team. 2024. http://www.foundationmodels.dk/blog/2024/01/11/releasing-munin-7b-alpha---a-danish-llm/ Releasing munin 7b alpha - a danish llm

  14. [22]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/1feb87871436031bdc0f2beaa62a049b-Paper-Conference.pdf Qlora: Efficient finetuning of quantized llms . In Advances in Neural Information Processi...

  15. [23]

    Christiansen, Mark Dingemanse, Malte H jmark-Bertelsen, Christer Johansson, Kristian Tyl \'e n, and Riccardo Fusaroli

    Christina Dideriksen, Morten H. Christiansen, Mark Dingemanse, Malte H jmark-Bertelsen, Christer Johansson, Kristian Tyl \'e n, and Riccardo Fusaroli. 2023. https://doi.org/10.1111/cogs.13387 Language-specific constraints on conversation: Evidence from danish and norwegian . C...

  16. [24]

    Longxu Dou, Qian Liu, Guangtao Zeng, Jia Guo, Jiahui Zhou, Wei Lu, and Min Lin. 2024. https://arxiv.org/abs/2404.03608 Sailor: Open language models for south-east asia . ArXiv preprint, abs/2404.03608

  17. [25]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. https://arxiv.org/abs/2407.21783 The llama 3 herd of models . ArXiv preprint, abs/2407.21783

  18. [26]

    Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk, Joey \"O hman, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Judit Casademont, and Magnus Sahlgren. 2024. https://aclanthology.org/2024.lrec-main.695 GPT - SW 3: An autoregressive language model for the S candinavia...

  19. [27]

    Kenneth Enevoldsen, Lasse Hansen, and Kristoffer Nielbo. 2021. https://arxiv.org/abs/2107.05295 Dacy: A unified framework for danish nlp . ArXiv preprint, abs/2107.05295

  20. [28]

    Julen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, and Aitor Soroa. 2024. https://aclanthology.org/2024.acl-long.799 Latxa: An open language model and evaluation suite for B asque . In Proceedings of the 62nd A...

  21. [29]

    Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tusha...

  22. [30]

    Suchin Gururangan, Ana Marasovi \'c , Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.acl-main.740 Don ' t stop pretraining: Adapt language models to domains and tasks . In Proceedings of the 58th Annual Meeting o...

  23. [31]

    Xiaochuang Han and Jacob Eisenstein. 2019. https://doi.org/10.18653/v1/D19-1433 Unsupervised domain adaptation of contextualized embeddings for sequence labeling . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Internation...

  24. [32]

    Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders S gaard. 2022. https://doi.org/10.186...

  25. [33]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lora: Low-rank adaptation of large language models . In The Tenth International Conference on Learning Representat...

  26. [34]

    Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted, Lasse Malm Lidegaard, and Anders S gaard. 2020. https://aclanthology.org/2020.lrec-1.565 D a NE : A named entity resource for D anish . In Proceedings of the Twelfth Language Resources and Evaluation Con...

  27. [35]

    Jaavid J, Raj Dabre, Aswanth M, Jay Gala, Thanmay Jayakumar, Ratish Puduppully, and Anoop Kunchukuttan. 2024. https://aclanthology.org/2024.acl-long.833 R oman S etu: Efficiently unlocking multilingual capabilities of large language models via R omanization . In Proceedings of...

  28. [36]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. https://arxiv.org/abs/2310.06825 Mistral 7b . ArXiv preprint, abs/2310.06825

  29. [37]

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. https://doi.org/10.18653/v1/2020.acl-main.560 The state and fate of linguistic diversity and inclusion in the NLP world . In Proceedings of the 58th Annual Meeting of the Association for Co...

  30. [38]

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. https://aclanthology.org/E17-2068 Bag of tricks for efficient text classification . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational Linguistics: Volume ...

  31. [39]

    Oliver Kinch. 2023. https://huggingface.co/datasets/alexandrainst/nordjylland-news-summarization Nordjylland news summarization . Accessed on 22-08-2024

  32. [40]

    Andreas Kirkedal, Marija Stepanovic, and Barbara Plank. 2020. https://doi.org/10.21437/Interspeech.2020-3164 FT speech: Danish parliament speech corpus . In Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai,...

  33. [41]

    Andrew V Knyazev and Merico E Argentati. 2002. Principal angles between subspaces in an A -based scalar product: algorithms and perturbation estimates. SIAM Journal on Scientific Computing, 23(6):2008--2040

  34. [42]

    Matthias Trautner Kromann and Stine Kern Lynge. 2004. The danish dependency treebank v. 1.0

  35. [43]

    Taku Kudo and John Richardson. 2018. https://doi.org/10.18653/v1/D18-2012 S entence P iece: A simple and language independent subword tokenizer and detokenizer for neural text processing . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processin...

  36. [44]

    Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. 2019. https://arxiv.org/abs/1910.09700 Quantifying the carbon emissions of machine learning . ArXiv preprint, abs/1910.09700

  37. [45]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili \'c , Daniel Hesslow, Roman Castagn \'e , Alexandra Sasha Luccioni, Fran c ois Yvon, Matthias Gall \'e , et al. 2023. Bloom: A 176b-parameter open-access multilingual language model

  38. [46]

    Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234--1240

  39. [47]

    LeoLM-Team. 2024. https://laion.ai/blog/leo-lm/ Leolm: Igniting german-language llm research

  40. [48]

    Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and ``Teknium''. 2023. Openorca: An open dataset of gpt augmented flan reasoning traces. https://https://huggingface.co/Open-Orca/OpenOrca

  41. [49]

    Pierre Lison and J \"o rg Tiedemann. 2016. https://aclanthology.org/L16-1147 O pen S ubtitles2016: Extracting large parallel corpora from movie and TV subtitles . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC '16) , pages 923-...

  42. [50]

    Hong Liu, Sang Michael Xie, Zhiyuan Li, and Tengyu Ma. 2023. Same pre-training loss, better downstream: Implicit bias matters for language models. In International Conference on Machine Learning, pages 22188--22214. PMLR

  43. [51]

    Shayne Longpre, Yi Lu, and Joachim Daiber. 2021. https://doi.org/10.1162/tacl_a_00433 MKQA : A linguistically diverse benchmark for multilingual open domain question answering . Transactions of the Association for Computational Linguistics, 9:1389--1406

  44. [52]

    Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki ...

  45. [53]

    Magnus Mabeck. 2024. Danish openhermes. https://huggingface.co/datasets/Mabeck/danish-OpenHermes

  46. [54]

    Guerreiro, Ricardo Rei, Duarte M

    Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. 2024. https://a...

  47. [55]

    Chenghao Mou, Chris Ha, Kenneth Enevoldsen, and Peiyuan Liu. 2023. https://doi.org/10.5281/zenodo.8364980 Chenghaomou/text-dedup: Reference snapshot

  48. [56]

    Max Müller-Eberstein, Dianna Yee, Karren Yang, Gautam Varma Mantena, and Colin Lea. 2024. https://doi.org/10.1162/tacl_a_00696 Hypernetworks for Personalizing ASR to Atypical Speech . Transactions of the Association for Computational Linguistics, 12:1182--1196

  49. [57]

    Tarek Naous, Michael Ryan, Alan Ritter, and Wei Xu. 2024. https://aclanthology.org/2024.acl-long.862 Having beer after prayer? measuring cultural bias in large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume ...

  50. [58]

    Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. 2020. https://doi.org/10.18653/v1/2020.emnlp-demos.2 BERT weet: A pre-trained language model for E nglish tweets . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, ...

  51. [59]

    Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A Rossi, and Thien Huu Nguyen. 2023. https://arxiv.org/abs/2309.09400 Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages . ArXiv pr...

  52. [60]

    Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Zhiqiang Hu, Chenhui Shen, Yew Ken Chia, Xingxuan Li, Jianyu Wang, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, Hang Zhang, and Lidong Bing. 2024. https://aclanthology.org/2024.acl-demos.28 ...

  53. [61]

    Dan Nielsen. 2023. https://aclanthology.org/2023.nodalida-1.20 S cand E val: A benchmark for S candinavian natural language processing . In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pages 185--201, T \'o rshavn, Faroe Islands. Universit...

  54. [62]

    Dan Saattrup Nielsen. 2024. https://huggingface.co/datasets/alexandrainst/danish-citizen-tests Danish citizen test . Accessed on 22-08-2024

  55. [63]

    NORA.LLM-Team. 2024. https://huggingface.co/norallm/normistral-7b-warm Instruction-tuned normistral-7b-warm . Accessed on 17-10-2024

  56. [64]

    Amalie Brogaard Pauli, Maria Barrett, Oph \'e lie Lacroix, and Rasmus Hvingelby. 2021. https://aclanthology.org/2021.nodalida-main.53 D a NLP : An open-source toolkit for D anish natural language processing . In Proceedings of the 23rd Nordic Conference on Computational Lingui...

  57. [65]

    Københavns Professionshøjskole. 2024. Skolegpt instruct. https://huggingface.co/datasets/kobprof/skolegpt-instruct

  58. [66]

    Giovanni Puccetti, Anna Rogers, Chiara Alzetta, Felice Dell ' Orletta, and Andrea Esuli. 2024. https://aclanthology.org/2024.acl-long.817 AI ` news ' content farms are easy to make and hard to detect: A case study in I talian . In Proceedings of the 62nd Annual Meeting of the ...

  59. [67]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. http://jmlr.org/papers/v21/20-074.html Exploring the limits of transfer learning with a unified text-to-text transformer . J. Mach. Learn. Res., ...

  60. [68]

    Rakuten Group , Aaron Levine, Connie Huang, Chenguang Wang, Eduardo Batista, Ewa Szymanska, Hongyi Ding, Hou Wei Chou, Jean-Fran c ois Pessiot, Johanes Effendi, et al. 2024. https://arxiv.org/abs/2403.15484 Rakutenai-7b: Extending large language models for japanese . ArXiv pre...

  61. [69]

    Edwin Rijgersberg and Bob Lucassen. 2023. https://github.com/Rijgersberg/GEITje Geitje: een groot open nederlands taalmodel

  62. [70]

    Oscar Sainz, Jon Ander Campos, García-Ferrero Iker, Julen Etxaniz, and Eneko Agirre. 2023. https://hitz-zentroa.github.io/lm-contamination/blog/ Did ChatGPT cheat on your test?

  63. [71]

    Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. https://openreview.net/forum?id=RIu5lyNXjT Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting . In The Twelfth International...

  64. [72]

    Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. https://doi.org/10.18653/v1/P16-1162 Neural machine translation of rare words with subword units . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages ...

  65. [73]

    Noam Shazeer. 2020. https://arxiv.org/abs/2002.05202 Glu variants improve transformer . ArXiv preprint, abs/2002.05202

  66. [74]

    SiloAI. 2024. https://www.silo.ai/blog/viking-7b-13b-33b-sailing-the-nordic-seas-of-multilinguality Viking 7b/13b/33b: Sailing the nordic seas of multilinguality

  67. [75]

    Shivalika Singh, Freddie Vargus, Daniel D ' souza, B \"o rje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O ' Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemi \'n ski, Haki...

  68. [76]

    Raivis Skadi n s , J \"o rg Tiedemann, Roberts Rozis, and Daiga Deksne. 2014. http://www.lrec-conf.org/proceedings/lrec2014/pdf/846_Paper.pdf Billions of parallel words for free: Building and using the EU bookshop corpus . In Proceedings of the Ninth International Conference o...

  69. [77]

    Leon Str mberg-Derczynski, Manuel Ciosici, Rebekah Baglini, Morten H. Christiansen, Jacob Aarup Dalsgaard, Riccardo Fusaroli, Peter Juel Henrichsen, Rasmus Hvingelby, Andreas Kirkedal, Alex Speed Kjeldsen, Claus Ladefoged, Finn rup Nielsen, Jens Madsen, Malte Lau Petersen, Jon...

  70. [78]

    Ǎguila Team. 2023. https://huggingface.co/projecte-aina/aguila-7b Introducing aguila, a new open-source llm for spanish and catalan

  71. [79]

    ``Teknium''. 2023. O pen H ermes dataset. https://huggingface.co/datasets/teknium/openhermes

  72. [80]

    J \"o rg Tiedemann. 2012. http://www.lrec-conf.org/proceedings/lrec2012/pdf/463_Paper.pdf Parallel data, tools and interfaces in OPUS . In Proceedings of the Eighth International Conference on Language Resources and Evaluation ( LREC '12) , pages 2214--2218, Istanbul, Turkey. ...

  73. [81]

    J \"o rg Tiedemann. 2016. https://aclanthology.org/L16-1559 Finding alternative translations in a large corpus of movie subtitle . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC '16) , pages 3518--3522, Portoro z , Slovenia. Eu...

  74. [82]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. https://arxiv.org/abs/2307.09288 Llama 2: Open foundation and fine-tuned chat models . ArXiv preprint, abs/2...

  75. [83]

    Christiansen

    Fabio Trecca, Kristian Tyl \'e n, Anders H jen, and Morten H. Christiansen. 2021. https://doi.org/10.1111/lang.12450 Danish as a window onto language processing and learning . Language Learning, 71(3):799--833

  76. [84]

    Ahmet \"U st \"u n, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D ' souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. ht...

  77. [85]

    Bram Vanroy. 2024. https://huggingface.co/BramVanroy/fietje-2 Fietje 2: An open and efficient llm for dutch

  78. [86]

    Daniel Varab and Natalie Schluter. 2020. https://aclanthology.org/2020.lrec-1.831 D a N ewsroom: A large-scale D anish summarisation dataset . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6731--6739, Marseille, France. European Language Res...

  79. [87]

    Leon Voukoutis, Dimitris Roussis, Georgios Paraskevopoulos, Sokratis Sofianopoulos, Prokopis Prokopidis, Vassilis Papavasileiou, Athanasios Katsamanis, Stelios Piperidis, and Vassilis Katsouros. 2024. https://arxiv.org/abs/2407.20743 Meltemi: The first open large language mode...

  80. [88]

    Wenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai, Jen-tse Huang, Zhaopeng Tu, and Michael Lyu. 2024. https://aclanthology.org/2024.acl-long.345 Not all countries celebrate thanksgiving: On the cultural dominance in large language models . In Proceedings of the 62nd Annual...

  81. [89]

    Smith, Daniel Khashabi, and Hannaneh Hajishirzi

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual Mee...

  82. [90]

    Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2024. https://aclanthology.org/2024.acl-long.820 Do llamas work in E nglish? on the latent language of multilingual transformers . In Proceedings of the 62nd Annual Meeting of the Association for Computationa...

  83. [91]

    Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzm \'a n, Armand Joulin, and Edouard Grave. 2020. https://aclanthology.org/2020.lrec-1.494 CCN et: Extracting high quality monolingual datasets from web crawl data . In Proceedings of the Twel...

  84. [92]

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. https://doi.org/10.18653/v1/2021.naacl-main.41 m T 5: A massively multilingual pre-trained text-to-text transformer . In Proceedings of the 2021 Conferenc...

  85. [93]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4...

  86. [94]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with BERT . In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April...

  87. [95]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36:46595--46623

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.