SnakModel, an open Danish 7B LLM, outperforms other Llama2-7B-based models on the ScandEval Danish benchmark, with analyses of training dynamics and data curation.
Training a T5 Using Lab-sized Resources
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Training large neural language models on large datasets is resource- and time-intensive. These requirements create a barrier to entry, where those with fewer resources cannot build competitive models. This paper presents various techniques for making it possible to (a) train a large language model using resources that a modest research lab might have, and (b) train it in a reasonable amount of time. We provide concrete recommendations for practitioners, which we illustrate with a case study: a T5 model for Danish, the first for this language.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SnakModel: Lessons Learned from Training an Open Danish Large Language Model
SnakModel, an open Danish 7B LLM, outperforms other Llama2-7B-based models on the ScandEval Danish benchmark, with analyses of training dynamics and data curation.