Continued pretraining with LoRA and a Dutch-specific tokenizer improves Llama-2's Dutch, but gives limited gains for already-multilingual Llama-3.
Adapting BigScience Multilingual Model to Unseen Languages
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
We benchmark different strategies of adding new languages (German and Korean) into the BigScience's pretrained multilingual language model with 1.3 billion parameters that currently supports 13 languages. We investigate the factors that affect the language adaptability of the model and the trade-offs between computational costs and expected performance.
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
ChocoLlama: Lessons Learned From Teaching Llamas Dutch
Continued pretraining with LoRA and a Dutch-specific tokenizer improves Llama-2's Dutch, but gives limited gains for already-multilingual Llama-3.