A tokenizer trained on more languages than the model's main pretraining set makes later language adaptation faster and better, with minimal loss on the pretraining languages.
Scalable Training of Language Models using JAX pjit and TPUv4
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
Modern large language models require distributed training strategies due to their size. The challenges of efficiently and robustly training them are met with rapid developments on both software and hardware frontiers. In this technical report, we explore challenges and design decisions associated with developing a scalable training framework, and present a quantitative analysis of efficiency improvements coming from adopting new software and hardware solutions.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
A tokenizer trained on more languages than the model's main pretraining set makes later language adaptation faster and better, with minimal loss on the pretraining languages.