A fixed-vocabulary tokenizer surgery method reduces Ukrainian token counts by 33-37% on Nemotron and GPT-OSS while preserving 77-78% of original token IDs and leaving English and EU token counts essentially unchanged.
Data-efficient adaptation of multilingual LLMs to Ukrainian
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Writing-System-Level Tokenizer Adaptation for Byte-Level BPE
A fixed-vocabulary tokenizer surgery method reduces Ukrainian token counts by 33-37% on Nemotron and GPT-OSS while preserving 77-78% of original token IDs and leaving English and EU token counts essentially unchanged.