A fixed-vocabulary tokenizer surgery method reduces Ukrainian token counts by 33-37% on Nemotron and GPT-OSS while preserving 77-78% of original token IDs and leaving English and EU token counts essentially unchanged.
Title resolution pending
1 Pith paper cite this work, alongside 3 external citations. Polarity classification is still indexing.
1
Pith paper citing it
3
external citations · OpenAlex
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2026 1verdicts
ACCEPT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Writing-System-Level Tokenizer Adaptation for Byte-Level BPE
A fixed-vocabulary tokenizer surgery method reduces Ukrainian token counts by 33-37% on Nemotron and GPT-OSS while preserving 77-78% of original token IDs and leaving English and EU token counts essentially unchanged.