A pretrained hypernetwork enables dynamic, batch-specific tokenization that compresses token sequences by 20% or more in multilingual models with under 2% average accuracy loss.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Retrofitting Large Language Models with Dynamic Tokenization
A pretrained hypernetwork enables dynamic, batch-specific tokenization that compresses token sequences by 20% or more in multilingual models with under 2% average accuracy loss.