Using a shortest-path search to split text with the same token vocabulary saves 3-5% of tokens on many languages, but downstream accuracy gains are mixed and confounded by the experimental design.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
When Every Token Counts: Optimal Segmentation for Low-Resource Language Models
Using a shortest-path search to split text with the same token vocabulary saves 3-5% of tokens on many languages, but downstream accuracy gains are mixed and confounded by the experimental design.