Entropy-guided pre-tokenization, especially a PMI plus left/right entropy combination with lambda equal to 4, raises BPE word-boundary F1 from 49.30 to 58.73 on a 2,255-sentence PKU subset.
A Stochastic Finite-State Word-Segmentation Algorithm for Chinese
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
We present a stochastic finite-state model for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names. We also evaluate the system's performance, taking into account the fact that people often do not agree on a single segmentation.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
Entropy-guided pre-tokenization, especially a PMI plus left/right entropy combination with lambda equal to 4, raises BPE word-boundary F1 from 49.30 to 58.73 on a 2,255-sentence PKU subset.