Pith. sign in

A Stochastic Finite-State Word-Segmentation Algorithm for Chinese

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We present a stochastic finite-state model for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names. We also evaluate the system's performance, taking into account the fact that people often do not agree on a single segmentation.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Entropy-Driven Pre-Tokenization for Byte-Pair Encoding

cs.CL · 2025-06-18 · conditional · novelty 4.0

Entropy-guided pre-tokenization, especially a PMI plus left/right entropy combination with lambda equal to 4, raises BPE word-boundary F1 from 49.30 to 58.73 on a 2,255-sentence PKU subset.

citing papers explorer

Showing 1 of 1 citing paper.

  • Entropy-Driven Pre-Tokenization for Byte-Pair Encoding cs.CL · 2025-06-18 · conditional · none · ref 12 · internal anchor

    Entropy-guided pre-tokenization, especially a PMI plus left/right entropy combination with lambda equal to 4, raises BPE word-boundary F1 from 49.30 to 58.73 on a 2,255-sentence PKU subset.