REVIEW 1 cited by
A Stochastic Finite-State Word-Segmentation Algorithm for Chinese
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a stochastic finite-state model for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names. We also evaluate the system's performance, taking into account the fact that people often do not agree on a single segmentation.
Forward citations
Cited by 1 Pith paper
-
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
Entropy-guided pre-tokenization, especially a PMI plus left/right entropy combination with lambda equal to 4, raises BPE word-boundary F1 from 49.30 to 58.73 on a 2,255-sentence PKU subset.
Discussion (0). Continue with ORCID to comment.