Pith. sign in

REVIEW 1 cited by

PKUSEG: A Toolkit for Multi-Domain Chinese Word Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.11455 v3 pith:G7UJG4NZ submitted 2019-06-27 cs.CL

classification cs.CL
keywords segmentationworddatapkusegtoolkitchinesedomaindomains
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Chinese word segmentation (CWS) is a fundamental step of Chinese natural language processing. In this paper, we build a new toolkit, named PKUSEG, for multi-domain word segmentation. Unlike existing single-model toolkits, PKUSEG targets multi-domain word segmentation and provides separate models for different domains, such as web, medicine, and tourism. Besides, due to the lack of labeled data in many domains, we propose a domain adaptation paradigm to introduce cross-domain semantic knowledge via a translation system. Through this method, we generate synthetic data using a large amount of unlabeled data in the target domain and then obtain a word segmentation model for the target domain. We also further refine the performance of the default model with the help of synthetic data. Experiments show that PKUSEG achieves high performance on multiple domains. The new toolkit also supports POS tagging and model training to adapt to various application scenarios. The toolkit is now freely and publicly available for the usage of research and industry.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A perceptual listening audit of 2,280 utterance pairs sets a cosine-similarity threshold of 0.354 for removing likely different-speaker utterances from Common Voice client IDs.

Pith tools