Pith. sign in

REVIEW 2 cited by

ChineseWebText: Large-scale High-quality Chinese Web Text Extracted with Effective Evaluation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.01149 v2 pith:SDGL4WXX submitted 2023-11-02 cs.CL

classification cs.CL
keywords dataqualitychinesetextcleanhigh-qualitylarge-scalellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

During the development of large language models (LLMs), the scale and quality of the pre-training data play a crucial role in shaping LLMs' capabilities. To accelerate the research of LLMs, several large-scale datasets, such as C4 [1], Pile [2], RefinedWeb [3] and WanJuan [4], have been released to the public. However, most of the released corpus focus mainly on English, and there is still lack of complete tool-chain for extracting clean texts from web data. Furthermore, fine-grained information of the corpus, e.g. the quality of each text, is missing. To address these challenges, we propose in this paper a new complete tool-chain EvalWeb to extract Chinese clean texts from noisy web data. First, similar to previous work, manually crafted rules are employed to discard explicit noisy texts from the raw crawled web contents. Second, a well-designed evaluation model is leveraged to assess the remaining relatively clean data, and each text is assigned a specific quality score. Finally, we can easily utilize an appropriate threshold to select the high-quality pre-training data for Chinese. Using our proposed approach, we release the largest and latest large-scale high-quality Chinese web text ChineseWebText, which consists of 1.42 TB and each text is associated with a quality score, facilitating the LLM researchers to choose the data according to the desired quality thresholds. We also release a much cleaner subset of 600 GB Chinese data with the quality exceeding 90%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

    cs.CL 2025-10 conditional novelty 5.0 of 10

    Chinese LLMs generate more positive continuations after 'we' prompts and more negative after 'they' prompts; the feminine 'they' intensifies negativity in several pretrained models.

  2. FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation

    cs.CL 2025-05 reject novelty 3.0 of 10

    FuxiMT combines a frozen BLOOMz model with sparse mixture-of-experts layers, Chinese-first pretraining, and curriculum learning to translate into Chinese from 65 languages, with claimed low-resource gains that the pap...

Pith tools