{"id":"1d57620e-d8d4-4185-989f-11e00c8f4f0b","arxiv_id":"2411.16387","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":12,"one_line_summary":"A multi-stage filtering pipeline produces a 14 GB Traditional Chinese web text dataset, and LLM-based scoring indicates it is cleaner and more educational than earlier pipeline stages.","lead":"This paper creates FineWeb-zhtw, a cleaned collection of Traditional Chinese text from the web, using a series of filters adapted from the FineWeb pipeline. It is one of the few public data resources for training or studying Traditional Chinese language models, and the authors report that their filtering improves judged text quality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central quality claim rests on unvalidated GPT-3.5 scores; without human agreement or a downstream training check, the reported gains may reflect length/formality artifacts rather than pretraining value.","rationale":"The paper is a clear, well-structured dataset curation effort with a reproducible pipeline built on datatrove, public release intent, and internally consistent ablations. The reader's conditional verdict is appropriate. My stress-test agrees with the reader's weakest assumption that LLM-based scores are the load-bearing proxy, and I sharpen it with a concrete mechanism: the filters remove short, fragmented, or boilerplate-heavy text, which can mechanically raise scores on a rubric that rewards coherent, structured, educational content without necessarily improving pretraining utility. The primary concern is not that the pipeline is wrong, but that the evidence does not yet establish the strong claim that the score gains translate to better language model pretraining. I also note the paper's own statement that no significant gains were achieved in the Sensitive Content category from the FineWeb filtering stage, which mildly undermines the 'consistent improvements across all categories' wording even though the main quality metrics show large, significant gains. A human annotation study with length matching directly tests the artifact hypothesis and would either validate or weaken the central claim. If the human raters confirm the GPT-3.5 ordering, the conditional verdict could be upgraded; if not, the paper would need downstream training evidence to support its quality claim. Since the reader already recommended conditions, the verdict should remain unchanged: accept only if the scoring proxy is validated or the claim is softened.","tokens_in":7730,"tokens_out":6712,"duration_ms":68413,"concrete_test":"Have two native Traditional Chinese annotators score a length-matched sample of 200 documents from each of the three stages (basic-filtered, language-identified, final FineWeb-zhtw) using the exact rubric in the Appendix, and compute annotator agreement and the correlation between human and GPT-3.5 scores. If human raters do not reproduce the GPT-3.5 ordering, or if the final-stage advantage disappears after matching on document length, then the Section 3.2 claim is an artifact of the scoring proxy rather than a real quality improvement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 uses GPT-3.5 as a scoring agent to rate naturalness, educational value, and sensitive content, but no human validation or score-consistency analysis is reported. The t-tests in Section 3.2 therefore show only that 1,000 sampled documents from the final pipeline score higher than samples from earlier stages on one unvalidated LLM rubric. Because the rubric explicitly rewards logically structured, coherent, curriculum-like writing, and because the pipeline's line- and document-level filters (Sections 2.3–2.5) remove short lines, high symbol ratios, boilerplate, and repetitive text, the score gains may be driven by mechanical properties such as document length and formatting rather than by content that improves pretraining. The paper also reports no significant gain in the Sensitive Content category, which complicates the wording 'consistent improvements across all categories.' In addition, filter thresholds were selected by manual inspection of a single Common Crawl dump (CC-MAIN-2024-26), with no held-out dump validation, so scalability to other dumps is not demonstrated. Thus the central claim that FineWeb-zhtw materially improves corpus quality for LLM pretraining remains conditional on an unvalidated proxy.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents FineWeb-zhtw, a pipeline for curating Traditional Chinese text from Common Crawl. It applies a cascade of filters—basic HTML/text extraction, a custom Traditional/Simplified language identifier, Gopher, C4, and FineWeb quality filters, and minhash deduplication—to one Common Crawl dump (CC-MAIN-2024-26), yielding 14.04 GB of text. The authors evaluate the final dataset by scoring 1,000 random samples with GPT-3.5 on three criteria (Traditional Chinese naturalness, educational value, sensitive content) and compare the scores against the intermediate outputs of the pipeline. They report statistically significant improvements in naturalness and educational value, and argue the filtering is effective. Code and data are released.","tokens_in":8012,"tokens_out":5255,"duration_ms":43630,"significance":"If the evaluation were valid, the dataset would fill a genuine gap: public Traditional Chinese pretraining corpora are scarce, and the authors' pipeline adapts established English filters to the linguistic properties of Traditional Chinese. The use of datatrove and the release of code/data are concrete community contributions. However, the paper's only evidence is an unvalidated LLM rubric on 1,000 samples, with no human agreement, no downstream training, no baseline comparison, and no cross-dump validation. The announced 'consistent improvements across all categories' is contradicted by the paper's own reporting for the sensitive-content category. The resource may be useful, but the central quality claim is not yet established.","major_comments":[{"comment":"The paper states that 'FineWeb-zhtw dataset has consistent improvements across all categories,' yet in the same section it reports that for Sensitive Content 'no statistically significant gains are achieved from the FineWeb filtering stage.' This is a direct contradiction in the paper's central claim. The authors should either revise the claim to 'all categories except sensitive content' or provide a different interpretation that is consistent with the reported p-values.","section":"Section 3.2"},{"comment":"The evaluation relies entirely on GPT-3.5 scores without human validation, inter-annotator agreement, or a reliability analysis. The rubric explicitly rewards logical structure, coherence, and educationally relevant content, while the filters in Sections 2.3–2.5 explicitly remove short lines, high symbol ratios, repetitive text, and poorly punctuated lines. Thus the observed score gains may be mechanical consequences of the filtering criteria rather than evidence that the data is better for pretraining. The paper should add a human-annotated subsample (e.g., 100–200 documents scored by native speakers) to validate the LLM scores, and ideally a downstream check such as language-model perplexity on a held-out Traditional Chinese benchmark or fine-tuning on a downstream task.","section":"Section 3.1"},{"comment":"All filter thresholds are chosen by grid search with manual inspection of filtered-in/filtered-out data on a single Common Crawl dump (CC-MAIN-2024-26), as stated in Section 2. No held-out dump or sensitivity analysis is provided, so the paper's title claim of 'scalable curation' is not supported by evidence. The authors should at least apply the pipeline to a second dump and report the resulting retention rates and quality scores, or provide a threshold sensitivity analysis.","section":"Section 2"},{"comment":"The evaluation compares FineWeb-zhtw only against its own intermediate stages (basic filtering, language identification). The paper does not compare against existing public Traditional Chinese corpora such as mC4-zh, CC-100-zh, or the Chinese split of FineWeb. Without such a comparison, the authors cannot justify the claim that FineWeb-zhtw is a notable advance over the current publicly available options for Traditional Chinese pretraining.","section":"Section 3.1"}],"minor_comments":[{"comment":"The dataset is referred to as both 'FineWeb-zhtw' (title, abstract) and 'FineWeb-TC' (Sections 2 and 5). Pick one name and use it consistently.","section":"Throughout"},{"comment":"There are several typos and formatting issues: 'suﬀicient' (Section 2.3), 'eﬀiciency' (Section 2.1 or discussion), 'Labatories' (author affiliation), 'eﬀicacy' (Section 5), and the Unicode range in Section 2.2 appears corrupted.","section":"Throughout"},{"comment":"The description of line-level deduplication is vague about how 'leading or trailing' lines are identified and whether the 100-occurrence threshold is per dump or cumulative. Please specify the algorithm.","section":"Section 2.6"},{"comment":"Figure 1 mixes document counts and byte counts within a single figure; annotate each stage with its unit or separate the two types of rates into different panels.","section":"Figure 1"},{"comment":"The evaluation prompt is provided only in Traditional Chinese. An English translation would help readers who do not read Chinese to understand the rubric.","section":"Appendix"},{"comment":"The phrase 'GPT3.5 on the DaVinci 3 platform' is unclear; specify the exact model identifier (e.g., text-davinci-003) and API settings used for scoring.","section":"Section 3.1"},{"comment":"The sentence 'For the parameters mentioned in the paragraph, we determined the parameters using grid search by manually inspecting the quality of filtered-in and filtered-out data' is redundant and should be rewritten.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The main obstacle is the mismatch between the strength of the claim and the evidence. The paper would be strengthened by framing it as a resource paper and adding at least a human-evaluation sanity check and a comparison with existing corpora."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this paper gives the community something it didn't have—a 14GB, cleaned Traditional Chinese corpus from a single Common Crawl dump, with a fully specified pipeline and the evaluation prompt in the appendix. That's a real resource. But the headline claim that the filters improve quality across all categories is not actually supported by the paper's own numbers, and the evaluation rests on an unvalidated GPT-3.5 rubric.\n\nWhat's genuinely new: the Traditional/Simplified Chinese identification step and the thresholds tuned for Traditional Chinese. The rest is Gopher/C4/FineWeb filters applied through datatrove, which is fine—reuse is sensible—but the novelty is the language-specific configuration, not the filter design. The ablation is transparent: you can see each stage's removal rate, and the t-tests on 1,000 sampled documents show real gains in the two categories that matter most for the paper's argument, Naturalness and Educational Value.\n\nThe soft spots are real but not fatal. First, the sentence in Section 3.2 says 'consistent improvements across all categories,' then the very next paragraph says Sensitive Content showed no statistically significant gains. That's an internal contradiction the authors need to fix. Second, the scores come from GPT-3.5 with no human agreement check. The rubric explicitly rewards logical structure, coherence, and curriculum-like writing—exactly what the filters enforce by cutting short lines, high symbol ratios, duplication, and boilerplate. So the gains may be partly mechanical. The paper would be much stronger with a human-rated subset or a small downstream training run. Third, the thresholds were tuned on one dump, CC-MAIN-2024-26, and there's no validation on a second dump. Fourth, the manuscript says code and datasets are public but no link appears; the reader can't actually grab the data right now.\n\nThe Discussion about the 40x gap between English and Traditional Chinese and the Chinchilla shortfall is honest and useful. That's the kind of context the field needs.\n\nVerdict: this deserves a serious referee, but the referee should ask for major revision—resolve the contradiction, add human or downstream validation, provide the link, and ideally test on another dump. In its current form it's a promising resource with a preliminary evaluation, not an established result.","headline":"A useful Traditional Chinese corpus with a clear pipeline, but its quality claims rest on an unvalidated LLM rubric and one self-contradictory sentence.","tokens_in":8569,"tokens_out":2647,"would_cite":true,"duration_ms":24986,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multi-stage filtering pipeline turns raw Common Crawl web pages into a Traditional Chinese pretraining corpus, and the paper argues the resulting samples score significantly higher on LLM-rated naturalness and educational value than…","keywords":["Traditional Chinese","pretraining dataset","web text filtering","Common Crawl","LLM-based evaluation","language identification","FineWeb-zhtw","data curation"],"falsifier":"Take about 200 documents sampled from the basic-filtered, language-identified, and final stages; have human annotators apply the same 0-5 rubric and compute agreement with the GPT-3.5 scores, or train a small Traditional Chinese language model on the final dataset and on the basic-filtered baseline and compare perplexity on a held-out set. If human ratings diverge from the LLM ratings, or the model trained on the final dataset shows no downstream gain, the paper's quality-improvement claim would be undercut.","tokens_in":7550,"feed_emoji":"📚","tokens_out":6982,"duration_ms":57526,"temperature":0.7,"pith_summary":"The paper sets out to build FineWeb-zhtw, a publicly available pretraining corpus for Traditional Chinese, by adapting the FineWeb curation recipe to a language that lacks spacing and has two character variants. Its central claim is that a multi-stage pipeline of basic filtering, custom Traditional Chinese language identification, Gopher/C4/FineWeb quality filters, and minhash deduplication yields a corpus whose samples score consistently higher than earlier-stage samples on GPT-3.5-rated naturalness, educational value, and total quality. If the claim holds, researchers get a ready-to-use dataset and a transferable recipe for a comparatively underserved language, and the reported 40x data-size gap between English and Traditional Chinese would urge new collection efforts beyond Common Crawl.","feed_headline":"Multi-stage filters lift Traditional Chinese web-corpus quality","feed_subtitle":"Filters cut Common Crawl to 0.5% while boosting naturalness and educational-value scores.","key_machinery":"The central object is the filter cascade: a fuzzy unicode-range pre-filter, a fastText-based language identifier augmented by phrase-level Traditional/Simplified Chinese discrimination, Gopher quality heuristics (document length, symbol ratio, ellipsis ratio, stop words), C4 line-level filters (JavaScript, policy boilerplate, bracket ratio), FineWeb document-level filters (line punctuation, short line, character duplication, new line ratios), and minhash deduplication. The evaluation machinery is GPT-3.5 used as a scoring agent with a specified 0-5 rubric for naturalness, educational value, and sensitive content, with t-tests comparing pipeline stages.","core_discovery":"The paper reports that after applying the full pipeline to Common Crawl dump CC-MAIN-2024-26, the final dataset retains about 0.5% of documents (214.04 GB of text) and that 1,000 randomly sampled documents from the final dataset score significantly higher on 0-5 GPT-3.5 ratings than samples from basic-filtered and language-identified stages: naturalness rises from 1.72 to 2.42, educational value from 1.54 to 2.04, and total score from 7.53 to 9.17, with t-test p-values below 0.05. Sensitive-content scores stay high and show no significant gain, indicating that Common Crawl under this pipeline is not a major source of harmful content. The paper also quantifies a roughly 40x gap in document volume between English and Traditional Chinese after language identification and argues on Chinchilla scaling grounds that Common Crawl alone is insufficient for a 70B-parameter Traditional Chinese model.","pith_inferences":["A direct downstream check should be run: pretrain a small Traditional Chinese model on FineWeb-zhtw versus a basic-filtered control and compare on a Traditional Chinese benchmark; the paper does not include this test.","Because the LLM scoring agent is GPT-3.5, the method's validity for other low-resource languages remains contingent on the scoring model's competence in those languages.","The grid-searched thresholds were tuned by manual inspection on one Common Crawl dump; re-validating them on a later dump would tell whether the recipe is stable over time.","The reported 40x English-to-Traditional-Chinese gap implies that simply re-running web-scale pipelines will not close the scale deficit; complementary curated sources or synthetic data would be needed."],"forward_implications":["If the reported quality gains are real, FineWeb-zhtw (214.04 GB for Common Crawl dump CC-MAIN-2024-26) is a ready-to-use pretraining corpus for Traditional Chinese.","The statistically significant t-test improvements across naturalness, educational value, and total score indicate that the full cascade removes more low-quality content than basic filtering or language identification alone.","The roughly 40x gap in document count between English and Traditional Chinese after language identification, combined with the Chinchilla scaling estimate, supports the paper's conclusion that Common Crawl alone is insufficient for a 70B-parameter Traditional Chinese model.","The public release of code and dataset makes the pipeline applicable to future Common Crawl dumps and adaptable to other Chinese language variants."],"supporting_citations":[{"why":"Supplies the FineWeb quality filters and overall curation paradigm that this work adapts to Traditional Chinese.","marker":"Penedo et al., 2024a"},{"why":"Provides the Gopher heuristic quality filter criteria (document length, symbol ratio, ellipsis ratio, stop words).","marker":"Rae et al., 2022"},{"why":"Provides the C4 line-level filtering criteria (JavaScript, policy boilerplate, bracket ratio).","marker":"Raffel et al., 2023"},{"why":"Establishes the language-model-as-scoring-agent methodology used for evaluating dataset samples.","marker":"Chiang and Lee, 2023"},{"why":"Supplies fastText, the language identifier that motivates the custom Traditional/Simplified Chinese filter.","marker":"Bojanowski et al., 2017"},{"why":"Provides datatrove, the implementation tool used to run the curation pipeline at scale.","marker":"Penedo et al., 2024b"},{"why":"Gives the Chinchilla scaling law used to estimate that Common Crawl alone is insufficient for a 70B Traditional Chinese model.","marker":"Hoffmann et al., 2022"}],"fun_headline_variants":["0.5% of Common Crawl becomes quality Traditional Chinese text","Multi-stage filters lift Traditional Chinese corpus ratings","FineWeb-zhtw: curating Traditional Chinese from 0.5% web","Filtering boosts Traditional Chinese naturalness, education scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claims rest on using GPT-3.5 ratings of naturalness and educational value as proxies for pretraining-data quality, with no human validation of those ratings and no downstream training run to confirm the proxy.","fun_headline_variants_meta":{"raw":{"variants":["0.5% of Common Crawl becomes quality Traditional Chinese text","Multi-stage filters lift Traditional Chinese corpus ratings","FineWeb-zhtw: curating Traditional Chinese from 0.5% web","Filtering boosts Traditional Chinese naturalness, education scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000525,"raw_usage":{"total_tokens":2486,"prompt_tokens":848,"completion_tokens":1638,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":1567}},"tokens_in":464,"tokens_out":1638,"duration_ms":14585,"temperature":1.0,"reasoning_tokens":1567,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:09:55.274151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take about 200 documents sampled from the basic-filtered, language-identified, and final stages; have human annotators apply the same 0-5 rubric and compute agreement with the GPT-3.5 scores, or train a small Traditional Chinese language model on the final dataset and on the basic-filtered baseline and compare perplexity on a held-out set. If human ratings diverge from the LLM ratings, or the model trained on the final dataset shows no downstream gain, the paper's quality-improvement claim would be undercut.","supporting_citations":[],"review_version":1}