A new 35 TB bilingual pretraining dataset with 4.5 billion chain-of-thought templates is described, but the evidence for its benefits is marginal, confounded, and contradicted by the paper's own tables.
CCI2-Data [Data set]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models
A new 35 TB bilingual pretraining dataset with 4.5 billion chain-of-thought templates is described, but the evidence for its benefits is marginal, confounded, and contradicted by the paper's own tables.