Pith. sign in

Paper Citation Record · LEDGER

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.07463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07463 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:41:14.636513Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1bb2cb67-650b-4bc9-bbeb-2e6be8b0b00d · outbound

This paper cites Rethinking reflection in pre-training, 2025.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Rethinking reflection in pre-training, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.061372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.508198Z digest=sha256:250570efbf911169d5ad373fc82b9d9658d5b06117092ca57cfd863be6a93c51

Observation ea004250-d416-4924-ade8-c31ed815c17a · outbound

This paper cites COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.511955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.511955Z digest=sha256:b313f17cb4631ee02e147255d7cd895c31768fdb48c62c99793d5d58831d3a3a

Observation 761eba16-ca57-421d-bf7b-81fca97c88fa · outbound

This paper cites Careful selection of knowledge to solve open book question answering.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Careful selection of knowledge to solve open book question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.052572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.515368Z digest=sha256:79a094cb48151e243585ca9534a0dbe374329472042a848f58586d2d8795f25d

Observation 5ffeffb1-d8a8-400a-bdcc-4bd59c1a7ac4 · outbound

This paper cites CCI-Data [Data set].

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models CCI-Data [Data set]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.042870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.518445Z digest=sha256:532b97f85ff8e4ec695e151380f48844345ec4d4cc18ee13a1e9a728fb7da51a

Observation 0db9bfae-f766-41b9-963b-b303ddd72d8b · outbound

This paper cites CCI2-Data [Data set].

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models CCI2-Data [Data set]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.032842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.521642Z digest=sha256:a4420bab2f70f9ef1640e21e646b38f784cddbcb69ee32e37887d860a3a43e6a

Observation 821bc742-0f46-4d42-ab6c-093a5f3eb79b · outbound

This paper cites WuDaoCorporaText [Data set].

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models WuDaoCorporaText [Data set]

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.022976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.524471Z digest=sha256:b67f9199cc620e9c92f8efd1d3577d8af311460fd238916aa4d6a6739f1ec103

Observation a77e3deb-fb75-4756-be2c-09d6d7bc5a95 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language, 2019.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Piqa: Reasoning about physical commonsense in natural language, 2019

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.527577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.527577Z digest=sha256:0f2fe4a2101956cdee5f009674a2b2133b9d27f0917b22501e991d7da21a2eda

Observation c76ddd3f-3f49-4c0c-bd78-aa2c8867756a · outbound

This paper cites an unresolved cited work.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:41:15.007881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.530537Z digest=sha256:503a2cc01efdf6bde4825a5292813f448fdd76a6612bcd6f9bc66ac48930a0ce

Observation eb896623-f4b0-4fca-92b1-f8139b77609c · outbound

This paper cites Data- juicer: A one-stop data processing system for large language models.Companion of the 2024 International Conference on Management of Data, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Data- juicer: A one-stop data processing system for large language models.Companion of the 2024 International Conference on Management of Data, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.999494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.533269Z digest=sha256:9a8f828077f5b23e09e6dd867da2e06545406c6305cf0ad61a0f278821f17d1f

Observation 6016aaf2-4ce7-44a1-a002-f5a9d56d1b65 · outbound

This paper cites Chinesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Chinesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.990030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.536439Z digest=sha256:26da028469af36e86f72dc120a63cd8f7f24d18fb58de549a9f081e10a836e18

Observation 1b0e76a6-13b4-4204-baec-7bb14c2f216c · outbound

This paper cites Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.539286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.539286Z digest=sha256:e7fcedfe21522309cd847795027dc7571ff70213145b2b198e1cb544d49cbd52

Observation 2e138668-d626-46a3-9a64-3ba295c84ae2 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unsupervised Cross-lingual Representation Learning at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.542023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.542023Z digest=sha256:e29bc10a9971a747ad028050330fc4b6ecaa37f6bd51bced3dfd848597708451

Observation 0d71b92f-79bb-419f-ae51-fcffb048db63 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.975953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.545952Z digest=sha256:080a2c7401952d5e94f34857365916cd21e61d953cc28ab602a61c913ac1f816

Observation 61725105-411f-4f92-a3ee-38d037407545 · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Lighteval: A lightweight framework for llm evaluation, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.965967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.549699Z digest=sha256:b264b1ef3488a0c039e89cbfe8254519ed86dc27cfae91eae4591c0bf1b4b4f8

Observation 491272bb-96a7-46ad-990d-94452d443c50 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.552614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.552614Z digest=sha256:6320cb0274ff5270669706636172d22b66823e65b17917416fd6a6801c6ff295

Observation eedb7d1d-8808-4cc4-afce-714e6be0fb58 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.555798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.555798Z digest=sha256:2a11a659c4eaa36567932b5052a949e220567b4d6b527ceeb02a2fa5338698c8

Observation 568e34bd-b5b9-4742-a8e1-d580cb3dea04 · outbound

This paper cites Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.956085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.559100Z digest=sha256:5dcc1d7149808993d41ddc0c9d6af6135932b394cbd0646bba5f8dc5eca3abb6

Observation 124c3a16-07df-4d3c-9311-ab3121e7cefe · outbound

This paper cites Measuring massive multitask language understanding, 2021.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Measuring massive multitask language understanding, 2021

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.562041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.562041Z digest=sha256:40031c276c32b97b031564fb10c90ad9aa16752695d3aa3cb4745bc5b0e339bc

Observation 21ec9a21-6c4c-414c-a150-2e4b90653a8e · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.565274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.565274Z digest=sha256:98960672e106c22e3fdd2e17a79bb97f83778df1c935074a32f4081f209c39eb

Observation 8a95f47b-695d-4749-9c5a-66e148574b86 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Bag of Tricks for Efficient Text Classification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.568240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.568240Z digest=sha256:115aa7348e674df67cf6df2aa7ff939dda4d2a2f8756d530587a7ac6d4a5845e

Observation 3d10ba18-7116-4154-add2-7ca26760cf8e · outbound

This paper cites Deduplicating training data makes language models better, 2022.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Deduplicating training data makes language models better, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.571434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.571434Z digest=sha256:b7ddf2cc3b506557f2cc8aff03750e95d2971b6d86a154f03990729582a9a707

Observation adeb30bd-1aff-4068-a773-356b85376e54 · outbound

This paper cites Levesque, Ernest Davis, and Leora Morgenstern.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Levesque, Ernest Davis, and Leora Morgenstern

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.930480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.574169Z digest=sha256:6fb8e1c9731fb86cbdfdf234e70c1d1835601e1afff508c5987bf68d8f343230

Observation f4787769-6224-46fb-bbf7-6ecd8ae82858 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cmmlu: Measuring massive multitask language understanding in chinese, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.921385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.577122Z digest=sha256:5212c3d59cab13a42984b5938893542ca485442d7034c1149ecfa468002cabda

Observation eecd1c91-24f1-4d04-af7c-e816aa38a308 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Datacomp-lm: In search of the next generation of training sets for language models, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.912583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.579822Z digest=sha256:c8495eed10ff2b2836c0450559d60d3612104c2a5c6ff3994e96e5ccd21bfd9e

Observation 5a40faad-a015-4d50-b153-f5e2fc448e83 · outbound

This paper cites Openhermes 2.5-zh: A partial chinese translation of openhermes-2.5, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Openhermes 2.5-zh: A partial chinese translation of openhermes-2.5, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.903381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.582561Z digest=sha256:a639604d0eaabf22d1e0eb48e50e46a5f04b405a28b62787cfd36c2f343b2a65

Observation f80f6c1e-32d9-412a-82e4-9f296db90657 · outbound

This paper cites Fineweb2: A sparkling update with 1000s of languages, December 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Fineweb2: A sparkling update with 1000s of languages, December 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.894325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.585479Z digest=sha256:b0952303024b6ce16a688715ce3a625f83c62a2b67db90b4e2152f973dbe1a52

Observation 170dd499-10b3-44f6-a796-c6a1bb4bd353 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.588446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.588446Z digest=sha256:5d7beb205f4831bd17c80c1f3c99462534e34b6c19856bfd856063e87cbebac4

Observation ec09a55b-c1db-4934-9655-83e2df2b2ec1 · outbound

This paper cites Deduplicate Text Datasets.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Deduplicate Text Datasets

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.884898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.592064Z digest=sha256:0ad6743b591a7ce512338e0a0b1d2f12b04e19ec82a4379b1d66c5aef53aee28

Observation 0bee756a-db61-495c-b69a-e91352b8ca2b · outbound

This paper cites Socialiqa: Com- monsense reasoning about social interactions, 2019.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Socialiqa: Com- monsense reasoning about social interactions, 2019

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.596279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.596279Z digest=sha256:9007cee1cfad03130a02f890dec7377a92ee7a0a13dd321d9d0141f85ec37400

Observation edbe1d46-58dd-4497-939c-f4d2edd9d3f6 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.600683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.600683Z digest=sha256:2989b7a5ee1a5d3ac798ee84492c597bcd72efced2e376eb09c1b7fc80a23c5b

Observation ae53137f-0530-4d27-a82b-da4d9d9ba88d · outbound

This paper cites Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.788616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.604175Z digest=sha256:3784d0553283f838ac7fab9da9739bb02917b79019e702f893664c35bf89dcc3

Observation d4c02cd1-5746-45b0-9ac3-8b0cc7ddabc6 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.610652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.610652Z digest=sha256:110e933a5ebccc964905f5baafcf27286570d4b2c58604d4b7e55ecd3bf3b5b5

Observation 5ff8f4b7-0c07-4cd2-83bb-0dc2b5b4905b · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.613803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.613803Z digest=sha256:61d776d7666cb087910a022d7264521833168340babf1ee2096b4d6404a57573

Observation 62abf318-fd8f-4daf-bd0b-cb4e3abf1b12 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Qwen2.5: A party of foundation models, September 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.616833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.616833Z digest=sha256:30017cc215be5c0a2c3c222d07a3fa3ac6e0a9323e1f9f770488d5222f1aa18f

Observation bfb42c5d-8ef1-4ffb-a744-bcb6f04f7312 · outbound

This paper cites Cci3.0-hq: a large-scale chinese dataset of high quality designed for pre-training large language models, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cci3.0-hq: a large-scale chinese dataset of high quality designed for pre-training large language models, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.773012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.619787Z digest=sha256:c960b764998fb76d3553af9aca1eccd405311df0d97e45f277ccfaaad570fb99

Observation ee8c4ba0-2457-4119-97b4-aa53e8736227 · outbound

This paper cites Qwen2 Technical Report.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Qwen2 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.622934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.622934Z digest=sha256:a43c819641a3c287092deb209f615a1baef78d1bf8e689cfbd26eedba33d7078

Observation fec69a07-0e43-46e1-91ea-2e38f3306176 · outbound

This paper cites Opencsg chinese corpus: A series of high-quality chinese datasets for llm training, 2025.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Opencsg chinese corpus: A series of high-quality chinese datasets for llm training, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.763329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.626461Z digest=sha256:567e46886e17f277f37fd575e6a0ba299c9804e8c2de3267e4568f99f9f21e20

Observation 094a470d-2194-4559-bdd4-ab9441d562f3 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.629752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.629752Z digest=sha256:a9d2e6c56d3170f52a08ccdcc69bc9cd757084cabd5024f323ec8382c690d652

Observation 730a1797-7f91-41a9-a855-48e534836ef8 · outbound

This paper cites Games" and.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Games" and

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.753261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.633110Z digest=sha256:8f22cf29ebeeaef98cb502e83028a30e532cce594f1dd4e5a949b4c1f2f0a88a

Observation 02d8a56d-45ad-4bdd-ac3c-98ac11d90f2d · outbound

This paper cites an unresolved cited work.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:41:14.743827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:41:14.636513Z digest=sha256:7a3f4b1fd1ce208256a62ed977636e66ab5ecf45c6ef51ec78886ee912df8d4a

Pith citing papers

No inbound Pith citation observations are available.