Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:41:14.636513Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.07463.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:41:14.636513Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1bb2cb67-650b-4bc9-bbeb-2e6be8b0b00d · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Rethinking reflection in pre-training, 2025
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea004250-d416-4924-ade8-c31ed815c17a · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 761eba16-ca57-421d-bf7b-81fca97c88fa · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Careful selection of knowledge to solve open book question answering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ffeffb1-d8a8-400a-bdcc-4bd59c1a7ac4 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models CCI-Data [Data set]
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0db9bfae-f766-41b9-963b-b303ddd72d8b · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models CCI2-Data [Data set]
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 821bc742-0f46-4d42-ab6c-093a5f3eb79b · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models WuDaoCorporaText [Data set]
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a77e3deb-fb75-4756-be2c-09d6d7bc5a95 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Piqa: Reasoning about physical commonsense in natural language, 2019
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76ddd3f-3f49-4c0c-bd78-aa2c8867756a · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb896623-f4b0-4fca-92b1-f8139b77609c · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Data- juicer: A one-stop data processing system for large language models.Companion of the 2024 International Conference on Management of Data, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6016aaf2-4ce7-44a1-a002-f5a9d56d1b65 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Chinesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b0e76a6-13b4-4204-baec-7bb14c2f216c · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e138668-d626-46a3-9a64-3ba295c84ae2 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unsupervised Cross-lingual Representation Learning at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d71b92f-79bb-419f-ae51-fcffb048db63 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61725105-411f-4f92-a3ee-38d037407545 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Lighteval: A lightweight framework for llm evaluation, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 491272bb-96a7-46ad-990d-94452d443c50 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eedb7d1d-8808-4cc4-afce-714e6be0fb58 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568e34bd-b5b9-4742-a8e1-d580cb3dea04 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 124c3a16-07df-4d3c-9311-ab3121e7cefe · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Measuring massive multitask language understanding, 2021
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ec9a21-6c4c-414c-a150-2e4b90653a8e · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a95f47b-695d-4749-9c5a-66e148574b86 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Bag of Tricks for Efficient Text Classification
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d10ba18-7116-4154-add2-7ca26760cf8e · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Deduplicating training data makes language models better, 2022
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adeb30bd-1aff-4068-a773-356b85376e54 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Levesque, Ernest Davis, and Leora Morgenstern
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4787769-6224-46fb-bbf7-6ecd8ae82858 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cmmlu: Measuring massive multitask language understanding in chinese, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eecd1c91-24f1-4d04-af7c-e816aa38a308 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Datacomp-lm: In search of the next generation of training sets for language models, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a40faad-a015-4d50-b153-f5e2fc448e83 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Openhermes 2.5-zh: A partial chinese translation of openhermes-2.5, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f80f6c1e-32d9-412a-82e4-9f296db90657 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Fineweb2: A sparkling update with 1000s of languages, December 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 170dd499-10b3-44f6-a796-c6a1bb4bd353 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec09a55b-c1db-4934-9655-83e2df2b2ec1 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Deduplicate Text Datasets
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0bee756a-db61-495c-b69a-e91352b8ca2b · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Socialiqa: Com- monsense reasoning about social interactions, 2019
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edbe1d46-58dd-4497-939c-f4d2edd9d3f6 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae53137f-0530-4d27-a82b-da4d9d9ba88d · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4c02cd1-5746-45b0-9ac3-8b0cc7ddabc6 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff8f4b7-0c07-4cd2-83bb-0dc2b5b4905b · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62abf318-fd8f-4daf-bd0b-cb4e3abf1b12 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Qwen2.5: A party of foundation models, September 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb42c5d-8ef1-4ffb-a744-bcb6f04f7312 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cci3.0-hq: a large-scale chinese dataset of high quality designed for pre-training large language models, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee8c4ba0-2457-4119-97b4-aa53e8736227 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Qwen2 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec69a07-0e43-46e1-91ea-2e38f3306176 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Opencsg chinese corpus: A series of high-quality chinese datasets for llm training, 2025
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 094a470d-2194-4559-bdd4-ab9441d562f3 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730a1797-7f91-41a9-a855-48e534836ef8 · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Games" and
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02d8a56d-45ad-4bdd-ac3c-98ac11d90f2d · outbound
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.