Pith. sign in

Paper Citation Record · LEDGER

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

As of 20 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2501.08197.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08197 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:33:20.690921Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:10:21.973692Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation d39cba6f-04f3-4814-a38f-b0bff6618873 · outbound

This paper cites write newline.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.501166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.501166Z digest=sha256:984c866c315ec4378a25eeef84e15d2446d03781aa87d2dc8dce3a0849615b57

Observation 4a42b0de-cbb5-434f-8140-355579ac362e · outbound

This paper cites @esa (Ref.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training @esa (Ref

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.506441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.506441Z digest=sha256:263a2e95c5191776a6985a02e09b9ff16f9ac80e0331354e70e865e3fab92731

Observation 6d9e9e2c-cec9-4447-a28b-61d80178b49f · outbound

This paper cites an unresolved cited work.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.510733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.510733Z digest=sha256:749853a85b28b72b9f16aeef1ec06904246bd6841765f837c8771e00ed5e3b19

Observation 1b9b8a48-5942-4341-a930-ffc257152ca8 · outbound

This paper cites an unresolved cited work.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:33:21.275063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.514821Z digest=sha256:56a58b61220f1a1daf9ec3a2dcdec1a212e58055dc493b65fe9559df8bd95795

Observation b0f9563d-2d1e-485b-9053-19e2cf3e0151 · outbound

This paper cites 5R _U^* mE,zq R((>xP D*.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training 5R _U^* mE,zq R((>xP D*

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.262245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.518466Z digest=sha256:473652c2c4167cd0346b34dd90055c397267fc19a49cac4b2d00d5561c269008

Observation 5d44c25e-8b08-4fa3-bb17-38dc9cbe3893 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Yi: Open Foundation Models by 01.AI

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.525624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.525624Z digest=sha256:c547d40989b67febfa0e473a09ff38226a0d01f351fa6259028122cebc18d739

Observation e91b199f-da3a-496f-9dbb-84e6f7ae1cc8 · outbound

This paper cites Smollm2 - with great data, comes great performance, 2024.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Smollm2 - with great data, comes great performance, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.249465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.530484Z digest=sha256:7a29722a9466c2b6e29bf687f72a61505b5e079d4d6cd68715f6b1b883617cce

Observation 878c8c13-fdc5-45e0-9965-d04a5e495711 · outbound

This paper cites Wudao corpus, 2023.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Wudao corpus, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.237489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.534725Z digest=sha256:93934f47aec6d744d2610fdc55e520e41f431cf0193e70cf779f49528a56f437

Observation 5d440660-c506-41c3-a414-3bd91df3fa32 · outbound

This paper cites LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.538968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.538968Z digest=sha256:c07b21bf40157e0068d67a6b2608416aa0fcc989decefdd997a86b38162f0d41

Observation a72c3cea-5801-4141-8840-c82b49e471be · outbound

This paper cites Cosmopedia, 2024.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Cosmopedia, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.543372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.543372Z digest=sha256:d3eb1371c4ab5aec92771efaea06787b961c9f758e880c38e1b5feb12b879fc0

Observation fd0dc9a8-8bfe-4812-9b4e-825a51c99796 · outbound

This paper cites an unresolved cited work.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:33:21.218309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.547203Z digest=sha256:7df3e693f475cb520c3ffce2a05d6be981d14175d8e48c282b0d9d5c5c43ef93

Observation 3114c927-10ed-44f3-91af-b41410b546a4 · outbound

This paper cites Chinesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model, 2023.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Chinesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.551107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.551107Z digest=sha256:e59c47e1e8b5c8c1cc07a53976e5515f745c109d2dbe9f64b26156d7637f4bda

Observation 2e5a9e7b-d082-451b-99a1-e180089c9f4f · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, 2023.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Redpajama: An open source recipe to reproduce llama training dataset, 2023

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.554853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.554853Z digest=sha256:32504f47d28b88d6560d59ce487feda482d6836775dcede468ee654dc2f7866d

Observation 7d8f7dc2-706e-425d-b3af-3c2f4f3db144 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.559243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.559243Z digest=sha256:c03ab449f5ccb8cfe49ff227249e471f6740914062953de3906f311bbbf4b5c4

Observation 28c2855a-1883-4524-b7cf-bcc3ff03dfc8 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.563771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.563771Z digest=sha256:3f5b68b54d0475cf89fde5f62d810a40e93c8327c0ea585727436c4ff75055cb

Observation ab185cee-e85d-4c20-95c8-b8a4405b4bce · outbound

This paper cites Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.568106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.568106Z digest=sha256:1769c662da012da94fd1b18fe78bf9c2e9025376743ceddad535b3c209384175

Observation a6ebadfb-c96e-4b08-8c86-57d014c902bb · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.572174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.572174Z digest=sha256:cb4eb6a41af750b3ef575f605bfa00a1440c671054299b4e02d9595448ff9ee7

Observation da212012-e2ab-429c-9b1b-2257f381b958 · outbound

This paper cites FlagAlpha /llama2-chinese, 2023.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training FlagAlpha /llama2-chinese, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.193842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.576249Z digest=sha256:10aac4a05b25f96962ea4b9b8c6ede0e355515639096957bb4c51ac2ac898e5d

Observation b728e479-3ebe-4075-8647-75e9ddf3d69f · outbound

This paper cites Textbooks Are All You Need.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Textbooks Are All You Need

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.579818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.579818Z digest=sha256:ca02508334d785a1ad3fa71c8a4c71c5bd7a675740dd90f8ebefa0bb89c0d6d5

Observation ee498378-86e0-4dbc-a6cd-8bd37125f4b0 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.583743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.583743Z digest=sha256:135f5b76639c650520f438e93c5824187c4271fa384af6a1e0c66409037e7e61

Observation 7aea41d5-7fab-4d90-8275-ec15ccb113c0 · outbound

This paper cites Mistral 7B.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.587234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.587234Z digest=sha256:e3f3901d0ffbe90072f22d1634929a5ec1ce5333463f23b1c82b2b2e8542a2da

Observation f6f91f34-d776-4ad7-bbb4-b742b92e8fa0 · outbound

This paper cites yangjianxin1/firefly, 2024.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training yangjianxin1/firefly, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.182150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.590639Z digest=sha256:219ac2af3073f121377d7ff936e7c8449c1bbd628b3e1692846805c9e1378b49

Observation 24bb2465-eb2e-4d68-b9b7-4d06977820b5 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese, 2023.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Cmmlu: Measuring massive multitask language understanding in chinese, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.593954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.593954Z digest=sha256:ffedeabdfb9c7cdbc6352164ba5d0434beb1810028f80490a6e9ec246a795dd2

Observation 6066c6a6-9ed7-4ece-93e1-1a47bde9d348 · outbound

This paper cites Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.597140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.597140Z digest=sha256:8bfe50fecc941b2b4382cbe6b5961eafd077662fac83d5636a0fe5c435d431d8

Observation f2eb2ecc-500a-4eff-985b-4bc52e64441f · outbound

This paper cites Alignbench: Benchmarking chinese alignment of large language models, 2023 a.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Alignbench: Benchmarking chinese alignment of large language models, 2023 a

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.156370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.600579Z digest=sha256:1d028b3d202cf588f05e31720c68faa7553e27b03e01db2ee8d4a3b6071008d8

Observation c3485399-cb3c-4c62-8265-829878092397 · outbound

This paper cites MiChao-HuaFen 1.0: A Specialized Pre-trained Corpus Dataset for Domain-specific Large Models.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training MiChao-HuaFen 1.0: A Specialized Pre-trained Corpus Dataset for Domain-specific Large Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:33:20.904151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.604501Z digest=sha256:8fb9cbc451265cff85b329dfb8f99a4bf9c0944777856deaee7090daa440f89f

Observation 1b2b70f9-030a-4a25-a09b-9944904b7aba · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.609878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.609878Z digest=sha256:f85b5ce8b9f0789790caf36f00c8b7aee6438e62731c51e0aa755b50c7cf64f9

Observation 72f56dcd-6efa-48dd-8c27-9f8e32da9b94 · outbound

This paper cites Infinity instruct, 2024.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Infinity instruct, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.143941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.615349Z digest=sha256:c6f8a611b148fd6f9081a7b070fda007fd85c22bdf083becde1ab1650073189e

Observation c0415548-f576-48ce-9fe5-49ce3369f513 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.619448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.619448Z digest=sha256:2b7a013febde7f8f51f316194dc7905cc9a8dc9b8449ffbe590167a4f44eb2b7

Observation a172eb7d-55fd-44ba-bfed-53294dd467ab · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training The fineweb datasets: Decanting the web for the finest text data at scale

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.130299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.624270Z digest=sha256:99a518ad5cf4b98ed64f4a3944bd7d80656b088ca1583d02bc7298679f096dbc

Observation 59d0637b-9a1c-4435-8a3b-ae56446432c7 · outbound

This paper cites Fineweb2: A sparkling update with 1000s of languages, December 2024 b.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Fineweb2: A sparkling update with 1000s of languages, December 2024 b

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.117060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.628069Z digest=sha256:2d179c7f9c0eddde25594b25b7e204c95f383b3ddd09f4adf6cd71b2b4389060

Observation 9cfb4d22-9e46-4d58-95ff-99993395e5d3 · outbound

This paper cites Language models and their data.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Language models and their data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.104149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.631904Z digest=sha256:0637e7789eee240b1150a4f7aa6a04af45bd17680f644cab67adf73eba5dd30f

Observation c5aaea64-6853-43ce-8e40-a612d50e6cc5 · outbound

This paper cites Sharegpt-chinese-english-90k bilingual human-machine qa dataset.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Sharegpt-chinese-english-90k bilingual human-machine qa dataset

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.092321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.635706Z digest=sha256:5ae21d650da0ea8562f949e25c00fdf3c53c8b595feb2d018cccdecad3031264

Observation 562c4900-cee3-4349-b29c-0efa4cff068e · outbound

This paper cites Industrycorpus2, 2024.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Industrycorpus2, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.081051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.639677Z digest=sha256:dd9031fffa00e6a02faf019c091dd6cb47a5505d9366457d42a0779412a59307

Observation aff800a6-727b-4dd7-9f56-df9c200df215 · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.643572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.643572Z digest=sha256:03d6d05e0901555349c9828098a5b8a8a2c59fb3e0bdff59610fbe256e5b2793

Observation ea67c111-26e3-4e87-a900-f0fb6053bc11 · outbound

This paper cites Common crawl - open repository of web crawl data, 2024 a.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Common crawl - open repository of web crawl data, 2024 a

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.063228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.647757Z digest=sha256:4d4b73c4a88294a0464a151ac21715da5f4a69bdc391861565293399b456ae3b

Observation e853389e-a6fc-47e3-b417-2ef59e45e133 · outbound

This paper cites Qwen2.5: A party of foundation models!, 2024 b.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Qwen2.5: A party of foundation models!, 2024 b

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.051310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.652258Z digest=sha256:420b9059ed77a294396f96aea49e6443e359b08457734f38a505229f6c64a103

Observation 622af388-ca6b-4f89-8b35-c2534e68f253 · outbound

This paper cites CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.656765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.656765Z digest=sha256:37af8a31f282dae5803cf4912a62a5479869ff4ca3d1e48f3b3138cbcf778f2f

Observation 2dd0659a-1236-4b1e-a333-cb2e4aca6d97 · outbound

This paper cites an unresolved cited work.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:33:21.039806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.661112Z digest=sha256:9326cdb82109b098d23a37d6d2a3b52266d0f2ba32d0cbf639a254f1ea76abf9

Observation af94c9d8-deb6-4ccb-9dea-9f7b0bc3ee82 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.664755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.664755Z digest=sha256:69bfaa0ada94a6c8dc20f5a4d49d0f3edd979a33416bbec6ce3055f78e5135b2

Observation e06830d6-68c1-47de-9357-149edc5570f1 · outbound

This paper cites Telechat technical report, 2024 b.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Telechat technical report, 2024 b

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:33:21.027623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T20:33:20.669612Z digest=sha256:a75247f07610719bc369eb15ce37e78b07c536d093e26c180781f9ca31244699

Observation 3d3a9309-cb42-4556-943a-8423d4eaba9f · outbound

This paper cites Skywork: A more open bilingual foundation model, 2023.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Skywork: A more open bilingual foundation model, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.673434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.673434Z digest=sha256:0dbbf13202f10a6223946200d857adba3ac32eda74c09e91ff76ad61cfa97105

Observation 9b5d76cc-8a1a-4ab4-a0ae-530411084e31 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.676873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.676873Z digest=sha256:603e2aa1d734975e65e20163e2d248ef424225bff46162c9c99291b74ddfe96e

Observation 2e4d847b-9c20-4bd8-a40c-7c75e8822fa3 · outbound

This paper cites Qwen2 Technical Report.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.681836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.681836Z digest=sha256:7b946c615d7aba488c38d40f20476b0e9effa9c09f7ecbad1eabc75c6860fff0

Observation 6a657123-0174-4fc9-904f-3f308254d554 · outbound

This paper cites Retrieve anything to augment large language models, 2023.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training Retrieve anything to augment large language models, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.686707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.686707Z digest=sha256:82bf6142506e12f748ffbb47b2a40c5e578f25d5529dccb896b36a0664748ab1

Observation 96068aa7-d718-434a-810d-999b2668c3b2 · outbound

This paper cites mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.690921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.690921Z digest=sha256:0ba5a0fd978953cc2c9b88afccbbdaa68d97f6d5baa1468b620c03fb22af7fa6

Pith citing papers

Observation 83ba508b-320e-4987-af8f-72a706b18958 · inbound

Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data cites this paper.

Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:10:21.973692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:10:21.973692Z digest=sha256:1fd9f6d2e2ca2f72b491a3a5ba8a6a1fd4da908bf2866621df4cb4afdc48295d

Observation d6180338-c7c0-4590-8908-2a9abcfcff23 · inbound

Assessing the Role of Data Quality in Training Bilingual Language Models cites this paper.

Assessing the Role of Data Quality in Training Bilingual Language Models OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:57.893420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:57.893420Z digest=sha256:48cfc835af232fc4423810ec86f45a68abde8a95467e2e1dbf8ed31d3be977b2

Observation 686c5cde-bfab-4615-81e1-9f26e7b14129 · inbound

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling cites this paper.

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-06T18:34:03.468882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:34:03.468882Z digest=sha256:c6922ddb5acd7ca3540aba470420056900a98b398984f4cf52661fdf3b2d9521

Observation e5127bb7-2320-4fb8-ad0c-eb7eac77e0c6 · inbound

Speculative Decoding and the Curse of Multilinguality cites this paper.

Speculative Decoding and the Curse of Multilinguality OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.180961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T07:15:35.673216Z digest=sha256:08aafc5e3fc6798b2278ba90ee4034616916d51de4086c7008c13f7e75c6a836

Observation ab6c9bea-e6cc-48d0-8928-d72dcd261db2 · inbound

Maestro: Workload-Aware Cross-Cluster Scheduling for LLM-Based Multi-Agent Systems cites this paper.

Maestro: Workload-Aware Cross-Cluster Scheduling for LLM-Based Multi-Agent Systems OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.101862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:09:06.904885Z digest=sha256:9fa68445a038a4f94469fc38a47255bf687d706d5b5c38565a6ec6eaf9983cbe

Observation 923645f8-decc-4409-8fd1-0c764f070ded · inbound

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation cites this paper.

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-27T00:40:18.530335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T00:33:18.513616Z digest=sha256:92be9ecb57ae0f03ae5900ba011959be49f98730b588318da7052980daeac830

Observation a1be480d-9b52-4913-8e55-650ef7ec0d5b · inbound

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators cites this paper.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.456364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.456364Z digest=sha256:69ae67159d4a80112eee42c878ff621512b2723c35614756996cabb56e9ca3eb