Pith. sign in

Paper Citation Record · LEDGER

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2501.08197.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08197 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:42:57.893420Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d6180338-c7c0-4590-8908-2a9abcfcff23 · inbound

Assessing the Role of Data Quality in Training Bilingual Language Models cites this paper.

Assessing the Role of Data Quality in Training Bilingual Language Models OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:57.893420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:57.893420Z digest=sha256:a98172be0d24e2b914ac3b795857087166e6cb38bce2559f6321c51cc0cfdd4d

Observation 686c5cde-bfab-4615-81e1-9f26e7b14129 · inbound

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling cites this paper.

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-06T18:34:03.468882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:34:03.468882Z digest=sha256:9c1ca5e559771d20d966ff3252367d5c264b018c2b36fe359c7735873ee71385

Observation e5127bb7-2320-4fb8-ad0c-eb7eac77e0c6 · inbound

Speculative Decoding and the Curse of Multilinguality cites this paper.

Speculative Decoding and the Curse of Multilinguality OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.180961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T07:15:35.673216Z digest=sha256:bce737381966e5f80fdd04981d01e3932d7991d964f9d9c5a7520e0f10cddb3c

Observation ab6c9bea-e6cc-48d0-8928-d72dcd261db2 · inbound

Maestro: Workload-Aware Cross-Cluster Scheduling for LLM-Based Multi-Agent Systems cites this paper.

Maestro: Workload-Aware Cross-Cluster Scheduling for LLM-Based Multi-Agent Systems OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.101862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:09:06.904885Z digest=sha256:0e418b44e3b3d48fefd966f7ac5d42c097ad470c1e96048c8e1c0df06a51898b

Observation 923645f8-decc-4409-8fd1-0c764f070ded · inbound

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation cites this paper.

ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-27T00:40:18.530335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:33:18.513616Z digest=sha256:03ffbf3ba69dd7bcdb0ad24f4f14638ae73f5628608ca06801cbc955f77f097f

Observation a1be480d-9b52-4913-8e55-650ef7ec0d5b · inbound

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators cites this paper.

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:29.456364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:29.456364Z digest=sha256:bc287e384616dd0f50e90bda21b1fd97d0f196de0cbb54184da9299c77d4978e