Pith. sign in

Paper Citation Record · LEDGER

CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.09400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.09400 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:37:10.214011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:51.301876Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca0e8629-1ad9-4c67-808a-1fe0e87f714d · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:27.886303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:3510cf3d6565a1a403fee0c887f146fa8d1d3c36c3784bf4bb598f92526de98d

Observation be8fd1a6-d041-4478-910c-ca51400aec8a · inbound

CamemBERT 2.0: A Smarter French Language Model Aged to Perfection cites this paper.

CamemBERT 2.0: A Smarter French Language Model Aged to Perfection CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:18:14.602724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:18:14.602724Z digest=sha256:200ab6cd7388b6e3b179c53defbb5a245172a5bdf20c8459d60a32f730f82f6b

Observation 8e9e0668-09e6-4bad-b50c-19e5ed4eeec1 · inbound

Yankari: A Monolingual Yoruba Dataset cites this paper.

Yankari: A Monolingual Yoruba Dataset CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:33:54.722020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:33:54.722020Z digest=sha256:26da66992f65ad26bccea3ccf2e7b53ffd060929af98fb85ed9314fb3a323744

Observation 140f1c8f-5874-4032-9b41-a7869637285c · inbound

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic cites this paper.

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:27.032045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:37:27.032045Z digest=sha256:30cde7f5af1d7c7a83094d99c50b2bccf46b35ac0fa98271473bf2545383b0d6

Observation cde37eea-8663-49b3-b9d6-d5ce405ff8c8 · inbound

Small Languages, Big Models: A Study of Continual Training on Languages of Norway cites this paper.

Small Languages, Big Models: A Study of Continual Training on Languages of Norway CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T19:40:51.677374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:40:51.677374Z digest=sha256:a38905471a0481b06d86292a7e45d26b46d243dbf16e3e49c110dfd4fe8f5060

Observation 3dac7b2e-0187-4538-833f-030e5daab34a · inbound

SnakModel: Lessons Learned from Training an Open Danish Large Language Model cites this paper.

SnakModel: Lessons Learned from Training an Open Danish Large Language Model CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T13:37:38.519174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:37:38.519174Z digest=sha256:5b5504fac6134a62dfc4ff183f4a8dc1d44e3f09dff37be46cf5508cdca8b26e

Observation fb0b91a7-a13d-48ab-82f4-774514d268fa · inbound

Analysis of Indic Language Capabilities in LLMs cites this paper.

Analysis of Indic Language Capabilities in LLMs CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:10.697511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:10.697511Z digest=sha256:3531fcf3ba4aff31693aaad7e9648082c2dd1fc3f4c04330969aa8b9cb732fe0

Observation 4d9717a0-9c4a-49ef-b92b-fce96e5a38ea · inbound

Kuwain 1.5B: An Arabic SLM via Language Injection cites this paper.

Kuwain 1.5B: An Arabic SLM via Language Injection CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:37:10.214011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:37:10.214011Z digest=sha256:80890be0b6562ff635482ceb4c2d320343412860f061dbc194062bebf0f6c22d

Observation f8786328-0db2-4dc0-8151-e9e2c3f9fbb5 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.325608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.325608Z digest=sha256:2df378abc20bbef18da2e8799a9467188fa1c0e78ed8e0229ffece01795f7d61

Observation f8619775-8d88-48de-aede-55cc71934158 · inbound

Full-Parameter Continual Pretraining of Gemma2: Insights into Fluency and Domain Knowledge cites this paper.

Full-Parameter Continual Pretraining of Gemma2: Insights into Fluency and Domain Knowledge CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:55:45.930374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:55:45.930374Z digest=sha256:d39102531f334e613a29ffe12f18170155f9ffd8160c8ed450dec4c91a969e1e

Observation 1d903287-defd-46c8-b5b5-b8873a81dfec · inbound

Synthetic Document Question Answering in Hungarian cites this paper.

Synthetic Document Question Answering in Hungarian CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:25.835410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:25.835410Z digest=sha256:d972c410a60026fba8fa6c00109900bef382b48f76c7844d198957d9e78af03c

Observation 8503dc19-4009-46da-8a76-2c49283237d5 · inbound

BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization cites this paper.

BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.714800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.714800Z digest=sha256:f31061c9c3bcbca042d3fa149c140f7633045aec217b1dde94e398a9a7885fad

Observation d26e90d6-4e49-4322-9b63-b76a565e415b · inbound

TokAlign: Efficient Vocabulary Adaptation via Token Alignment cites this paper.

TokAlign: Efficient Vocabulary Adaptation via Token Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:28.059185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:06:28.059185Z digest=sha256:3f6d076c85622ab71b17008d38dc897a55ac9ee97b8222250640826cdded9cfa

Observation 53bcebbb-a4be-48b0-a779-c31f70122194 · inbound

GeistBERT: Breathing Life into German NLP cites this paper.

GeistBERT: Breathing Life into German NLP CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:59.129461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:59.129461Z digest=sha256:b17faa695d0bb9da4c301ab1422ef34194bb97c936e5499e02f120524edbdebe

Observation abadba50-c7a0-4135-aae2-4cb28c47fe71 · inbound

Semantic Outlier Removal with Embedding Models and LLMs cites this paper.

Semantic Outlier Removal with Embedding Models and LLMs CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:25:36.702359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:25:36.702359Z digest=sha256:1fe76040ba67d8e1b0a6e18dce4b94e64fab80e50aa0e11e12967baeacd8c040

Observation 8a241d9d-2f6e-4ed5-aec9-41ee2fe130fa · inbound

Mangosteen: An Open Thai Corpus for Language Model Pretraining cites this paper.

Mangosteen: An Open Thai Corpus for Language Model Pretraining CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:56:23.835272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:56:23.835272Z digest=sha256:03a7f2361a1f2c5635db57ae1bbb280c92347c00b7c4a63bc3954ec751f787bf

Observation e2d4e365-3e12-4033-8bb8-f0b0d237ee85 · inbound

Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms cites this paper.

Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T21:15:37.405672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:15:37.405672Z digest=sha256:620e6bd6d0fdb837bdf0cdc7bda735eda36e11c062235b0b32b46782dfdda6c5

Observation ab296b74-5501-4841-819f-f5cf2a67d0c3 · inbound

Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance cites this paper.

Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:16:09.169879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T15:02:54.374120Z digest=sha256:189bb0f68f2c7c10be3af4ef29f9bb32c2a3401f9413e68a6119853cb9dee6a6

Observation a8873b35-05f4-4f30-9077-b5adf23bd945 · inbound

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment cites this paper.

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:37:52.257765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-14T19:35:10.026293Z digest=sha256:d62d49212da33e1f2586cda240a602d9e367be60133be45442350b657451aed6

Observation 97dccc55-727c-44b0-ae83-fb2624bdc845 · inbound

TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale cites this paper.

TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:41.547085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T17:06:12.280460Z digest=sha256:78618fbe8b290d2393cf04e61e18503c229ff0f9982bb08976b82ec2e2963dd1

Observation f76e6b07-e605-426c-8c2b-353fc618c1fe · inbound

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE cites this paper.

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:13.613153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T11:09:22.027588Z digest=sha256:a836de4d4fd608d460ad3a6cc224261654dffe89f6539cd534f8e58027c7cc62

Observation 255e4fe2-17bc-4f35-9f87-869f4ac4fa87 · inbound

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty cites this paper.

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:14:40.882589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T13:08:33.172423Z digest=sha256:e69e6297c9b1c98383120471c9e99a68523e0f5f1874e42c10b4a3de2671483c

Observation 3d6325ca-f10c-4b6b-85a8-1563608b8c54 · inbound

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet cites this paper.

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:39:34.631117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T16:53:25.556222Z digest=sha256:cf673455636f41ffa5f190d11bd8b25543ff0d4cbd8ac7c3aba1f58c0648f5c0

Observation 945dcc70-a317-4826-b16e-5b7992a4ca69 · inbound

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment cites this paper.

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.304281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T04:59:05.057552Z digest=sha256:4f5f064f76f25486fb52abe70df2de550aca9331f2c007162ace5c8f5ab93c50

Observation 1f33e711-3560-4865-ac4a-3ee19212a5bc · inbound

Explicit Boundary Markers for Subword Vocabularies cites this paper.

Explicit Boundary Markers for Subword Vocabularies CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:08.671910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:32:08.671910Z digest=sha256:a225f580fdddec4a3a6d33cc4a4820fd3b01cc7c34d23e377d993a03e1aa3d26