Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.09400.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:37:10.214011Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:39:51.301876Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ca0e8629-1ad9-4c67-808a-1fe0e87f714d · inbound
Yi: Open Foundation Models by 01.AI CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation be8fd1a6-d041-4478-910c-ca51400aec8a · inbound
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9e0668-09e6-4bad-b50c-19e5ed4eeec1 · inbound
Yankari: A Monolingual Yoruba Dataset CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140f1c8f-5874-4032-9b41-a7869637285c · inbound
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde37eea-8663-49b3-b9d6-d5ce405ff8c8 · inbound
Small Languages, Big Models: A Study of Continual Training on Languages of Norway CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dac7b2e-0187-4538-833f-030e5daab34a · inbound
SnakModel: Lessons Learned from Training an Open Danish Large Language Model CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0b91a7-a13d-48ab-82f4-774514d268fa · inbound
Analysis of Indic Language Capabilities in LLMs CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d9717a0-9c4a-49ef-b92b-fce96e5a38ea · inbound
Kuwain 1.5B: An Arabic SLM via Language Injection CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8786328-0db2-4dc0-8151-e9e2c3f9fbb5 · inbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8619775-8d88-48de-aede-55cc71934158 · inbound
Full-Parameter Continual Pretraining of Gemma2: Insights into Fluency and Domain Knowledge CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d903287-defd-46c8-b5b5-b8873a81dfec · inbound
Synthetic Document Question Answering in Hungarian CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8503dc19-4009-46da-8a76-2c49283237d5 · inbound
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d26e90d6-4e49-4322-9b63-b76a565e415b · inbound
TokAlign: Efficient Vocabulary Adaptation via Token Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53bcebbb-a4be-48b0-a779-c31f70122194 · inbound
GeistBERT: Breathing Life into German NLP CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abadba50-c7a0-4135-aae2-4cb28c47fe71 · inbound
Semantic Outlier Removal with Embedding Models and LLMs CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a241d9d-2f6e-4ed5-aec9-41ee2fe130fa · inbound
Mangosteen: An Open Thai Corpus for Language Model Pretraining CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d4e365-3e12-4033-8bb8-f0b0d237ee85 · inbound
Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab296b74-5501-4841-819f-f5cf2a67d0c3 · inbound
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a8873b35-05f4-4f30-9077-b5adf23bd945 · inbound
TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 97dccc55-727c-44b0-ae83-fb2624bdc845 · inbound
TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f76e6b07-e605-426c-8c2b-353fc618c1fe · inbound
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 255e4fe2-17bc-4f35-9f87-869f4ac4fa87 · inbound
The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3d6325ca-f10c-4b6b-85a8-1563608b8c54 · inbound
Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 945dcc70-a317-4826-b16e-5b7992a4ca69 · inbound
MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1f33e711-3560-4865-ac4a-3ee19212a5bc · inbound
Explicit Boundary Markers for Subword Vocabularies CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.