Pith. sign in

Paper Citation Record · LEDGER

RedPajama: an Open Dataset for Training Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2411.12372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12372 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:09.304268Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4824c5f2-d689-44d9-b29b-4d9ac1db0003 · inbound

Diagnosing our datasets: How does my language model learn clinical information? cites this paper.

Diagnosing our datasets: How does my language model learn clinical information? RedPajama: an Open Dataset for Training Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:09.304268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:09.304268Z digest=sha256:58b90fbd6d5cb1e990ff7e254c511abfcd95d497f591e8ad640973c2ef0d2504

Observation e15629f2-7a49-41eb-85b5-df5ef3621fc5 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding RedPajama: an Open Dataset for Training Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:33.391740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:33.391740Z digest=sha256:e01f9d728ee9c3576318b02a28b3fe7a7d1ba01a3b34c873811158bd69997ecd

Observation 37b1cbea-2d9b-41f2-a510-3b6ce62ca490 · inbound

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text cites this paper.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text RedPajama: an Open Dataset for Training Large Language Models

Reference 196

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:45.444268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:45.444268Z digest=sha256:94d3d87a3ccd530645b4d98c13af224d4da27cb5450f8df8b2008ffa7ea05578

Observation 2a633f51-b59d-4224-8915-6a1fad4c7ad1 · inbound

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training cites this paper.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RedPajama: an Open Dataset for Training Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.350194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.350194Z digest=sha256:da82366b61d26793f4e32ec440025d476eef69bd3ae1c3a8d09310619f7214f4

Observation 566ef656-7006-4020-8388-65edc1ea2007 · inbound

Essential-Web v1.0: 24T tokens of organized web data cites this paper.

Essential-Web v1.0: 24T tokens of organized web data RedPajama: an Open Dataset for Training Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:47.623955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:47.623955Z digest=sha256:d8562a02c840b3365dc4efb9652f04f38d30f5caa7840c53d3ab100e0944ac70

Observation e8d4a238-444a-4897-981e-b963eab614df · inbound

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models cites this paper.

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models RedPajama: an Open Dataset for Training Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:44:26.591858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T23:44:01.953344Z digest=sha256:b994bed94a6b9e9ff6ba7090c441ce6151281abc4fc335be31ef2bf0aa3026da

Observation 9bbf3e5f-ac51-4f10-b045-a2f423eafe06 · inbound

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models cites this paper.

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models RedPajama: an Open Dataset for Training Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:35.952702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:49:35.952702Z digest=sha256:a75c13f089f22dfd9fb62a2325c3b024229dfb84014dbb866c93ac4dbbc46011

Observation c4457a45-cf0d-497e-b605-89b5accb6f50 · inbound

Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling cites this paper.

Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling RedPajama: an Open Dataset for Training Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:36.827145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:36.827145Z digest=sha256:1c95814d91172c748203dc508b4a1d1709012cb5a6a7d69ecc39695fdfdcb62d

Observation 98500d58-abc6-4f10-a38d-02f91e33dfbc · inbound

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining cites this paper.

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining RedPajama: an Open Dataset for Training Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:49:03.027674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T04:46:33.641714Z digest=sha256:6a6d40593a7a02017fd5cacc6e5ccce060d8c66230a035f3f5d9ce8e894a6247

Observation 6b449b42-ea8b-4d18-969d-63fcdb5c7cf9 · inbound

The Effect of Scripts and Formats on LLM Numeracy cites this paper.

The Effect of Scripts and Formats on LLM Numeracy RedPajama: an Open Dataset for Training Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T08:58:51.336919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:58:51.336919Z digest=sha256:a147ba80203b03252436a3c41b52696768b895f53ba4b407a24c46dad50bfbbf

Observation dff2e6a1-fa65-4449-b2f9-f9606e2091bd · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data RedPajama: an Open Dataset for Training Large Language Models

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:34.574812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:34.574812Z digest=sha256:e558526638327dc9a4cb87ae62ec0698601b45e880bb20b0b7925b76cf8f352e

Observation 590a810a-bea8-41b1-9df8-106cc74348e1 · inbound

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings cites this paper.

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings RedPajama: an Open Dataset for Training Large Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.427174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:01:37.613094Z digest=sha256:40b2b60d1106224203ed3413479c3f9052de5a6402f0a896f928f12674fb7166

Observation f10efb69-6579-48cc-bbba-5edf701d6a07 · inbound

Small edits, large models: How Wikipedia advocacy shapes LLM values cites this paper.

Small edits, large models: How Wikipedia advocacy shapes LLM values RedPajama: an Open Dataset for Training Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:35.771485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T09:03:09.977131Z digest=sha256:3ce94c89d21b1c7a67a9ecadb13730e99c0299185f36208d4ab50b77d3a8b294

Observation 32e897c5-9350-474a-a03d-39cf2adbf4ec · inbound

Small edits, large models: How Wikipedia advocacy shapes LLM values cites this paper.

Small edits, large models: How Wikipedia advocacy shapes LLM values RedPajama: an Open Dataset for Training Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T19:23:22.815333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:23:22.815333Z digest=sha256:9918f3c52d964bd1cf45a99cb0d12e21296c9f67cf821395d4848dbb533666bb

Observation 9729b55e-f368-4e28-bc2c-9fd2bdb8eb9a · inbound

Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing cites this paper.

Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing RedPajama: an Open Dataset for Training Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:06:02.750903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T01:14:50.756453Z digest=sha256:54a2ea6d4d8c303ca2b5898f06630658533b569f7e60aa4d010a3af01d0513a1

Observation 152f7a68-96f4-42be-912d-4711e32a7b26 · inbound

$\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space cites this paper.

$\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space RedPajama: an Open Dataset for Training Large Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:55.872510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T12:35:58.613973Z digest=sha256:643cd66b917a7d104bea99275dbbbf2ea1581c7ba757dea16bb030e571d97076

Observation 2a65a684-4ba9-4249-bc4c-5dbb900df86c · inbound

A First-Principles Theory of Slow Thinking and Active Perception cites this paper.

A First-Principles Theory of Slow Thinking and Active Perception RedPajama: an Open Dataset for Training Large Language Models

Reference 167

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:37:03.328504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T11:32:24.374377Z digest=sha256:da1eec825fedbb0902cb4682d6bba03f8061c9c0936c655767d3eb8387c9dee7

Observation 69a6c735-9747-426f-aa56-d90025b34803 · inbound

Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool cites this paper.

Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool RedPajama: an Open Dataset for Training Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:06.880355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:06.880355Z digest=sha256:708bd1c62997dc6cff8f0381f0624dfd5ff6f8a85ef4fcf413dc80e05abcb7d2

Observation 2acd6937-8313-480a-9af2-15d5e44ef445 · inbound

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference cites this paper.

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference RedPajama: an Open Dataset for Training Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T11:19:36.131568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:19:36.131568Z digest=sha256:15b280f22144a4053ff904fc3448772822d7b02ce842a52239f196c81ecd18a6