Pith. sign in

Paper Citation Record · LEDGER

Position: The Most Expensive Part of an LLM should be its Training Data

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 5 inbound Pith citation observations for arXiv:2504.12427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12427 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:34:53.589828Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:44.273599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:39:46.610216Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ffee95a-85f2-44e2-8647-575ce81577e8 · outbound

This paper cites https://huggingface.co/collections/r-three/common-pile-665a13e48528df6b00416dc0, 2025.

Position: The Most Expensive Part of an LLM should be its Training Data https://huggingface.co/collections/r-three/common-pile-665a13e48528df6b00416dc0, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.222109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.416080Z digest=sha256:4cae719cfa940e4638589f1e5ffdefb5e932f464941d590e0baf37e8c2e297bc

Observation c0836bd0-436d-4232-a7b5-5f1de43df576 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Position: The Most Expensive Part of an LLM should be its Training Data Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.420387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.420387Z digest=sha256:b0dcc77c3d1ff09ec0ad415d91c1a53e19c10ef56ff12d927c064f9adb1a4ffb

Observation 29a1ac42-8e9a-48f9-bc48-2966b1f50e70 · outbound

This paper cites Phi-4 Technical Report.

Position: The Most Expensive Part of an LLM should be its Training Data Phi-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.424961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.424961Z digest=sha256:fee0e0dc2cf8843c35a4ecdd2b40e880bcee3b6e915563c5688861db2d7c4494

Observation 94a2aa78-8802-47da-b446-a0bfd1abc6e2 · outbound

This paper cites Alden newspapers v.

Position: The Most Expensive Part of an LLM should be its Training Data Alden newspapers v

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-16T12:34:53.976485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.429188Z digest=sha256:f65fadc29819ddc684aab02887114461aa52156237bb50fc0c2b50a2a9e96beb

Observation 60d5f8ad-9b5f-4da0-bfcd-50cdfb622b42 · outbound

This paper cites M., and Weber, G.

Position: The Most Expensive Part of an LLM should be its Training Data M., and Weber, G

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.433029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.433029Z digest=sha256:3c707b522521a55dabd69a2026be63f007fd1103eec072b67f74f13dfd83e4f5

Observation 7adfd76b-b4c3-45f7-b7cc-788c976aecab · outbound

This paper cites Authors guild v.

Position: The Most Expensive Part of an LLM should be its Training Data Authors guild v

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.206262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.437311Z digest=sha256:2a31083d18018946484d2f35483f8effd8ca263b208e869d31301a9b175a03a4

Observation 72bb5d94-82ea-4e6d-abb9-10b563348510 · outbound

This paper cites What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions.

Position: The Most Expensive Part of an LLM should be its Training Data What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.441434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.441434Z digest=sha256:540152229fc69e73f5dcfa79991bb175401be0c957a0c6ec0281c224d9353a44

Observation 558c1ceb-306d-4caa-86a9-1ef6e5f74c7d · outbound

This paper cites Common crawl dataset.

Position: The Most Expensive Part of an LLM should be its Training Data Common crawl dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.196771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.445464Z digest=sha256:e56403b756fd31ba587e10478fac3f40d302323e8eebfb3ab252ac5a08d12344

Observation 3518d52d-387a-4a8d-9541-f23ebcd8f2f9 · outbound

This paper cites Concord music group v.

Position: The Most Expensive Part of an LLM should be its Training Data Concord music group v

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.187530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.449194Z digest=sha256:4d4f4f2d0c16b75c6e1ea315fc391f905859ada2e7c2f1ad3eb637399cb90d9d

Observation 252eac39-7388-4717-be9e-6a70a8dc6d8e · outbound

This paper cites The rising costs of training frontier AI models.

Position: The Most Expensive Part of an LLM should be its Training Data The rising costs of training frontier AI models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.452722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.452722Z digest=sha256:b253dbd879f7f133f5d35c759a15796179dd729eb82bb55af17ce560717a461a

Observation cbefb914-e2ca-4d11-93c0-b8d709d75485 · outbound

This paper cites Ai is a lot of work.

Position: The Most Expensive Part of an LLM should be its Training Data Ai is a lot of work

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.177432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.457066Z digest=sha256:f7629a6ff677dbd339c58f4abb9015973d5861444bbd0f6c62cffc377ccc1c78

Observation 0b9777a9-53d8-4c17-988e-4a47343008d5 · outbound

This paper cites Encyclopedia britannica 15th edition.

Position: The Most Expensive Part of an LLM should be its Training Data Encyclopedia britannica 15th edition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.167706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.460931Z digest=sha256:901c8d360f91e1d5bdfa9e64e195ddf8b829fa43f48159002107820ebc189649

Observation 9ac7fb88-d47c-487b-9be8-f0936f3ab916 · outbound

This paper cites Data on notable ai models, 2024.

Position: The Most Expensive Part of an LLM should be its Training Data Data on notable ai models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.157863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.464492Z digest=sha256:cca92046b99082a8239d70cd2894173e7a14e066a5a92fc0db3a3c52d54f32f4

Observation c957d152-6d94-4247-b56b-21a7ecc89e37 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

Position: The Most Expensive Part of an LLM should be its Training Data DataComp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.468171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.468171Z digest=sha256:e2c6bbb51c9c5b6394a1b3422d8c33691628a490d3ab9754a41b2c83196d6ac3

Observation 01f09962-6c0c-46ab-a7ef-e7905f4ce083 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Position: The Most Expensive Part of an LLM should be its Training Data Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.471844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.471844Z digest=sha256:694835b28b15ab70953d34afc9148ad6a7ec4e702c8d91aa83926586e9cb2451

Observation cba7a639-6920-48ad-a543-dbde1f5b53a7 · outbound

This paper cites Data Shapley: Equitable Valuation of Data for Machine Learning.

Position: The Most Expensive Part of an LLM should be its Training Data Data Shapley: Equitable Valuation of Data for Machine Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.475484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.475484Z digest=sha256:307a8b0ba4697d8e36049e656088cf5b715542e64a2d9bb7d58a99e1799ae3e1

Observation 03d72fee-497e-469b-a1cd-5e46266b1faa · outbound

This paper cites Evaluation of Similarity-based Explanations.

Position: The Most Expensive Part of an LLM should be its Training Data Evaluation of Similarity-based Explanations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.479173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.479173Z digest=sha256:e7f07f265f8d110f046231baa472c06defd68df32da611084ebb2215638b1ff9

Observation 95e6842e-d01d-40fb-95c2-875558897655 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.482888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.482888Z digest=sha256:4e8cde81efa8fe082da674fb72423e8812c4c8fbdf4d19591aab92d093b9611f

Observation e0e7818b-ac76-4f46-a3a7-5c125e939bf3 · outbound

This paper cites Statistics on wages.

Position: The Most Expensive Part of an LLM should be its Training Data Statistics on wages

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.147863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.487229Z digest=sha256:8f07cc971485c709f9e58647d7fc768c41f1d78315b17dfab3e156b2401465ae

Observation 6bad627c-5881-4537-af10-2065619543c6 · outbound

This paper cites Scaling Laws for Neural Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Scaling Laws for Neural Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.490580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.490580Z digest=sha256:9b76725955f64cf6f08dca4e1c8ab23ab4044eb8824eadfa3ec263865b79669e

Observation b779aefe-5e3e-4bf2-9355-d243a13effb7 · outbound

This paper cites Understanding Black-box Predictions via Influence Functions.

Position: The Most Expensive Part of an LLM should be its Training Data Understanding Black-box Predictions via Influence Functions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.494514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.494514Z digest=sha256:22d34c2edc636784b860b9052e300557c8135710e8256f9ae0e4be1f2dc43117

Observation dc20d4d5-255e-42cc-80ea-f82e155d7c85 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Position: The Most Expensive Part of an LLM should be its Training Data OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.498197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.498197Z digest=sha256:9a2984d4402f8f7a8036d503370a624aa531d856d25d9e35507ba0da6950afab

Observation 2d526e30-7940-40ee-9d3b-5bf491476136 · outbound

This paper cites Releasing Common Corpus: the largest public domain dataset for training LLMs.

Position: The Most Expensive Part of an LLM should be its Training Data Releasing Common Corpus: the largest public domain dataset for training LLMs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.138207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.501877Z digest=sha256:372d78c3c7f264d1dac33f1e6ba9153ffafeeab4374cfdba027c6d60c2c1c7ce

Observation b173a557-7554-43ba-9c09-4b38c6b0d91f · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Position: The Most Expensive Part of an LLM should be its Training Data DataComp-LM: In search of the next generation of training sets for language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.505422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.505422Z digest=sha256:b3f24c670b69c1a37e4a95ca2f3d9238468a3574895f8b1a7ca9317681574599

Observation 589d0d14-d539-4794-a32c-62452bcf0ab2 · outbound

This paper cites Consent in Crisis: The Rapid Decline of the AI Data Commons.

Position: The Most Expensive Part of an LLM should be its Training Data Consent in Crisis: The Rapid Decline of the AI Data Commons

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.509363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.509363Z digest=sha256:50dc12314a74f51199ca57e6dacb6fe49ebfafe845c12cccbc6512f4a8847cbf

Observation b9175506-313c-4a0e-88e3-645e166e0898 · outbound

This paper cites SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore.

Position: The Most Expensive Part of an LLM should be its Training Data SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.513176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.513176Z digest=sha256:effd3de015d62cfd66adc4167c23f53d6cf2fa627a9821335883100a25cc4ea6

Observation 50021800-f7b3-4840-8bfa-76e96383ddf3 · outbound

This paper cites AI models that cost \ 1 billion to train are underway, \ 100 billion models coming — largest current models take “only” \ 100 million to train: Anthropic CEO.

Position: The Most Expensive Part of an LLM should be its Training Data AI models that cost \ 1 billion to train are underway, \ 100 billion models coming — largest current models take “only” \ 100 million to train: Anthropic CEO

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.128186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.516963Z digest=sha256:a530df2c8d193d5ce919db4399c571e33ae69061c333f594f4006a1d53bf6704

Observation 34d09960-edb5-43cd-87e2-dfe3162eaac7 · outbound

This paper cites Scaling Data-Constrained Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Scaling Data-Constrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.520465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.520465Z digest=sha256:9b192f17afeced62cf9f1f71722094c364c423852d5738c8f14819f684ae3d20

Observation 51e8f85b-56e4-42ca-806b-f3d77227f000 · outbound

This paper cites New york times v.

Position: The Most Expensive Part of an LLM should be its Training Data New york times v

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.117493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.524159Z digest=sha256:9d730167704cf99227c2418c109065cab63ce0933092270dd80bf52eb312fab1

Observation 42b77bc3-054d-4116-bc9e-b719feca545f · outbound

This paper cites TRAK: Attributing Model Behavior at Scale.

Position: The Most Expensive Part of an LLM should be its Training Data TRAK: Attributing Model Behavior at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.527627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.527627Z digest=sha256:0d80c228fc9f0812afbf39d1e2cabf719cd4d8c415c56413e5fe70e1b820da39

Observation 1c16cddf-21d2-4a9b-abad-7c7caead9e09 · outbound

This paper cites an unresolved cited work.

Position: The Most Expensive Part of an LLM should be its Training Data Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:34:54.107235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.531585Z digest=sha256:db362f6698822c9ee521ffc961a20c06f8835adf0f98d2c6b87154161bc038ee

Observation 4e90d540-21fa-4d3f-a022-c765ce57d782 · outbound

This paper cites Multiple ai companies bypassing web standard to scrape publisher sites, licensing firm says.

Position: The Most Expensive Part of an LLM should be its Training Data Multiple ai companies bypassing web standard to scrape publisher sites, licensing firm says

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.096999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.535354Z digest=sha256:12db27c3a8d4cffb180023f560c58b68f5e74ec1ed8b198234660d05b1802ed4

Observation 6bc0e423-6cf3-4234-9b54-9de913a1b9e9 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Position: The Most Expensive Part of an LLM should be its Training Data The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.538809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.538809Z digest=sha256:649693a9928bc954b6134bd3a16aec2b2327d6b10826d50f634bc30ba120f462

Observation 117ed8c8-2b11-4ca1-9282-57fa9ad774a7 · outbound

This paper cites Estimating Training Data Influence by Tracing Gradient Descent.

Position: The Most Expensive Part of an LLM should be its Training Data Estimating Training Data Influence by Tracing Gradient Descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.542296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.542296Z digest=sha256:57082d85653c01c73c3e340106d30a8a140dcf881b9154d1c6a3c6186ac0e46b

Observation 8f50bfd0-16cd-479b-85d2-b1164801cce4 · outbound

This paper cites Reddit and openai build partnership.

Position: The Most Expensive Part of an LLM should be its Training Data Reddit and openai build partnership

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.086656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.546327Z digest=sha256:72af39d831073d0d5d957b9f1adc5287bb8fcc1bbb937900b7b46256ff403807

Observation a27c7275-92b1-4929-a0ce-de699e929c1d · outbound

This paper cites Thomson reuters' adjusted eps beats expectations, ai boosts results.

Position: The Most Expensive Part of an LLM should be its Training Data Thomson reuters' adjusted eps beats expectations, ai boosts results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.075904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.549780Z digest=sha256:0d62d7332fca572638a184d815926a1bab449bb3853efa6e1b8663178c54375e

Observation 22b3cdde-bdef-4cd2-b62d-524a157ec2fd · outbound

This paper cites Shutterstock expands partnership with openai, signs new six-year agreement to provide high-quality training data.

Position: The Most Expensive Part of an LLM should be its Training Data Shutterstock expands partnership with openai, signs new six-year agreement to provide high-quality training data

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.064402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.553071Z digest=sha256:b307b62b927577a7225b9c2690e67d4fe224aa350687bc9613d8309ea6f1d4e6

Observation 56e88207-67cb-4cda-b630-7c7928ad4f1e · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Position: The Most Expensive Part of an LLM should be its Training Data Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.556356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.556356Z digest=sha256:a445b11211ccd2778f527fadfd50e6bc4c7acaa13bed7eba3a48e71dcbbf9adc

Observation 5c5e27f6-786f-482f-8476-120d83c243b4 · outbound

This paper cites The atlantic announces product and content partnership with openai.

Position: The Most Expensive Part of an LLM should be its Training Data The atlantic announces product and content partnership with openai

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.052727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.560435Z digest=sha256:8783f103b43b6055de21f1d2946a30a3b8056b06652729d33c969684e3284e85

Observation 02da8847-5bea-46ec-90e4-ab9716fc5df4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.563650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.563650Z digest=sha256:3ed1d6fdbf9920e5e179ee685d51c4c7df0cf3cd3ead15cf2b85061c13ebd971

Observation dde8342c-841c-44e7-9816-8d01bed89acc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Position: The Most Expensive Part of an LLM should be its Training Data Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.567083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.567083Z digest=sha256:6dcc8da27ae87b3dd3a9dfca15678781799c920ae4a10429697cc0b2fb72b83b

Observation cbf0bcf5-7040-43f4-ae36-73e83a968d80 · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Position: The Most Expensive Part of an LLM should be its Training Data Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.570304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.570304Z digest=sha256:7d4a0ff289ee7d653f4da799b33331c9c8c8655f2b53581c4a70d5065152e5e6

Observation d74ac334-eff0-41e9-a858-6d7d0eb6f15b · outbound

This paper cites Vox media and openai form strategic content and product partnership.

Position: The Most Expensive Part of an LLM should be its Training Data Vox media and openai form strategic content and product partnership

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.042052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.573730Z digest=sha256:d20fb8b5408aa77365c90b8754af0e509f937ef1d8a1dd4838fbc8920436d82f

Observation 3b58b3df-6705-40f0-bd7c-321f26fd55e5 · outbound

This paper cites Html standard.

Position: The Most Expensive Part of an LLM should be its Training Data Html standard

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.030794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.576828Z digest=sha256:b3d4cc64345b869d1980f4389673fbfdc09728eca94e31286bd42e7d126e91d8

Observation 78226ba2-210d-4061-9217-5374472983b3 · outbound

This paper cites Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus.

Position: The Most Expensive Part of an LLM should be its Training Data Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.580062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.580062Z digest=sha256:f61f488819b0eca72db6bf9a3c0a8ea22634087b4d0c9d864eeb47f65e1c6429

Observation d9761ef1-9d90-41fb-b3ac-18d1a352da19 · outbound

This paper cites Breaking down emerging segments in the ai content licensing landscape.

Position: The Most Expensive Part of an LLM should be its Training Data Breaking down emerging segments in the ai content licensing landscape

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.019864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.583589Z digest=sha256:27d1776dcd67504389199195fb9879efebd1d472f629c6cbe045ff4f6d2d325a

Observation e321c3e6-f2cd-44a7-b3d7-749b33eb06b1 · outbound

This paper cites Articulating value from data.

Position: The Most Expensive Part of an LLM should be its Training Data Articulating value from data

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.007464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.586790Z digest=sha256:863d6aa0399ebbeb5ff511acdbd1520afa36ce8ec9ec083a91b919846cede417

Observation d301cb7b-75c5-43ff-bab2-da3c74223efc · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Position: The Most Expensive Part of an LLM should be its Training Data WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.589828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.589828Z digest=sha256:ec6225427e30a90d8f8035b4ead86b4788d4c059cf46ef1187d6a9f6e78c082c

Pith citing papers

Observation 179fac02-40ff-4f0c-a983-85f16f6422af · inbound

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text cites this paper.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text Position: The Most Expensive Part of an LLM should be its Training Data

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.273599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:44.273599Z digest=sha256:fc2c6da1f3d0115ec2bd97a43109b6f74df6631956c07e4d7f146202a5ca144d

Observation 7895a035-1447-469b-a995-8b83e5474abe · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:25:55.503750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T14:21:53.325788Z digest=sha256:3a8ec45cd09abebc691f90e9d8e04c75aa8cb1b5425d56cf924961358625226e

Observation de189bdf-7b49-496d-8753-7ab89b8483e5 · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T12:17:47.123519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:17:47.123519Z digest=sha256:07abb9b56f377dcf9e621215f27a9eb504815fdfc14cc219d615ad872aad672e

Observation f2438a06-593e-4df0-9cc9-bd8ff63152eb · inbound

On the Fragility of Data Attribution When Learning Is Distributed cites this paper.

On the Fragility of Data Attribution When Learning Is Distributed Position: The Most Expensive Part of an LLM should be its Training Data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:17:39.452630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T15:16:32.125215Z digest=sha256:acb9d9389d3e53237064fbdd783e4b20a7e48d09067b26eed64560cffe0e1929

Observation fc6c92ef-bb74-4ea7-8367-1d7781beffa3 · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:46.611909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T07:52:24.053501Z digest=sha256:3b0a8ce254597f16beb53187836a1e90c557d8ed256f4f5760b5d40ebbf7df3b