Pith. sign in

Paper Citation Record · LEDGER

Position: The Most Expensive Part of an LLM should be its Training Data

As of 22 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 5 inbound Pith citation observations for arXiv:2504.12427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12427 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:34:53.589828Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:44.273599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:39:46.610216Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ffee95a-85f2-44e2-8647-575ce81577e8 · outbound

This paper cites https://huggingface.co/collections/r-three/common-pile-665a13e48528df6b00416dc0, 2025.

Position: The Most Expensive Part of an LLM should be its Training Data https://huggingface.co/collections/r-three/common-pile-665a13e48528df6b00416dc0, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.222109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.416080Z digest=sha256:f2a2c20b34d15b8dcbe816dfb87f4eac2a6fa34d772f63568c8771b341c509fe

Observation c0836bd0-436d-4232-a7b5-5f1de43df576 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Position: The Most Expensive Part of an LLM should be its Training Data Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.420387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.420387Z digest=sha256:174922bbe1f99ae97cebc9cd80d2b35401f32a8aa48cb86f19de5012c4666f56

Observation 29a1ac42-8e9a-48f9-bc48-2966b1f50e70 · outbound

This paper cites Phi-4 Technical Report.

Position: The Most Expensive Part of an LLM should be its Training Data Phi-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.424961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.424961Z digest=sha256:4e4ecf9722d0d84f632a4b8dc04aa24bebde84ff00438799303b3138f5e22cb3

Observation 94a2aa78-8802-47da-b446-a0bfd1abc6e2 · outbound

This paper cites Alden newspapers v.

Position: The Most Expensive Part of an LLM should be its Training Data Alden newspapers v

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-16T12:34:53.976485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.429188Z digest=sha256:7545c7ed168550c1123f57264cfd5c018fbabdd331cf8fa006186632c181c716

Observation 60d5f8ad-9b5f-4da0-bfcd-50cdfb622b42 · outbound

This paper cites M., and Weber, G.

Position: The Most Expensive Part of an LLM should be its Training Data M., and Weber, G

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.433029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.433029Z digest=sha256:aafa43118788b13515a299634ec796c47586807ff0a43b62081f752ea0b2e2a9

Observation 7adfd76b-b4c3-45f7-b7cc-788c976aecab · outbound

This paper cites Authors guild v.

Position: The Most Expensive Part of an LLM should be its Training Data Authors guild v

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.206262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.437311Z digest=sha256:5a1e30e01b55226bb45273cc817c9ba2b9b5269a3c901ea288f6cf5dbdbd5d18

Observation 72bb5d94-82ea-4e6d-abb9-10b563348510 · outbound

This paper cites What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions.

Position: The Most Expensive Part of an LLM should be its Training Data What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.441434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.441434Z digest=sha256:0f0e3a3025dd9317f462e8fbbd29fd8e6a64461078ce81b0c4e4457b0216c2ac

Observation 558c1ceb-306d-4caa-86a9-1ef6e5f74c7d · outbound

This paper cites Common crawl dataset.

Position: The Most Expensive Part of an LLM should be its Training Data Common crawl dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.196771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.445464Z digest=sha256:4496dd2c55fe3addff47199705554df44e39ca92cee9edbbfed0250a3d796d6f

Observation 3518d52d-387a-4a8d-9541-f23ebcd8f2f9 · outbound

This paper cites Concord music group v.

Position: The Most Expensive Part of an LLM should be its Training Data Concord music group v

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.187530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.449194Z digest=sha256:fd4488cd0a802d4c9263b5fedfde2c0e5d75ffafb78049c35bfb509175493ff4

Observation 252eac39-7388-4717-be9e-6a70a8dc6d8e · outbound

This paper cites The rising costs of training frontier AI models.

Position: The Most Expensive Part of an LLM should be its Training Data The rising costs of training frontier AI models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.452722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.452722Z digest=sha256:1f41903b0b5e3fd354e7c18bba480a50ce755b6394d9f8f9f58a8c6a2d110955

Observation cbefb914-e2ca-4d11-93c0-b8d709d75485 · outbound

This paper cites Ai is a lot of work.

Position: The Most Expensive Part of an LLM should be its Training Data Ai is a lot of work

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.177432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.457066Z digest=sha256:8a7de494d0f1589d220bf5333c7589a42a8e9f917cba43c19f993c3adc3d9120

Observation 0b9777a9-53d8-4c17-988e-4a47343008d5 · outbound

This paper cites Encyclopedia britannica 15th edition.

Position: The Most Expensive Part of an LLM should be its Training Data Encyclopedia britannica 15th edition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.167706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.460931Z digest=sha256:7f6778d8790a3d333258b2694c964ce28a9befc44222234a9b7c95189e4bd0a7

Observation 9ac7fb88-d47c-487b-9be8-f0936f3ab916 · outbound

This paper cites Data on notable ai models, 2024.

Position: The Most Expensive Part of an LLM should be its Training Data Data on notable ai models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.157863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.464492Z digest=sha256:2b3bd130496b5d6a821f64bc625359294853eeb41cc6cd6363ad0f88b9581d5e

Observation c957d152-6d94-4247-b56b-21a7ecc89e37 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

Position: The Most Expensive Part of an LLM should be its Training Data DataComp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.468171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.468171Z digest=sha256:3292f4f7f9c3487639a2ff6ab57e3fbc0662855aa6b9815538e228beb240caa7

Observation 01f09962-6c0c-46ab-a7ef-e7905f4ce083 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Position: The Most Expensive Part of an LLM should be its Training Data Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.471844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.471844Z digest=sha256:4cf21e3c37d8bb9308b07891a35266d3452f06cd84403d86cc5a88d3a129e6e9

Observation cba7a639-6920-48ad-a543-dbde1f5b53a7 · outbound

This paper cites Data Shapley: Equitable Valuation of Data for Machine Learning.

Position: The Most Expensive Part of an LLM should be its Training Data Data Shapley: Equitable Valuation of Data for Machine Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.475484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.475484Z digest=sha256:b25c011334d1732e63291ef15c02cc1c0b2401c77774d79e3d8e98abf05f456b

Observation 03d72fee-497e-469b-a1cd-5e46266b1faa · outbound

This paper cites Evaluation of Similarity-based Explanations.

Position: The Most Expensive Part of an LLM should be its Training Data Evaluation of Similarity-based Explanations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.479173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.479173Z digest=sha256:1e0ab374695458cfb66cab36bcb87c6d2fc5cb3fd2e195d1addaed2d09687944

Observation 95e6842e-d01d-40fb-95c2-875558897655 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.482888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.482888Z digest=sha256:6995ca851d9cfdc4cb4ec52fea83d95964f56b5446edcdeeb02205de5d43ad1f

Observation e0e7818b-ac76-4f46-a3a7-5c125e939bf3 · outbound

This paper cites Statistics on wages.

Position: The Most Expensive Part of an LLM should be its Training Data Statistics on wages

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.147863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.487229Z digest=sha256:a195ba5c15f0e1a905aa11aa00a72ba535d29730f8d80289aef9dffacf7b7fee

Observation 6bad627c-5881-4537-af10-2065619543c6 · outbound

This paper cites Scaling Laws for Neural Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Scaling Laws for Neural Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.490580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.490580Z digest=sha256:f7e8342a6b3ffca9852a88f4646523336b0762b1874a7228e52c258c2497d769

Observation b779aefe-5e3e-4bf2-9355-d243a13effb7 · outbound

This paper cites Understanding Black-box Predictions via Influence Functions.

Position: The Most Expensive Part of an LLM should be its Training Data Understanding Black-box Predictions via Influence Functions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.494514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.494514Z digest=sha256:da8cd0d3019a1c5150f0510afdb97fa1e9bd4f27e6854779f77e031e5e5d37b8

Observation dc20d4d5-255e-42cc-80ea-f82e155d7c85 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Position: The Most Expensive Part of an LLM should be its Training Data OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.498197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.498197Z digest=sha256:824155f2459fea5703454544565883c9a7dcbde8dbdeef0d7786b8feaf98e29d

Observation 2d526e30-7940-40ee-9d3b-5bf491476136 · outbound

This paper cites Releasing Common Corpus: the largest public domain dataset for training LLMs.

Position: The Most Expensive Part of an LLM should be its Training Data Releasing Common Corpus: the largest public domain dataset for training LLMs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.138207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.501877Z digest=sha256:05119a48da7d5cb641f4eca143a25eaf4055358afe0530c969f7f1acdf6bae90

Observation b173a557-7554-43ba-9c09-4b38c6b0d91f · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Position: The Most Expensive Part of an LLM should be its Training Data DataComp-LM: In search of the next generation of training sets for language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.505422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.505422Z digest=sha256:d25071ab0c9bd7873e2db57708446f5314f0cd479fccc7506aaa0a2f3d735f81

Observation 589d0d14-d539-4794-a32c-62452bcf0ab2 · outbound

This paper cites Consent in Crisis: The Rapid Decline of the AI Data Commons.

Position: The Most Expensive Part of an LLM should be its Training Data Consent in Crisis: The Rapid Decline of the AI Data Commons

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.509363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.509363Z digest=sha256:7f8a90ca7b8234b918b7c08d3a3af89dc011964309ac346bc92c19d69ff098ed

Observation b9175506-313c-4a0e-88e3-645e166e0898 · outbound

This paper cites SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore.

Position: The Most Expensive Part of an LLM should be its Training Data SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.513176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.513176Z digest=sha256:1abadc67ad438e2e59ab34445d6ddd00a91775bd1f174e2c50fd4287a5087711

Observation 50021800-f7b3-4840-8bfa-76e96383ddf3 · outbound

This paper cites AI models that cost \ 1 billion to train are underway, \ 100 billion models coming — largest current models take “only” \ 100 million to train: Anthropic CEO.

Position: The Most Expensive Part of an LLM should be its Training Data AI models that cost \ 1 billion to train are underway, \ 100 billion models coming — largest current models take “only” \ 100 million to train: Anthropic CEO

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.128186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.516963Z digest=sha256:ae2b368ef21d4083e876d9299887e853cfc1885daeacc20f902bfd3a3d2005ce

Observation 34d09960-edb5-43cd-87e2-dfe3162eaac7 · outbound

This paper cites Scaling Data-Constrained Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Scaling Data-Constrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.520465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.520465Z digest=sha256:74305d09c0619f7c162a342bc2d374951a5869a94ac0c90e27c1e9fc0305c651

Observation 51e8f85b-56e4-42ca-806b-f3d77227f000 · outbound

This paper cites New york times v.

Position: The Most Expensive Part of an LLM should be its Training Data New york times v

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.117493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.524159Z digest=sha256:227d174b7fd951a039c29dddd858bd74b7b17906f88394d69b1214895bcf3cc6

Observation 42b77bc3-054d-4116-bc9e-b719feca545f · outbound

This paper cites TRAK: Attributing Model Behavior at Scale.

Position: The Most Expensive Part of an LLM should be its Training Data TRAK: Attributing Model Behavior at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.527627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.527627Z digest=sha256:e04183af042167460ce7e0e8c2febc1488026592e7bf65a48518b8bc191d6c34

Observation 1c16cddf-21d2-4a9b-abad-7c7caead9e09 · outbound

This paper cites an unresolved cited work.

Position: The Most Expensive Part of an LLM should be its Training Data Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:34:54.107235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.531585Z digest=sha256:34fe1883e891d840e44fe066e2a86cf0ee88b6b52c05ce441109b1a75f6e8edb

Observation 4e90d540-21fa-4d3f-a022-c765ce57d782 · outbound

This paper cites Multiple ai companies bypassing web standard to scrape publisher sites, licensing firm says.

Position: The Most Expensive Part of an LLM should be its Training Data Multiple ai companies bypassing web standard to scrape publisher sites, licensing firm says

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.096999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.535354Z digest=sha256:05749eb93167d0b0225a847b87d708f955043a83838d01bdd19a7752183058a0

Observation 6bc0e423-6cf3-4234-9b54-9de913a1b9e9 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Position: The Most Expensive Part of an LLM should be its Training Data The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.538809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.538809Z digest=sha256:4b21ded1a731e62f7254d7081cab3ff3fa6d8a3b079aae87c4f67a0be8ee8c21

Observation 117ed8c8-2b11-4ca1-9282-57fa9ad774a7 · outbound

This paper cites Estimating Training Data Influence by Tracing Gradient Descent.

Position: The Most Expensive Part of an LLM should be its Training Data Estimating Training Data Influence by Tracing Gradient Descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.542296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.542296Z digest=sha256:670dbee7ad386b6ee28cbb271b56a2791d7d137df20f272419312f2eb1b74e95

Observation 8f50bfd0-16cd-479b-85d2-b1164801cce4 · outbound

This paper cites Reddit and openai build partnership.

Position: The Most Expensive Part of an LLM should be its Training Data Reddit and openai build partnership

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.086656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.546327Z digest=sha256:5fc4a62afdbfc825f56b5de800c8b9a9f6e06589d5bbd9aca47754b7ede94c97

Observation a27c7275-92b1-4929-a0ce-de699e929c1d · outbound

This paper cites Thomson reuters' adjusted eps beats expectations, ai boosts results.

Position: The Most Expensive Part of an LLM should be its Training Data Thomson reuters' adjusted eps beats expectations, ai boosts results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.075904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.549780Z digest=sha256:36a9c4dbd14e22b38bf3a3069205b156ce43017b1a941ad21c8313f5a00fea86

Observation 22b3cdde-bdef-4cd2-b62d-524a157ec2fd · outbound

This paper cites Shutterstock expands partnership with openai, signs new six-year agreement to provide high-quality training data.

Position: The Most Expensive Part of an LLM should be its Training Data Shutterstock expands partnership with openai, signs new six-year agreement to provide high-quality training data

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.064402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.553071Z digest=sha256:e6f81e7e7d2cbfbece54f05ca2262ab9daa8fc54ed5cb2ab893d3369ee76ed4e

Observation 56e88207-67cb-4cda-b630-7c7928ad4f1e · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Position: The Most Expensive Part of an LLM should be its Training Data Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.556356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.556356Z digest=sha256:86058a0151d426a8d62e7c48b74a9e1e4f6921606bb48259172bf2aa39b476af

Observation 5c5e27f6-786f-482f-8476-120d83c243b4 · outbound

This paper cites The atlantic announces product and content partnership with openai.

Position: The Most Expensive Part of an LLM should be its Training Data The atlantic announces product and content partnership with openai

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.052727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.560435Z digest=sha256:df69d90796ec36c2bbc0c927b38d98bccb0a31b8e380022e1e2aa39fe77d2743

Observation 02da8847-5bea-46ec-90e4-ab9716fc5df4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.563650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.563650Z digest=sha256:f48776a52f14971c92bc129f7a3796642f67c1b311b4d46fded15344e94e1457

Observation dde8342c-841c-44e7-9816-8d01bed89acc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Position: The Most Expensive Part of an LLM should be its Training Data Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.567083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.567083Z digest=sha256:938ce4e4feabde88e1fe71883747dce803448ac65aa62f6d4b9418bbe11f89d7

Observation cbf0bcf5-7040-43f4-ae36-73e83a968d80 · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Position: The Most Expensive Part of an LLM should be its Training Data Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.570304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.570304Z digest=sha256:6e526291f1a63434f7d716f770766991f411d9c92d3138e4403e00e6f59fa379

Observation d74ac334-eff0-41e9-a858-6d7d0eb6f15b · outbound

This paper cites Vox media and openai form strategic content and product partnership.

Position: The Most Expensive Part of an LLM should be its Training Data Vox media and openai form strategic content and product partnership

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.042052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.573730Z digest=sha256:555b47536ff594ff44b5b290d35638e38ac2c223cdb07e30aa1c21e93d80ee62

Observation 3b58b3df-6705-40f0-bd7c-321f26fd55e5 · outbound

This paper cites Html standard.

Position: The Most Expensive Part of an LLM should be its Training Data Html standard

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.030794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.576828Z digest=sha256:21f84da7bb6340b5539f9185ff87c40b66a9e2b892faae617e7fb722daca0913

Observation 78226ba2-210d-4061-9217-5374472983b3 · outbound

This paper cites Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus.

Position: The Most Expensive Part of an LLM should be its Training Data Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.580062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.580062Z digest=sha256:90ba1af1f0e0a9574626a667b8b833fcccd3e40ad6ac13e3817e94ff8fc3cb3d

Observation d9761ef1-9d90-41fb-b3ac-18d1a352da19 · outbound

This paper cites Breaking down emerging segments in the ai content licensing landscape.

Position: The Most Expensive Part of an LLM should be its Training Data Breaking down emerging segments in the ai content licensing landscape

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.019864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.583589Z digest=sha256:c82aabd64fc3d5bf95dbe06cbcac2fce49e04498e247d5edaa2c5fef21d69e89

Observation e321c3e6-f2cd-44a7-b3d7-749b33eb06b1 · outbound

This paper cites Articulating value from data.

Position: The Most Expensive Part of an LLM should be its Training Data Articulating value from data

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.007464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.586790Z digest=sha256:ac84723692f6a548d4d4cc4c6e8f8e675b2f5e27b971d8b6d4ca41ebbd5ec596

Observation d301cb7b-75c5-43ff-bab2-da3c74223efc · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Position: The Most Expensive Part of an LLM should be its Training Data WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.589828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.589828Z digest=sha256:58c2694e68cb90b533170ca89a14bd622460872e45a9210854f35677c3d1aed0

Pith citing papers

Observation 179fac02-40ff-4f0c-a983-85f16f6422af · inbound

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text cites this paper.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text Position: The Most Expensive Part of an LLM should be its Training Data

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.273599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:44.273599Z digest=sha256:82f06f40b6085cbfa3e370c2461ae9bffa6bd75c6bebfc6bd5a0313970f11dd6

Observation 7895a035-1447-469b-a995-8b83e5474abe · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:25:55.503750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T14:21:53.325788Z digest=sha256:c0dd536c3ff49f3e25819c81deabac18a8068044abee1a3d78f8efb5bcde4d6f

Observation de189bdf-7b49-496d-8753-7ab89b8483e5 · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T12:17:47.123519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:17:47.123519Z digest=sha256:8ffe57e4fe9135192b5a517a5877d90a1679ca42c3288624c6fdcaae1e62b903

Observation f2438a06-593e-4df0-9cc9-bd8ff63152eb · inbound

On the Fragility of Data Attribution When Learning Is Distributed cites this paper.

On the Fragility of Data Attribution When Learning Is Distributed Position: The Most Expensive Part of an LLM should be its Training Data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:17:39.452630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T15:16:32.125215Z digest=sha256:071417d492f6bcf3c0f9a74571a2fa0f8538b953fe8d4eb7fcc7e5fc3b06c68e

Observation fc6c92ef-bb74-4ea7-8367-1d7781beffa3 · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:46.611909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T07:52:24.053501Z digest=sha256:b122528f1be3da27f35c1fbcde7462a17901efd1f08541d1479965edaa9d07ab