Pith. sign in

Paper Citation Record · LEDGER

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling

As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2501.13779.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13779 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:37:28.031565Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:30:13.859525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:30:14.503066Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57aa3840-cf97-490f-8bf0-3165d9225760 · outbound

This paper cites Data curation via joint example selection further accelerates multimodal learning.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Data curation via joint example selection further accelerates multimodal learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.987061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.987061Z digest=sha256:b42803251d74d2d725326910bcee5467fc9768209c860fe652bf776c157a1695

Observation 76e21ff7-a526-4546-8e6c-a79829391353 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.991160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.991160Z digest=sha256:d4d34be22b832123a847596d8e02b4360281dce9eaa8288816132ed45c208660

Observation 8f0e6cfd-f91a-4dca-ba43-7acdf99c0232 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Training Compute-Optimal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.995473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.995473Z digest=sha256:ef20dd933566b3cd1afaa163c74d21f51fc0ecd43e338dbce3b54ff5cb2f9e57

Observation cf1cbce5-c727-467e-a0c7-3f0cbf45055c · outbound

This paper cites A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.007556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.007556Z digest=sha256:6e8119bb40a80589b79d965d42827d779dab96604a270c831831937ce5efa881

Observation df4be504-2cc3-45a3-8968-7a680def5720 · outbound

This paper cites Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.015364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.015364Z digest=sha256:d3369f28e688e14571067dcc8bb85789f5b516b9e2c6037fcde91e0badde540c

Observation 493bc0ec-9dd7-4b54-aaff-a0627e07ee4e · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.019203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.019203Z digest=sha256:8251142948779192b5199b2806f5142d2b9e502e6a90718fc188d90322be1523

Observation fa7f149c-1842-4da7-b1d9-d485e5ff0547 · outbound

This paper cites Topological Data Analysis Applications in Natural Language Processing: A Survey.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Topological Data Analysis Applications in Natural Language Processing: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.023924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.023924Z digest=sha256:64eb115ba2bb14b3e12e2ab035d4e7220b9f0a02a5f113bf638f8844b3e34b93

Observation 2a9d285f-a993-48d1-99f8-40abf99fe9a7 · outbound

This paper cites Distilling System 2 into System 1.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Distilling System 2 into System 1

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.031565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.031565Z digest=sha256:7e720b0f83311119fe36a03785ed1c360bff0684921b78879f43dba747c78254

Observation 1d68e051-9436-4202-af41-b71856f84b7b · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.027916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.027916Z digest=sha256:f173f96b89aad193be8d6da65ee883665116db14095642a64fd6b58e490842d1

Observation e72f662c-45d7-4076-836a-49b06127408c · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Quantifying Memorization Across Neural Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.978597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.978597Z digest=sha256:b7013d94c39b3043e37125fca3a4d2da96a1aa713500f551834cd71cf3561606

Observation 48e7baaa-e522-43e2-b2cc-0783039d3c92 · outbound

This paper cites Deduplicating Training Data Makes Language Models Better.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Deduplicating Training Data Makes Language Models Better

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.003718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.003718Z digest=sha256:b1fa93e4eaf495af805b3b5f41aa17ccc63e93f8d64164ebbb0dea25d93a9b6e

Observation 27396786-c410-4a16-864d-89ac43ba7668 · outbound

This paper cites Scaling Laws for Neural Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Scaling Laws for Neural Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.999507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.999507Z digest=sha256:5eab661f0841e891eb69dfc4be0b58ed1616949d8f02f54cf1344ef697d82aca

Observation de5fd0ff-ecc9-49c5-a025-124ef5306c59 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:28.011614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:28.011614Z digest=sha256:5bf65a96ca0679c83fb199f0738ec3cf39a42b6b7ce5195138f2d83a86b82b92

Observation ab030787-fe67-4c77-9d86-88cae4a6308c · outbound

This paper cites The Llama 3 Herd of Models.

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:27.983122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:37:27.983122Z digest=sha256:c316b7b2b87bdbf1412836c77997ac0ba96adb34099d8620f2f9173ec716191d

Pith citing papers

Observation 4fc270a1-8595-4ca1-8907-7105aa6dcf02 · inbound

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training cites this paper.

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:30:14.508240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T10:30:13.859525Z digest=sha256:2725c8722936668a7066fb1b9976fd61ff5217e03ae82720d131188ea258665e