Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:37:28.031565Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2501.13779.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:37:28.031565Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:30:13.859525Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T10:30:14.503066Z
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57aa3840-cf97-490f-8bf0-3165d9225760 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Data curation via joint example selection further accelerates multimodal learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e21ff7-a526-4546-8e6c-a79829391353 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f0e6cfd-f91a-4dca-ba43-7acdf99c0232 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Training Compute-Optimal Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf1cbce5-c727-467e-a0c7-3f0cbf45055c · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4be504-2cc3-45a3-8968-7a680def5720 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 493bc0ec-9dd7-4b54-aaff-a0627e07ee4e · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7f149c-1842-4da7-b1d9-d485e5ff0547 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Topological Data Analysis Applications in Natural Language Processing: A Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a9d285f-a993-48d1-99f8-40abf99fe9a7 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Distilling System 2 into System 1
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d68e051-9436-4202-af41-b71856f84b7b · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Will we run out of data? Limits of LLM scaling based on human-generated data
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72f662c-45d7-4076-836a-49b06127408c · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Quantifying Memorization Across Neural Language Models
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e7baaa-e522-43e2-b2cc-0783039d3c92 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Deduplicating Training Data Makes Language Models Better
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27396786-c410-4a16-864d-89ac43ba7668 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling Scaling Laws for Neural Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5fd0ff-ecc9-49c5-a025-124ef5306c59 · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab030787-fe67-4c77-9d86-88cae4a6308c · outbound
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling The Llama 3 Herd of Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc270a1-8595-4ca1-8907-7105aa6dcf02 · inbound
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.