Pith. sign in

Paper Citation Record · LEDGER

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 7 inbound Pith citation observations for arXiv:2502.12120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12120 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T02:47:37.492619Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:12.517211Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact34
  • verified fuzzy8
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e36ddbe5-0711-4b57-bb6e-f18d93399fa1 · outbound

This paper cites Exploring The Landscape of Distributional Robustness for Question Answering Models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Exploring The Landscape of Distributional Robustness for Question Answering Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.150102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:7003dd6c858667c73396bceb72bcead17f8d8d558864a04d557e81cfe4302476

Observation c7cea62b-03fc-4c28-b0ed-fea36cc1fd3b · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.153689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:8c8f83c63be5772fad2778270c8704293663f54eee6a4fcbabb311cb778ff564

Observation d06baeb1-d551-4cc1-b11b-31006ccf408e · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.142978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:ec07ce140779c5475011cc65c5de0dbf09ffc38dc071ffda1fdbb1685fc6b6e5

Observation 6aace21d-bf84-4ef1-8893-bdbb6402fbe7 · outbound

This paper cites Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.098836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:9f1392a0358eacca0c9563468294cbba3b265715aaa12b97bb085806fb8c8a01

Observation e2ddf806-0f71-4c3b-aaff-ea600b6737c1 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.125494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:6f6a6414d209c84258c9a2fa3b1fae00256a234d0326726a2ee171d60c5b64c3

Observation 301b63cf-4ca6-43db-9fd3-45446de0fef7 · outbound

This paper cites Loss-to-Loss Prediction: Scaling Laws for All Datasets.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Loss-to-Loss Prediction: Scaling Laws for All Datasets

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.118254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:9326817c9be2282c8edbf4bd1f7c9d9dd6c939e556518e774159a34898699580

Observation fbdb04fe-61c0-4e19-8fdf-1792bcc767d6 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.121737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:79351f9b7f0c4a442c1ef5d3eb0524c2eb2e968cd5478eac6257942cbb95f930

Observation 363b004a-1ffb-4daf-bc06-6e9c0cfe343d · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.128976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:235f677c80ee86f6990f49003abb396656c3d1049049e01d09f389f249860977

Observation 084b1a00-7f2d-49f3-b08d-4b2d743be9d5 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.135687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:0a6d3e04c9a62be232245c10746c767b144a81972a6b5ad31508735a772f8e31

Observation 1b5153ba-a9e3-4f99-8924-8726896eb432 · outbound

This paper cites Understanding Emergent Abilities of Language Models from the Loss Perspective.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:52:27.109203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:775e6d6fddddc1e71d344d598038b6c56cf92faf32d597b2911b98756eeebce4

Observation 364ec55a-d6e0-462f-9dfc-01d21d19c0c5 · outbound

This paper cites Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP).

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.114874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:472a30b6131b29d1e9cb5f2d30c0938d9e260c89dc40afb26b3d803ebcaff316

Observation a75009f9-5d6a-4cf2-b1ff-5b5c1803bba2 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Language models scale reliably with over-training and on downstream tasks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.106052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:ef733fc053aa82d3f2723065a1167dc42d19277dace10d9d01cba43ac1b9b5e2

Observation 4cf67fdd-c60e-4d83-a24f-b30aa6f3fc25 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.098574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:a622f42f35c295e78057294b77dff090356d8a53016ba1a38d9f0b34222ed9e1

Observation cda4c4ca-532c-4873-b0c1-0a612c7ca51d · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:52:27.102560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:e8830feb0ef6d7a8e80b4215eb4e0d6158282c03f9e23483142ca1af85e34ca3

Observation 81d5c581-4990-43f7-abd9-756b2bcf6f4f · outbound

This paper cites S., Kozareva, Z., and Roemmele, M.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws S., Kozareva, Z., and Roemmele, M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.108621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:33dccc2cc5eac7a88dcf839c248958b01accac057b17eaa6d2894b7cf3d17812

Observation d1538cda-641c-40ad-a48e-900c7d8de0d8 · outbound

This paper cites The Llama 3 Herd of Models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws The Llama 3 Herd of Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.111926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:0143ab698a5a8a0d726303c3bb6317d9bd1e189b4d5fd6c86e2b6b8ed237c096

Observation 25ad0d65-43b7-4e46-b480-ad677be85622 · outbound

This paper cites OLM o: Accelerating the science of language models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws OLM o: Accelerating the science of language models

Reference 17

Resolution
verified exact
doi, observed 2026-05-23T02:52:26.337254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:ac4fd60b87eec3dc443eb3e5634c4fff62568747ec768b0e514e084d1f1fcf30

Observation fcc93f31-4bbd-486f-bd71-73b88014a7ff · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.139211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:91f9a93784c52a92a156c9b703d8374c666c3020f062309e786c7129d6429b78

Observation c8dc67b7-acfa-4e7b-ac8c-aae2bacf76c8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Measuring Massive Multitask Language Understanding

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.084002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:57adc9e4781799d46abdd48582d6d9f5ed7e84f4d612358f911120af8a2e823e

Observation 58e3fb3e-fadf-4442-90c2-2ea0e79218b3 · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Deep Learning Scaling is Predictable, Empirically

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.146660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:684ea82de37c99d065aa98e02f5d0efcb0fc25cf5747a9812e406715d69b7821

Observation 6b70340b-707e-4ee4-955a-5fa53d4ac107 · outbound

This paper cites Training Compute-Optimal Large Language Models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Training Compute-Optimal Large Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.077172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:af1e5f6b03d80ea6bbf1f2c928799afa3afaa0ab306aa7a7f5931534517e8e97

Observation 629d0798-9ebd-46e2-9a83-6010ee3ac1ed · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.094710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:2567ec49a1bcc1ae93123b0fa9a939ecd3fe30517854ef1e4270abf2f25bc2aa

Observation 078c30a2-9096-4c4c-b0c7-95988dbad9d3 · outbound

This paper cites Scaling laws for downstream task performance of large language models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Scaling laws for downstream task performance of large language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.087842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:24c5e092a1d7cfcc209699f128e0789da06e1dfedd8572c5ab2a5bf3d5584b10

Observation fbae82ac-aae8-499b-8a10-f6fdc2595eda · outbound

This paper cites Scaling Laws for Neural Language Models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Scaling Laws for Neural Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.090961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:4deee1984337fd4c68abbc610cba2be01f33a3cec891a91289164f6ce3e2a311

Observation d936a0bf-0910-4dc8-a2fe-0c7c5e64c97a · outbound

This paper cites nanogpt.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws nanogpt

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.115411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:fdc1ed2be792f64734eb83665eb3ed3433d124e7e6c4a379e27c939293eb64df

Observation c358675d-1e29-4408-b9ce-8ada8219276c · outbound

This paper cites Adam: A Method for Stochastic Optimization.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Adam: A Method for Stochastic Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.132399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:aa553c5d442143314d0299b3a482a0fceb9f90e32145ac43771a39c3021d2bb9

Observation 8d3247ff-e86f-4033-97c3-2edf744f7ba8 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.073705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:ba91451776d176d770ac30b0562e0bb9d941718cc79ed36aa545d13eba662c4e

Observation 2df8dae7-c30f-4812-b7d4-9205bce9803e · outbound

This paper cites Decoupled Weight Decay Regularization.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Decoupled Weight Decay Regularization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.059391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:0592fef90f1a8beb98af4b9bd5671445b24c6e1fd1553a502d60ba4743aa3699

Observation 0b1e145f-0b0c-4cbf-938e-5aa28133c018 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Quantifying Variance in Evaluation Benchmarks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.051526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:f9764d6d41c5924662cb092eb97ad6bbf0559e2724f3de9abff86202b374c6d6

Observation eb592648-1229-459d-983a-d2cd8215eae8 · outbound

This paper cites Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.055543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:e65f8c3f96e89bd6c4b6d387c7ce1224584c3d160b01512b025900b1806ecf7c

Observation eb1ea2a2-7574-4ae8-80ef-4b181286e205 · outbound

This paper cites In Search of Forgotten Domain Generalization.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws In Search of Forgotten Domain Generalization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.063822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:b9b7ab0dc1327f9bbc600e897e381f9318e00f3ee7a9fda038ab0725a7d0b887

Observation d0ac1c9a-2327-4dd7-88ea-771f6bac3209 · outbound

This paper cites Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.067063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:aab204fa287c2574111bc6d19981fb42425e0f453fce9fc7ed7eee51602924b6

Observation fe364a94-f893-4468-ab85-1d61454841d2 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.008419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:432767a0429f7e1bf54b002a6690449afc41ddc80ec3b7cd7b0a5c6de4c13d55

Observation 16246f4b-6745-49a5-8b97-5e93b98c865f · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.040547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:a5f96f59c5543401200436d1b2ebabf049711a1c26e29f87fff59fae67ca18b5

Observation 67dd0421-289c-416f-89ea-365cd9570f7d · outbound

This paper cites Resolving Discrepancies in Compute-Optimal Scaling of Language Models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.048190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:fd6699e95374ad54062586f8b9beb191841ab83f568f84f80d838d8c674e2557

Observation 85eaf020-83ab-46d9-9930-5a1b8df51300 · outbound

This paper cites Language models are unsupervised multitask learners.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Language models are unsupervised multitask learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.118336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:e915b07fdcaf9373d2f1d74fe4e275549068eaaaff3e14b8b3c86ae637cc2e17

Observation 18ec24e3-6e4c-4537-9c4c-9f0d19ccef1e · outbound

This paper cites On Linear Identifiability of Learned Representations.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws On Linear Identifiability of Learned Representations

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.025945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:3118f955db8708cdc29316ba193afb37e272ca6508d2613ef1e3e97cacaa5363

Observation 95f63f9d-15ea-4528-9fec-3598a2e9c887 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.029445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:35e2b1b63d8d98b304ad99c2e19487cae6e9fc5964662ffc98eeb05f5de4d350

Observation db2ab083-36a2-4b7f-a885-cb06416fa40b · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws SocialIQA: Commonsense Reasoning about Social Interactions

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.036979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:96cfbf95f4c53d5b7b73ae2e57807e97c3a1e672dd633fb36b7ef58791c931e1

Observation 24b30f4c-6ab0-4dd5-8bb9-fb22db13d759 · outbound

This paper cites On the Inductive Bias of Stacking Towards Improving Reasoning.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws On the Inductive Bias of Stacking Towards Improving Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.044287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:62766f6102e957a9ed0b45828aefb90438593942bef080dcc34df6657776709b

Observation 7068dbf6-0708-4ce8-8de9-7d697bdceed7 · outbound

This paper cites Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.070448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:51934514e6ab587a0fb423e3ec4410338331e4fa663d7f85277be722b6ac77ba

Observation a9c4ac4a-f922-4458-8966-c4436408edf6 · outbound

This paper cites SlimPajama-DC: Understanding Data Combinations for LLM Training.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws SlimPajama-DC: Understanding Data Combinations for LLM Training

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:52:27.015453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:8583d184be792c9badf5461755bcef08ad2874b4917cbee342f16e0714f16df7

Observation 8d15554c-d71a-4ef7-8c1c-fdf1cccd2c44 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.022377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:ef4ff52befea2c54243bfb64a04705ea9248fb9df683eb744f2fcaf072f94d32

Observation 66c487a5-5417-4570-8154-86fe5729766b · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.018724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:a91e99908c0f2586d617fe3c50752d46e55e6247d4c6504898bef053bd86e7a9

Observation aff85cf0-1cc3-48ad-9117-65df32209e0c · outbound

This paper cites Measuring Robustness to Natural Distribution Shifts in Image Classification.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Measuring Robustness to Natural Distribution Shifts in Image Classification

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.011832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:ae99c77e4842e8ded4e4174128a6d9dd9d755bd65e0d2c20099333d6250f50f6

Observation 91de4426-7f01-48ba-ad49-6892bf231d93 · outbound

This paper cites Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.033438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:f42616f2b4bf4cebbfb5f9a7bc22ef83159332ad938b9a8412107a96e5cc1710

Observation 9387ca80-436e-48bc-9825-6b3459f90f8c · outbound

This paper cites Y., Haziza, D., Wehrstedt, L., Copet, J., Teytaud, O., and Lopez-Paz, D.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Y., Haziza, D., Wehrstedt, L., Copet, J., Teytaud, O., and Lopez-Paz, D

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.112060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:815da42a93d02d6f6553e2ccbb022137713df8847fc4c2904a673efed2009f66

Observation 5e3d9051-8da0-4201-b34a-7055c581d87e · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St´ efan J.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St´ efan J

Reference 48

Resolution
verified exact
doi, observed 2026-05-23T02:52:26.333993Z

Source-reported events for the cited work

correction dated 2020-03-04. Source: crossref record 10.1038/s41592-020-0772-5->10.1038/s41592-019-0686-2:correction, observed 2026-07-11T03:08:19.430942+00:00. This notice travels one citation hop only.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:db91c25f065638ab3c6de4f23fa3fdd25e9a06b3f16a6dbdd44707f28ec0c7a8

Observation cbfe1370-971c-4fea-b12c-716dd4ef8d09 · outbound

This paper cites and Komatsuzaki, A.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws and Komatsuzaki, A

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.105234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:70738eb0d9c98360d5ea61d148e0b5d31952f37a03d31c27a0780c995745ee91

Observation 73e00f2a-6b80-4a37-9734-c94fd4c1993a · outbound

This paper cites Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.080480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:57c2461703d34f1ec3f668e13a33520374bd915814f981fd3111f1e0ed4de883

Observation ad57ad34-54a2-4996-b754-2a1c99fdde05 · outbound

This paper cites Pretraining frequency predicts compositional generalization of CLIP on real-world tasks.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Pretraining frequency predicts compositional generalization of CLIP on real-world tasks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.102181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:b4485b3457768c4e3fb2d3699f6cdf59b276f0f0aafe15c07950f59a5ccb0cc8

Observation 749b280b-6aec-42a7-afbe-1319b694d4a1 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:27.004648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:b359a9d9baa0d4df4d271e89c9dc6f5bc7ba6e0c87fa33727f8702f159e75714

Observation 7ade2323-c95b-4630-ae9f-567be4c1e44d · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:27.001197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:48ca20e9507ff662c50dbf2d430bf07ba67c018557243a63f674d6e2f18b5e4f

Observation 7ae70e27-f2c3-4dd1-8a0b-762fbf1f3215 · outbound

This paper cites write newline.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws write newline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:27:39.121286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:27aeacd5b02886b1e70f1034cabb7507452990585cdbe5b7cb24df042a554660

Pith citing papers

Observation 89f4ec0a-15e4-410c-9400-07e39d7e2ef8 · inbound

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection cites this paper.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.154116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.154116Z digest=sha256:ba555c0e30f029d70beac1c106b39896181cbf1527dc3f4205da29ef416837c9

Observation 7fcf85e1-298e-4c01-8cb1-422e54e41ba0 · inbound

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining cites this paper.

ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:48.398461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:09:48.398461Z digest=sha256:d53ec9233b61660769ccdd6da72290b0495f2dee25e7f1cd1a4d5e139fe72809

Observation 1846d484-bc65-4f03-a5e3-cebc18861ac0 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:11.706808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:11.706808Z digest=sha256:b5d5189eb630b69280a29718479110bc5070b1019d352f08d825e064b039e79f

Observation 3c516a2c-cfca-4dda-b51e-369d1236d499 · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:34.355258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:34.355258Z digest=sha256:43c15a579fae15e5c31485008677334aa69a2980ae10654fd49c8b1bf7fdcf9f

Observation 9833058b-abf7-4bad-b952-1a77eea611cc · inbound

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks cites this paper.

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:31:48.273075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:31:48.273075Z digest=sha256:89fd671d9ddb73b72a4269aaeff2aa330ed99ef2852c9582713e09a3b0636c4e

Observation e59d2c6f-c038-46e0-af4a-c47f3d7bce59 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 103

Resolution
metadata mismatch
local_arxiv, observed 2026-08-01T03:08:35.287832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-01T03:02:07.540907Z digest=sha256:eb8c10084a4c3fdff4fc1031b23aff020b0ba9b64565a4eaa1fdc74eebe93da8

Observation b4333c37-c8e1-46c5-88ff-8055cf3b5f51 · inbound

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure cites this paper.

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:12.517211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:12.517211Z digest=sha256:633bde19c7590b9fde7a8ea4a29914eea7252d6935255f31b236693db7dfa2eb