Pith. sign in

Paper Citation Record · LEDGER

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2607.14306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14306 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:45:19.127391Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa6f2505-82ca-46f3-a3ac-67c21dfb0bc3 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.540645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.540645Z digest=sha256:6029fe3611618f644bef7ade55e24c622b335e2c8c6f69fb441add0cd869a597

Observation 150c4578-eaf8-4766-b1d3-ced5db3b2652 · outbound

This paper cites Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.695606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.695606Z digest=sha256:7f467cf1f2123fcd17ea4b47c9ae54475f1a7639cc040e8948f34944e8210da8

Observation 780bf20f-282f-46c5-85ef-4a661bf8fd7a · outbound

This paper cites In-context Learning and Induction Heads.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions In-context Learning and Induction Heads

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.749882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.749882Z digest=sha256:bdb8838bf4110c195481165a34ae6e6a46085a19645afe46863d73a1c813c913

Observation 0cf14fd0-76dc-4043-90c9-9f4180ce1986 · outbound

This paper cites Backtracking mathematical reasoning of language models to the pretraining data.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Backtracking mathematical reasoning of language models to the pretraining data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.821579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.821579Z digest=sha256:520e1497dde66493c6b351737cbef50bc7e3a4b9f63de8fcae3aece21f7fd481

Observation 46f28f8b-a6e8-48f7-bf91-c2189e23b6f3 · outbound

This paper cites Polypythias: Stability and outliers across fifty language model pre-training runs.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Polypythias: Stability and outliers across fifty language model pre-training runs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.971744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.971744Z digest=sha256:486be114d03bb5726ebc7448c3839494539d58497fd5c0ddeb957c2626c91c62

Observation 3676417d-e443-41f8-927c-9cc3d4a0c35f · outbound

This paper cites Jiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan, Yunhua Zhou, and Xipeng Qiu.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Jiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan, Yunhua Zhou, and Xipeng Qiu

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:19.127391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:19.127391Z digest=sha256:73c3e6d8fd295c0977aa024b10f42c534cba96fe57492464b238822c57465379

Observation 5838fa0b-fa07-42dd-8f97-28d6d0efeb9d · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Studying Large Language Model Generalization with Influence Functions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.619376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.619376Z digest=sha256:3583b21c1ec16cc0877b91fb09a8898e0afa94ce90e2f590cad62cb32c415aa7

Observation 13dc8bbc-f492-4229-9cb5-4ada2d9b8f04 · outbound

This paper cites Stealing Part of a Production Language Model.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Stealing Part of a Production Language Model

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.338994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.338994Z digest=sha256:3002b4f4fb09707dfb4303e72276bb95f84731dc12cfa6c218e324834fd85280

Observation 90fcf031-a76e-4a02-b970-b99ba6125a6a · outbound

This paper cites Language contamination helps explains the cross-lingual capabilities of english pretrained models.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Language contamination helps explains the cross-lingual capabilities of english pretrained models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.298168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.298168Z digest=sha256:612b05a1c5392716a5be1326c11691c5a8e3eb54312f83599223c87f5c24399a

Observation c3aec7f0-926f-49ce-8ecf-62598d29d059 · outbound

This paper cites Pretraining data statistics shape the phases of learning entity comparison in language models.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Pretraining data statistics shape the phases of learning entity comparison in language models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.393265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.393265Z digest=sha256:24a6a4f22aa8847a55c3f1e4183b35dc1a3ee5dc46217956009dd6204351e915

Observation 08e6c2ee-e8ef-488e-968a-145f1d8e271b · outbound

This paper cites Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:19.044382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:19.044382Z digest=sha256:7e4671b62a00644a0274bb3ba06e913f5eb24957e8d039b16d32edeb2141640d

Observation df60d6c8-5a36-4e5f-9d43-e4f19ada691b · outbound

This paper cites Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.458866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.458866Z digest=sha256:12fff23376590660fe7db3a8e6cecfc6c0b57f6e449df70bab45fb4b6f3529e2

Pith citing papers

No inbound Pith citation observations are available.