Pith. sign in

Paper Citation Record · LEDGER

Predicting the Emergence of Induction Heads in Language Model Pretraining

As of 19 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 4 inbound Pith citation observations for arXiv:2511.16893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.16893 v3

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:06:40.995944Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:46:08.533454Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:18:13.088871Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 742b60f9-0d32-401c-bc77-9c58a9f97b5b · outbound

This paper cites Language models grow less humanlike beyond phase transition.

Predicting the Emergence of Induction Heads in Language Model Pretraining Language models grow less humanlike beyond phase transition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:38.665206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:38.665206Z digest=sha256:9a79f159b5e9472071afc5058fe91ad4b81b1e2d0f2d9b33c649ef519f0fa695

Observation 077ff1f1-71c1-42bd-80a0-5134499e5ee1 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Predicting the Emergence of Induction Heads in Language Model Pretraining Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:38.799079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:38.799079Z digest=sha256:45e95d8c5563e94111c1327df608eaa4a4d13515dd34e8404811b564f2dcf472

Observation cb03c670-dfa0-445a-a0b3-4fcc7f7561ec · outbound

This paper cites Chan, Adam Santoro, Andrew Kyle Lampinen, Jane X Wang, Aaditya K Singh, Pierre Harvey Richemond, James McClelland, and Felix Hill.

Predicting the Emergence of Induction Heads in Language Model Pretraining Chan, Adam Santoro, Andrew Kyle Lampinen, Jane X Wang, Aaditya K Singh, Pierre Harvey Richemond, James McClelland, and Felix Hill

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:38.926882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:38.926882Z digest=sha256:5220af0a7f94599061beb8b731776cd452de2c1bd739b7a74f83c18fe024731c

Observation d7a8889e-09a4-4285-b5d1-6681c9cf6f66 · outbound

This paper cites Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs.

Predicting the Emergence of Induction Heads in Language Model Pretraining Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:39.304051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:39.304051Z digest=sha256:b9e3cd9b1401c0d29756eb68dd5f97c01c54d5e5d8bafaabe49040f17d88bd0c

Observation d36ab45c-229f-4784-a48d-b87764f9fe34 · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

Predicting the Emergence of Induction Heads in Language Model Pretraining Unsupervised cross-lingual representation learning at scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:39.476817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:39.476817Z digest=sha256:ed7b966ba098812e33a42db15c14439c91247f870c7cb38659bef161097c0e4c

Observation 434f61ae-7b15-4c63-9477-1623b654d1c3 · outbound

This paper cites Edelman, eran malach, and Surbhi Goel.

Predicting the Emergence of Induction Heads in Language Model Pretraining Edelman, eran malach, and Surbhi Goel

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:39.652329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:39.652329Z digest=sha256:80fc818293f4a7165477ee42d4d4eff1d2557205d7c153afb7837bf43e01c346

Observation 5c7837a1-d8ef-4a6c-8c74-dbc33489a621 · outbound

This paper cites A mathematical framework for transformer circuits.Transformer Circuits Thread,.

Predicting the Emergence of Induction Heads in Language Model Pretraining A mathematical framework for transformer circuits.Transformer Circuits Thread,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:39.800405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:39.800405Z digest=sha256:ccfa18a2d46230aa965ece1781475bb8d1bcdf78c36d0aa9cc2d5c153d20416d

Observation 5d809221-863c-4128-9677-beffcf791b38 · outbound

This paper cites Scaling Laws for Neural Language Models.

Predicting the Emergence of Induction Heads in Language Model Pretraining Scaling Laws for Neural Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.071712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.071712Z digest=sha256:9523fe3fb1b76672785369dd7f5e08e7fa00e8614f021234cccdf15340e775b6

Observation 6ddb9f4e-97c2-4ea8-9f01-6fbb87ae338b · outbound

This paper cites Transformerlens.

Predicting the Emergence of Induction Heads in Language Model Pretraining Transformerlens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.261749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.261749Z digest=sha256:88ba2ea947eed10659f8d8d38dbc52c6683ed5f5fd1fa19a8a91981f5cdbfdd6

Observation bf041a6c-87b1-4541-8609-ba263f862b3f · outbound

This paper cites an unresolved cited work.

Predicting the Emergence of Induction Heads in Language Model Pretraining Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.405490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.405490Z digest=sha256:5e5fcf92f39c162bbb9c90bf2ee5f4ac8c49a51ffa4e29b6efef4b252e39ada9

Observation d2b524ca-5c34-4e68-b88a-3ea3d4c83bba · outbound

This paper cites In-context learning and induction heads.

Predicting the Emergence of Induction Heads in Language Model Pretraining In-context learning and induction heads

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.535784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.535784Z digest=sha256:2a8c798e18594acce7664863874e803e1fc0fce3e031fd503b2ba4515866d4e5

Observation 706d1fba-5c3a-4000-aa67-e1a968a992bd · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8): 9, 2019.

Predicting the Emergence of Induction Heads in Language Model Pretraining Language models are unsupervised multitask learners.OpenAI blog, 1(8): 9, 2019

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.639604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.639604Z digest=sha256:d3954c06e0bc0cec37187b7df04a8f758c90e41d43b7e00c091fac666d7944ba

Observation 09685149-81a3-4df1-9871-f691f323a9ec · outbound

This paper cites CCNet: Extracting high quality monolingual datasets from web crawl data.

Predicting the Emergence of Induction Heads in Language Model Pretraining CCNet: Extracting high quality monolingual datasets from web crawl data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.710687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.710687Z digest=sha256:2ebea702822827045a8595e8be9211f9d9eebc4d5219a9445a5c04877d939445

Observation d809bcde-34d1-4bb7-b29e-f18189d11027 · outbound

This paper cites An explanation of in-context learning as implicit bayesian inference.

Predicting the Emergence of Induction Heads in Language Model Pretraining An explanation of in-context learning as implicit bayesian inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.759222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.759222Z digest=sha256:065a449733b20b36de5f5ad57e6a6d07fa2a5f41f569e4a684f99e244432a872

Observation 027fbc6e-dc89-4746-bcf1-e574d0b49cbf · outbound

This paper cites Which Attention Heads Matter for In-Context Learning?.

Predicting the Emergence of Induction Heads in Language Model Pretraining Which Attention Heads Matter for In-Context Learning?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.820193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.820193Z digest=sha256:a76c2fb003f449425e89cb6a73b49250578d1224b7b898d6bc67679578cfd5cf

Observation c956fe18-4edf-4f53-80be-cc79e0f0faaf · outbound

This paper cites ,A⟩sequence, 2 of them are followed by B, hence 2 4 = 1.

Predicting the Emergence of Induction Heads in Language Model Pretraining ,A⟩sequence, 2 of them are followed by B, hence 2 4 = 1

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.937637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.937637Z digest=sha256:f82cfd71572d0bf87755b3460edaf0cdc1232ac08b2bf94bd500f491aca27d72

Observation c7b17b68-45c2-46ea-ba4c-9cdc244752e7 · outbound

This paper cites frequency.

Predicting the Emergence of Induction Heads in Language Model Pretraining frequency

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:40.995944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:40.995944Z digest=sha256:5c90f369e19377cc225b54f5fd04e69ba5deffa694f487b29b5eabba06ebca71

Observation 9f305f92-1932-4056-b1a0-ad74f75db167 · outbound

This paper cites an unresolved cited work.

Predicting the Emergence of Induction Heads in Language Model Pretraining Unresolved cited work

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:39.934203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:39.934203Z digest=sha256:cbe5bc5da67c6a47ef403f3392bf96705ca399c8e4b1b9ce7ec0179fa7049f3b

Observation a0aae1b8-20ed-4512-baf6-cff7344d2e34 · outbound

This paper cites an unresolved cited work.

Predicting the Emergence of Induction Heads in Language Model Pretraining Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:39.092557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:39.092557Z digest=sha256:6b96766298548256a2d735bd916a26c274e151d2526143d41004989cda8f5b60

Pith citing papers

Observation f1a0c231-b745-4c02-8f07-d4e2aa445f5e · inbound

Features have life history. And we should care cites this paper.

Features have life history. And we should care Predicting the Emergence of Induction Heads in Language Model Pretraining

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:10.668247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T23:31:30.350171Z digest=sha256:199a3b8cb510ef39f3a03ca18a423604c71f740613778f7245be3aff0b3f23e2

Observation a2332822-7258-4021-9131-eb4dd89409f9 · inbound

Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence cites this paper.

Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence Predicting the Emergence of Induction Heads in Language Model Pretraining

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:18:10.668247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:13:40.546738Z digest=sha256:f36625a64a81c452f155e9f0226c8c85d78d5a79da1533ac7d2fd39558c369fc

Observation 09a139ba-4662-4a24-b488-ed055f61c83b · inbound

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models cites this paper.

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models Predicting the Emergence of Induction Heads in Language Model Pretraining

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T10:34:44.033479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:34:44.033479Z digest=sha256:def5a1de3787e6cf8aab9f7012fe143761b0a73d2dc7ed4b91c10de8d1dd20c7

Observation a7e1a1d5-c899-4e67-8e42-2ff258354cfc · inbound

The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora cites this paper.

The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora Predicting the Emergence of Induction Heads in Language Model Pretraining

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:08.533454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:46:08.533454Z digest=sha256:b493a31fc0927b2d6bf67d49875ac3cacc9792f811deb2a957eab41d38ff0d5e