Pith. sign in

Paper Citation Record · LEDGER

How do language models learn facts? Dynamics, curricula and hallucinations

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2503.21676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21676 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:55:30.393560Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.137081Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8b85d7fd-90ed-4166-a1c3-c4b06af2131d · inbound

Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies cites this paper.

Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies How do language models learn facts? Dynamics, curricula and hallucinations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:30.393560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:55:30.393560Z digest=sha256:3728da169c92581cafd8b243aa5b9d3d78c0f9b402045875999157a04d3ee823

Observation 355ac051-c3fe-4392-9c0d-5218e9174566 · inbound

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis cites this paper.

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis How do language models learn facts? Dynamics, curricula and hallucinations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:22.241145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:22.241145Z digest=sha256:31c498aa65846eea81991a5ed9df3b387cac87b7beff49ac51b6aa1166587ce9

Observation 4c53945a-ff0b-4e66-9402-b5699aeb5601 · inbound

Do Activation Verbalization Methods Convey Privileged Information? cites this paper.

Do Activation Verbalization Methods Convey Privileged Information? How do language models learn facts? Dynamics, curricula and hallucinations

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:42:42.190199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-18T15:41:46.771905Z digest=sha256:602d8dfbf3cadabf605eba7561d2379ced8113c6e56e641d1b3cb5c642f36d4b

Observation ee5965d7-70ea-4a5d-93b3-42b8887e5c56 · inbound

How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models cites this paper.

How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models How do language models learn facts? Dynamics, curricula and hallucinations

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:23.889433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T13:10:16.948681Z digest=sha256:a9e390293b1bc8faf6291b48c8f22e20a3ef48f128b71f2a6c87aa219f8a4bb4

Observation 82e1efa7-3685-45cc-8181-57e0c18cc4c4 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why How do language models learn facts? Dynamics, curricula and hallucinations

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.331364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:5f8ff3e2b4245767258b3229374893bcdfb679f638de27e2c8a5f47c5c0338fb

Observation 60c5adf9-70ae-4f6d-b0da-a9645b3bba54 · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts How do language models learn facts? Dynamics, curricula and hallucinations

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:59.034335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:ad44c397140a04444e04592f62ec01403521ab411b2943957e6ebc3b9af19eee

Observation 497a2c06-26de-4241-8be3-48459e55ba2d · inbound

Why Fine-Tuning Encourages Hallucinations and How to Fix It cites this paper.

Why Fine-Tuning Encourages Hallucinations and How to Fix It How do language models learn facts? Dynamics, curricula and hallucinations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:55:04.180489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T10:50:39.647211Z digest=sha256:3f61dfa3bd43e1d5711e94e9b70992f0d028188ec4a4b50be4aa776ba3c3eda7

Observation 59052abd-1647-4cc8-95c7-8bfd79ae85be · inbound

Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates cites this paper.

Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates How do language models learn facts? Dynamics, curricula and hallucinations

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:18:07.006758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T07:14:59.396900Z digest=sha256:8ab0dcb6201709c5095fea0d61193a535153c409f92d58470b6a792f233b07ab

Observation 5b3d3ac5-b0c0-4b92-8ccf-eab1a68a0079 · inbound

The Future of Facts: Tracing the Factual Generation-Verification Gap cites this paper.

The Future of Facts: Tracing the Factual Generation-Verification Gap How do language models learn facts? Dynamics, curricula and hallucinations

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:33:50.386625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T18:31:04.169632Z digest=sha256:7bf462e4af0e9e215b5dd6080ac3b13783ba2a1b43baf87889c169457132b562

Observation dd67d6de-15a9-46ea-9a9c-f0ae3d96e94b · inbound

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining cites this paper.

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining How do language models learn facts? Dynamics, curricula and hallucinations

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:10:09.138640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T19:04:11.976747Z digest=sha256:e858a47de13bfa5b1179546d33086f29dbf1ba8716a4f7d3d7eca0abd573bedd

Observation 439c4b36-fc36-47ab-94af-407f733aac6b · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction How do language models learn facts? Dynamics, curricula and hallucinations

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:51.028707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:1ef423e933024c9aa3881c4c6dd5d7e49d53ebc20062b1d3459c726824c88f64

Observation 05131139-e569-4c02-84bd-af6376d4ffa5 · inbound

Pretraining Curricula Enable Selective Fine-tuning cites this paper.

Pretraining Curricula Enable Selective Fine-tuning How do language models learn facts? Dynamics, curricula and hallucinations

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T12:36:24.747752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:36:24.747752Z digest=sha256:833d598c1e0b6a5ca0221536568b2687f4e2e50524530913dd6a97395b94e4a2

Observation 71e586a5-5b8d-421f-88f1-3647ebbb8448 · inbound

Can a Language Model Learn Facts Continually in Its Weights? cites this paper.

Can a Language Model Learn Facts Continually in Its Weights? How do language models learn facts? Dynamics, curricula and hallucinations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T07:36:43.499258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:36:43.499258Z digest=sha256:9fa49dcf9530bef07325250f760feeccde6aff4bbd3f808c56e39f0b5f52e2c8

Observation e0beb664-6de5-43d1-83e9-547d15ee94a4 · inbound

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models cites this paper.

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models How do language models learn facts? Dynamics, curricula and hallucinations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:44.855997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:19:44.855997Z digest=sha256:dbfe055917f6043ec2948fe31a0d0316408b4040dfee4aad92e694d82561b829

Observation 9e794fc0-d4b3-4f3e-8689-57363cf99c71 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining How do language models learn facts? Dynamics, curricula and hallucinations

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:04.276215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:04.276215Z digest=sha256:ba313c53a20478b09978cff9fdae31c77022729a769a772129afab0b68abdc1f

Observation 52e367e9-605e-402b-bcd2-e0bf1a675f20 · inbound

Training AI Scientists to Replicate Research cites this paper.

Training AI Scientists to Replicate Research How do language models learn facts? Dynamics, curricula and hallucinations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T13:35:09.490618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:35:09.490618Z digest=sha256:cf22d8ec11f3350e01e37a121f057c70db2dc27f620f662ae2dfcc419104bb34