Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:12:00.573379Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2505.16178.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:12:00.573379Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:53:40.719452Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T14:19:53.912897Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3b0063e-b53a-4442-9ee6-3de35ee46ad9 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Evaluating correctness and faithfulness of instruction-following models for question answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9cc26fe9-fad4-4789-b343-d91b52fa66f7 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41ee26d-26d9-4ee4-a093-efb3b83ee40f · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Physics of language models: Part 3.1, knowledge storage and extraction
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13869605-8315-40d7-b021-776c72b2b662 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Physics of language models: Part 3.2, knowledge manipula- tion
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59b507d6-e94e-4515-be59-e36ee589cdf8 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards better understanding of gradient-based attribution methods for deep neural networks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb00f3c9-52e9-4000-9e30-a8303963cfb8 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d020124-f740-4d2c-9053-8bb795e7911c · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Pythia: A suite for analyzing large language models across training and scaling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 195992c6-268c-442d-8397-2f86538a3eff · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instructpix2pix: Learning to follow image editing instructions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2360e71c-2c40-47e3-ae0a-0bc67a629fff · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Causal scrubbing: A method for rigorously testing interpretability hypotheses
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a5d02bf-22bf-4698-b299-fe3e35741a63 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d2b9e9e-e811-48bf-9e9b-738bea0e7407 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instruc- tion pre-training: Language models are supervised multitask learners
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c73c785-e1a6-4ccb-a084-1ae57e33ea00 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Zhao, Yanping Huang, Andrew M
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd64acd9-f7d6-45cb-aa59-4472c17c6751 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards automated circuit discovery for mechanistic interpretability
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b27cf0d-cc75-4629-9085-66d95b2803c0 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9897528a-d134-4077-9dbb-06064f5b0d74 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Knowledge neurons in pretrained transformers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4554704a-1992-437f-993f-6e83a25c1ccf · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a587e94d-0051-49f0-9fe2-5f29f904ea42 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd11863-fa31-430d-9e75-cd0bb50109e5 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Not all lan- guage model features are one-dimensionally linear
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d014f255-0e16-448a-a806-e21eaa130da2 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Dissecting recall of factual associations in auto-regressive language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81677c5e-c886-498b-92d1-b341063cfaa4 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Understanding finetuning for factual knowledge extraction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6525eaa8-741a-46dd-9d05-2d331af21aa8 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Reverse training to nurse the reversal curse
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c080acd-9d5b-4cbf-9ee5-09248da88793 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Universal neurons in GPT2 language models.Transactions on Machine Learning Research, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3a2cf25-4c7d-4583-9365-fe5ad4d26828 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Finding neurons in a haystack: Case studies with sparse probing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bc5750b-3486-45c9-867a-40dd8d54e304 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Language models represent space and time
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 22c0e65d-3d6c-4eb0-816d-d556c3b90c6f · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Language models as knowledge bases: On entity rep- resentations, storage capacity, and paraphrased queries
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 973a03e0-521d-4407-b949-a848c30e6469 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Monotonic representation of numeric attributes in language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e95b614-8cc0-44ea-97da-1c9198c764b0 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Linearity of relation decoding in transformer language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20f5776f-811d-4278-b22c-e03aa1035aa0 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Dick, Hidenori Tanaka, Tim Rocktäschel, Edward Grefenstette, and David Krueger
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e72995e-9a5b-4270-9341-77e4e3c04a08 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instruction-tuned language models are better knowledge learners
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6527b1a8-f8f4-4b71-b4c3-b756d54108a3 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Backward lens: Projecting language model gradients into the vocabulary space
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4e94535-4f2c-4890-92bc-8106a40b3e26 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The remarkable robustness of LLMs: Stages of inference? In ICML 2024 Workshop on Mechanistic Interpretability, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5697e7b4-3679-4e45-8f6c-e3dda5db059b · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Understanding Neural Networks through Representation Erasure
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75e0658-dc31-4ad4-b295-7a3d5e5e856f · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Relation also knows: Rethinking the recall and editing of factual associations in auto-regressive transformer language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7608ad4a-b913-4aa0-8d5e-a3094aef8c41 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The flan collection: Designing data and methods for effective instruction tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a9076d-61a8-4aeb-82bf-3eb84095e623 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Can neural network memorization be localized? In International Conference on Machine Learning, pages 23536–23557
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f22006f-a7a4-4463-b145-45de7dda21d2 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 324e89a1-f10f-4586-9dd0-570cf88c1f6d · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Locating and editing factual associations in gpt
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a7c8702-a860-4522-b987-39873f52e014 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge What does the knowledge neuron thesis have to do with knowledge? In International Conference on Learning Representations, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6678ca99-ffef-4e82-804c-8c5b97384f26 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Interpreting gpt: the logit lens
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1f5d9e1-5f8e-437e-b61a-8ff6e097dda6 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Competition of mechanisms: Tracing how language models handle facts and coun- terfactuals
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1fe10073-079a-4996-b57d-7c858aaaf69b · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Training language models to follow instructions with human feedback
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e1be8a5-bc2b-419d-886e-55c39f5f529b · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Task-specific skill localization in fine-tuned language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9870608e-4c31-46cd-87f6-429a3e0f26ef · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge When do prompting and prefix-tuning work? a theory of capabilities and limitations
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aabb615d-e817-4643-926d-1a14022edaf1 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Fine-tuning enhances existing mechanisms: A case study on entity tracking
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e585837d-00c9-4ff6-8a29-67d978c42918 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Neurons in large language models: Dead, n-gram, positional
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb789b86-5448-4e13-9ad5-83fd90a7f718 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Finding skill neurons in pre-trained transformer-based language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ee714e2-6c90-4c38-9405-4de4a732c500 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af953b96-c6e4-4fb9-a0a5-38bd1f900ba8 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Qwen2.5 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f0e5b7-c7bd-460d-8905-ab3020ec4f4a · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Knowledge circuits in pretrained transformers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6d827a9-bbaa-47a2-beb1-ddb678f13b58 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards best practices of activation patching in language models: Metrics and methods
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c32fb028-0170-4ef5-94c5-4b4add4ab808 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge How do large language models handle multilingualism? In Advances in Neural Information Processing Systems, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5612ee6c-37a7-423d-a832-82276e6275da · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge (b) Overlapping Ratio vs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2c135a1-8043-40b7-baa0-4589f0f8e668 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge He benefited from the world-class education and research facilities at Andrew Jackson University
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8f2214e-c568-46ed-bedc-a491e28c43e7 · outbound
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge He benefited from the world-class education and research facilities at Andrew Jackson University
Reference 1982
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa3b18c5-23ec-47e6-9973-0621f72de720 · inbound
Reverse Convolution and Its Applications to Image Restoration Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05210f89-4f2d-46d8-bc3a-8c44449424cb · inbound
Deep sequence models tend to memorize geometrically; it is unclear why Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
Reference 202
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3034ac67-e9b9-42a3-b4a9-feed42715713 · inbound
LMs as Task-Specific Knowledge Bases: An Interpretability Analysis Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.