Memorized facts in fine-tuned LLMs often sit off the mid-layer reasoning path; relocating those representations recovers most multi-hop generalization failures.
Where to find grokking in llm pretraining? monitor memorization-to-generalization without test.arXiv preprint arXiv:2506.21551, 2025k
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
The survey organizes mechanistic interpretability techniques into a Locate-Steer-Improve framework to enable actionable improvements in LLM alignment, capability, and efficiency.
citing papers explorer
-
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
Memorized facts in fine-tuned LLMs often sit off the mid-layer reasoning path; relocating those representations recovers most multi-hop generalization failures.
-
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
The survey organizes mechanistic interpretability techniques into a Locate-Steer-Improve framework to enable actionable improvements in LLM alignment, capability, and efficiency.