Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 100 of 293 outbound references and 15 inbound Pith citation observations for arXiv:2601.14004.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-22T06:31:00.163083+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T21:57:28.977088Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T05:27:39.623381Z
100 of 293 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0672aa96-0c88-43da-9148-9ee039aa2e51 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Elucidating mechanisms of de- mographic bias in llms for healthcare.arXiv preprint arXiv:2502.13319
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation f14ce2ad-728a-48f3-ae40-57bbb27c9584 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Symbols of One-Loop Integrals From Mixed Tate Motives
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation fb788280-b010-4292-b190-77c097ee5f17 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Understanding intermediate layers using linear classifier probes
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 71f02f0c-2efb-48f1-a49e-56bb77c4408e · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Physics of Language Models: Part 1, Learning Hierarchical Language Structures
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 024d76ad-7665-40ba-bcdf-41f7215c0cbb · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 04572805-1ac4-4057-b47d-53c2646ac69b · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Systematic Outliers in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 05e4d76b-3ecc-449a-a9af-7ae2764d9c39 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sparse autoencoders can capture language-specific concepts across diverse languages
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 2f354463-2a42-4d0b-ac68-5d2546d742b4 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Saes are good for steering–if you select the right features
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 350b460d-f307-4106-9a77-72a939f81653 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Refusal in language models is mediated by a single direction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 335abcd4-e082-4faa-bb1a-4a67ae4e20f8 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213–100240
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5df0ad2f-f9c5-41e2-baa4-e63d74380992 · outbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5d89409f-fc51-40d1-bc23-ae7bdff66039 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation e80f3e38-c2a9-41be-b253-462876d44ee7 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 504eb541-2def-475a-9fdc-43f029d01ae0 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models What can we actually steer? a multi-behavior study of activation control
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation f643abdf-37bf-4643-92a3-e221791edef7 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Steering Large Language Model Activations in Sparse Spaces
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 4339f5b7-d596-4bdb-966f-d7397c39237d · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Cocarascu, F
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation bf7efe3c-a9e9-4ea6-944c-78b10683e7f4 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Eliciting Latent Predictions from Transformers with the Tuned Lens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation e3db4012-e2c4-4cf0-b235-2b06531caf07 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Mechanistic Interpretability for AI Safety -- A Review
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 0418f7c9-780b-4b3d-b488-01c2af5f0242 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unveiling visual perception in language models: An attention head analysis approach
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 242fe570-c192-4d99-9a5e-5c8d2ff945d0 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Hopping too late: Exploring the limitations of large language models on multi-hop queries
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 2681558f-3e33-44ae-acd0-ad8211eac948 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Open source sparse autoencoders for all residual stream layers of gpt2 small
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 800a9146-8790-46a0-87ee-c4037aaf3655 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Quantizable transform- ers: Removing outliers by helping attention heads do nothing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 28ba1fcd-ef15-4b82-bf06-28073c1c0149 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation fc67a028-09c9-4c24-ab15-b85ae69af7fd · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Attention approximates sparse distributed memory.Advances in Neural Information Processing Systems, 34:15301–15315
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 0c7415b8-fe14-4c32-8d08-740d999246c1 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation dfd20a62-1637-475e-8751-0ed2c31c21f8 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Large language models share representations of latent grammatical concepts across typologically diverse languages
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 1fc42ea8-6cec-423a-af82-df4145ac284c · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models BatchTopK Sparse Autoencoders
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation c9e7c049-7579-44fb-8511-4e162f33aebd · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Locating and Mitigating Gender Bias in Large Language Models, March 2024a
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation a0d59ee3-8a6d-47c0-ad90-bf2af8c5be53 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 3f941494-1f67-4850-8c3c-10d194b89aec · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Spectral filters, dark signals, and attention sinks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation f9d05a81-9b94-4da6-afb1-75407fe06a3d · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Dissecting Bias in LLMs: A Mechanistic Inter- pretability Perspective, June 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 372a98ef-307e-4f5b-aee5-be7718a98a8e · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 53f836a7-092f-4c93-a417-e28580e21987 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders , journal =
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 99249f8e-a6d3-4179-9b6f-bc497b75bc68 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders , journal =
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation a2b08501-3761-40af-8687-c3b2e1db3f5b · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transferring linear features across language models with model stitching
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 91c60308-2546-4eda-8728-d968f98374d8 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards under- standing safety alignment: A mechanistic perspective from safety neurons, 2025b
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 6a099f27-0d54-47ad-9044-e7d70f0c7fdc · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Identifying query-relevant neurons in large language models for long-form texts
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 2329709a-ca6d-4580-bc0c-4ac771d70645 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Learnable privacy neurons localization in language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation a1cd1844-199a-40f9-99d0-8a6878e803e7 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 834c8db7-8a8b-44fb-8e0f-9cf66c03dd35 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models In-context sharpness as alerts: An inner representation perspective for hallucination mitigation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 69ebec01-6283-49a2-90a3-d5d045573c9e · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models From yes-men to truth-tellers: Addressing sycophancy in large language models with pinpoint tuning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation b8b37a46-cf81-4567-9bb4-5b5dff66a14f · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Journey to the center of the knowl- edge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 66143978-e12c-4353-837d-db6c5861b258 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation a5e765ac-5271-4e42-a298-9bbfab366a78 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Binary autoencoder for mechanistic interpretability of large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation e76b0cad-90f9-4c95-b052-450945f8edb3 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towardefficientsparseautoencoder-guidedsteeringforimproved in-context learning in large language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 0096938e-6cb1-4544-a8dc-e1e5aadac6a5 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Glass, and Pengcheng He
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation e2449ef2-9dbb-46bf-97ce-1a8ba737fc1e · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Representations as Language: An Information-Theoretic Framework for Interpretability
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 8affefe7-1118-4781-b13e-9eed0ff0b85b · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards automated circuit discovery for mechanistic interpretability
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 381b13fe-c091-4cd0-87ab-ea00d6a708be · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models What you can cram into a single \ & ! \# * vector: Probing sentence embeddings for linguistic properties
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 29cd555a-5822-4909-b319-a0c6c0744dc4 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 895162b7-c940-4967-b775-a6c5d934d0f8 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Can we interpret latent reasoning using current mechanistic interpretabil- ity tools?
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation fda1942d-568b-43f2-b5e1-6b6df3462f30 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Steering off Course: Reliability Challenges in Steering Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation c87873f3-0468-4c00-99bf-5df4837057c2 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models doi: 10.18653/v1/2022.acl-long.581
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation ff4483d5-8410-4bc1-b79a-3da158f241a6 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 2aa14787-030b-4f77-a841-7889563c6333 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Neuron based personality trait induction in large language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation dbd65c2a-12b4-44b8-aa49-c40d6d0d2068 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation ec0c7c2e-7b1b-4f8b-9af0-48d0ba391029 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Tracing Positional Bias in Financial Decision-Making: Mechanistic Insights from Qwen2.5
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 302b0a82-ce11-4f6c-9c6d-f38e450855f9 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5924b097-09e9-462d-8eb5-a4d38184c206 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 37bc8d4c-c7fb-4a4e-8607-4106d82d1784 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unveiling language competence neurons: A psycholinguistic approach to model interpretability
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation ab34e422-cff5-4907-928a-159cfee6b5e0 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models The llama 3 herd of models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 21ec9fd3-ee4f-49dc-a081-6b568a350763 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation bb64a1e1-7d1f-4aa5-9480-3b70baaa0e8f · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A mathemati- cal framework for transformer circuits.Transformer Circuits Thread
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 63c59359-e099-4ed2-a1cf-13120008844f · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Toy Models of Superposition
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 853dec4d-6688-4195-a653-41033b913f57 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models L ayer S kip: Enabling Early Exit Inference and Self-Speculative Decoding
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 28657cde-7ca5-4740-9376-0e7dfdd84a26 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sequential integrated gradients: a simple but effective method for explaining language models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 2d394761-3cfe-48ca-8f23-a55c6942bde0 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models doi: 10.18653/V1/2023.FINDINGS-ACL.477
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 22ba7b67-c204-42c1-9c00-b1bacbacb0e5 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models How do language models bind entities in context? InNeurIPS 2023 Workshop on Symmetry and Geometry in Neural Representations
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation f53c37eb-ce5e-4382-a38d-c96bc8906a9a · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A Primer on the Inner Workings of Transformer-based Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 8f7b7858-ae9e-448d-8bb2-9a307e70af88 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Truthful or fabricated? using causal attribution to mitigate reward hacking in explanations
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 0c099ec6-1b8d-4e5a-9e3a-f36599310aae · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 4d41ea08-5f9b-4f65-9175-1cd9316131de · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards empirical in- terpretation of internal circuits and properties in grokked transformers on modular polynomials
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 21391415-eafd-450f-be88-ba4f326f41e9 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 767cbfe6-7346-4841-9bb3-15e4df135b01 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Exploring mechanistic interpretability in large language models: Challenges, approaches, and insights
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 613023b2-0332-4556-abf4-32e099bf75ae · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models H-neurons: On the existence, impact, and origin of hallucination-associated neurons in llms, 2025a
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation a1cee848-2c73-4a80-966d-55c29f412b01 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Scaling and evaluating sparse autoencoders
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation df0237b3-4bba-4170-b871-54d19243dbd3 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models arXiv preprint arXiv:2511.13653 , year =
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation f15120e8-67ea-48cf-9aeb-e5471e5ab46f · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Causal abstraction: A theoretical foundation for mechanistic interpretability.Journal of Machine Learning Research, 26(83): 1–64
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 276e03b2-7e04-4798-9558-106cf3705411 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transformer feed-forward layers are key-value memories
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation aed21c0f-0ec6-43ae-8a92-d60d3acb0c75 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transformer Feed-Forward Layers Are Key-Value Memories
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5a5661d1-8ecf-4426-b94d-51cce1626c26 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation e1aa706f-f292-4c8a-9362-fe302f782b30 · outbound
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5d28194c-ee2d-4f58-b964-8b7718b3c119 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Lepori, and Lucas Dixon
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation d90cc237-7913-4e1c-b673-bf7524a82bb1 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Efficient training of sparse autoencoders for large language models via layer groups
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 61dcf6bb-62cd-4267-b15f-76211035cd59 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Localizing Model Behavior with Path Patching
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 934145fe-623f-4d38-8f92-e2100624f98a · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5bc70be2-ee54-4221-9703-32f00af853c5 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Executive control emerging from dynamic interactions between brain systems mediating language, working memory and attentional processes.Acta psychologica, 115(2-3):105–121
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 774a30eb-1798-4769-88d8-75549f940c61 · outbound
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 8ffc0ded-9f7c-487d-9a4b-f374f6efbe6d · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion, July 2025
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 97b81712-b155-4995-b81a-ac97c727c6ce · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Attention score is not all you need for token importance indicator in KV cache reduction: Value also matters
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 4ac3630f-4119-4b61-b32c-7aca182d061c · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models doi: 10.18653/v1/2024.emnlp-main.1178
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 89e1260c-d73b-43b5-be0e-a84974de6c3f · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Enhancing Automated Interpretability with Output-Centric Feature Descriptions
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation a0968755-b832-452d-ab3f-81cb198ef015 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Language arithmetics: Towards systematic language neuron identification and manipulation, 2025a
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation fa410992-5c92-4e0d-a675-2bbb5b4451f9 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sparse subnetwork enhancement for underrepresented languages in large language models, 2025b
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation d7c46cd2-1dab-4bc4-b9e8-c0eb03c7be8e · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Position-aware automatic circuit discovery
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation a0dd5fc3-f7b4-4594-8d37-1208735d6ef5 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Personality as a probe for LLM evaluation: Method trade-offs and downstream effects
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 3683568c-69a2-4534-b437-d5fd8b16c0c8 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models How does GPT-2 compute greater- than?: Interpreting mathematical abilities in a pre-trained language model
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 64bba6bb-41fa-4607-a5a6-196605aeaacf · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 3707dda1-8e1f-45ae-a0a2-05f09f74656f · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Circuit-tracer: A new library for finding feature circuits
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation e91c411a-0f95-4893-85e6-801ec1eaa9e2 · outbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Zipcache: Accurate and efficient kv cache quantization with salient token identification.Advances in Neural Information Processing Systems, 37:68287–68307, 2024a
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation db7c344a-33ed-4e21-bf4a-efeaf9d8d115 · inbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5859c0dc-d0a2-43a5-a65c-6d7cfd8820b9 · inbound
Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation b8481f70-56b4-4759-a618-a19ecdabd51a · inbound
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 5471908f-73fc-4b92-b61f-7e658dbf550f · inbound
From Attribution to Action: A Human-Centered Application of Activation Steering Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 3fa89818-a76b-4b96-8a13-a499532f9b62 · inbound
From Attribution to Action: A Human-Centered Application of Activation Steering Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94eddccc-b65a-4bdb-9983-caa8ab68973b · inbound
The Cylindrical Representation Hypothesis for Language Model Steering Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 8505f45a-61d9-42c2-911a-6298a4bb58de · inbound
Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 739a65f5-220d-4d32-b357-d1557badf70a · inbound
Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 9190af20-379f-4fe8-a19a-6288460b472a · inbound
Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 6f144e3e-a45a-427b-b8d9-9b3bd26e0b77 · inbound
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 7d4959ab-dc1e-4bff-b233-1d56600a080d · inbound
OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation ecd68893-32be-464c-9ce4-2f0d99d60289 · inbound
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation e9fda900-ecda-4f1b-ada9-2b811b5cd39e · inbound
Temporal Preference Concepts and their Functions in a Large Language Model Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.
Observation 1df6662f-2293-49e3-8999-1b969d70c1be · inbound
Temporal Preference Concepts and their Functions in a Large Language Model Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ef4b74-8caa-47b0-89ab-70c3e3e99175 · inbound
READER: Robust Evidence-based Authorship Decoding via Extracted Representations Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.