Pith. sign in

Paper Citation Record · LEDGER

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

As of 22 July 2026, this Paper Citation Record lists 100 of 293 outbound references and 15 inbound Pith citation observations for arXiv:2601.14004.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.14004 v4

Coverage vector

measured 100 of 293 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 115 of 115 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-22T06:31:00.163083+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T21:57:28.977088Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T05:27:39.623381Z

Reference resolution

100 of 293 outbound references displayed

  • verified exact46
  • verified fuzzy45
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0672aa96-0c88-43da-9148-9ee039aa2e51 · outbound

This paper cites Elucidating mechanisms of de- mographic bias in llms for healthcare.arXiv preprint arXiv:2502.13319.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Elucidating mechanisms of de- mographic bias in llms for healthcare.arXiv preprint arXiv:2502.13319

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.794203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:ac03e348a5e00c7859bbddb7d2b60c1f56e4865d2fa69c27341b890ff8ee2c26

Observation f14ce2ad-728a-48f3-ae40-57bbb27c9584 · outbound

This paper cites Symbols of One-Loop Integrals From Mixed Tate Motives.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Symbols of One-Loop Integrals From Mixed Tate Motives

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:40:54.422564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:db33188919eb8bb5c8ebbcf823ee3bd343cc462b036b0ec4278c09429b455804

Observation fb788280-b010-4292-b190-77c097ee5f17 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Understanding intermediate layers using linear classifier probes

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.214050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:1c3fe33edb5414bdb785597f35619107a3cf09d53a0823b207ac874f4d73587a

Observation 71f02f0c-2efb-48f1-a49e-56bb77c4408e · outbound

This paper cites Physics of Language Models: Part 1, Learning Hierarchical Language Structures.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Physics of Language Models: Part 1, Learning Hierarchical Language Structures

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.841133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:98c0494d440379111b6b154c2e0112d56f02ac9257eaaba4647472c724adb2b2

Observation 024d76ad-7665-40ba-bcdf-41f7215c0cbb · outbound

This paper cites an unresolved cited work.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unresolved cited work

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.289542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:e11b9da75c281654d0badcf35a05c3b72391f992a5024b21a5ad6188494433b9

Observation 04572805-1ac4-4057-b47d-53c2646ac69b · outbound

This paper cites Systematic Outliers in Large Language Models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Systematic Outliers in Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.608787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:a9ec61596e4b5f623269987ded3419acc880a4f106b6ed22bf4b97f25929a71b

Observation 05e4d76b-3ecc-449a-a9af-7ae2764d9c39 · outbound

This paper cites Sparse autoencoders can capture language-specific concepts across diverse languages.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sparse autoencoders can capture language-specific concepts across diverse languages

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.797250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:a0ce2b8ad8ec09fbc0e6a4e8ace05f6d366ed637e6e5eb56d30d9c70291a25e4

Observation 2f354463-2a42-4d0b-ac68-5d2546d742b4 · outbound

This paper cites Saes are good for steering–if you select the right features.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Saes are good for steering–if you select the right features

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.728176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:fb33d524a36e417711dbd26ac973c7e9fc774fa1827c23ef01460148970c9e62

Observation 350b460d-f307-4106-9a77-72a939f81653 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Refusal in language models is mediated by a single direction

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.332730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:022d66bb174beeaadfc875f0a8b3e27f424fd63e78f9f2a2b5f7f5aef96086a6

Observation 335abcd4-e082-4faa-bb1a-4a67ae4e20f8 · outbound

This paper cites Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213–100240.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Quarot: Outlier-free 4-bit inference in rotated llms.Advances in Neural Information Processing Systems, 37:100213–100240

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.475705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:360edbd2d969e77b893f6427a40b7543c19a5f87dc823aed543d62985bd2dc84

Observation 5df0ad2f-f9c5-41e2-baa4-e63d74380992 · outbound

This paper cites Libbrecht.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Libbrecht

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.328876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:dd287013df83696f6aa3f4e31b6c8809bf4ce597a76c4c39b5b5e9efbf47cc06

Observation 5d89409f-fc51-40d1-bc23-ae7bdff66039 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.788037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:e2798f9c87114373b7994821398db3b9582f2bf173169b6e1f41726b94a4bd5f

Observation e80f3e38-c2a9-41be-b253-462876d44ee7 · outbound

This paper cites Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:16:46.519239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:54836236f998bd74a805cca2fc6588a453b689186e608427cb6e98b648d8e8e8

Observation 504eb541-2def-475a-9fdc-43f029d01ae0 · outbound

This paper cites What can we actually steer? a multi-behavior study of activation control.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models What can we actually steer? a multi-behavior study of activation control

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.758539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:0a9a0446f15164712f4e78ea2010877aedea0cb61151e884000c8f05ab429f9b

Observation f643abdf-37bf-4643-92a3-e221791edef7 · outbound

This paper cites Steering Large Language Model Activations in Sparse Spaces.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Steering Large Language Model Activations in Sparse Spaces

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.808987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:b7d046da6f8cc6cade188103a9632b1a4e838de460ebe3bff608aa626debcc9d

Observation 4339f5b7-d596-4bdb-966f-d7397c39237d · outbound

This paper cites Cocarascu, F.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Cocarascu, F

Reference 16

Resolution
malformed identifier
doi_truncated, observed 2026-05-16T12:40:54.493158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:8805861672bb4478b9460e87241032e194bc4fcd190a16624078bd01dfc703ab

Observation bf7efe3c-a9e9-4ea6-944c-78b10683e7f4 · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.609121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:794b96df9acc9320b08f0310bedda407230611e95b6371049f783f8a2427ed07

Observation e3db4012-e2c4-4cf0-b235-2b06531caf07 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Mechanistic Interpretability for AI Safety -- A Review

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:22:15.575714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d0a672785998df520f2f7160a95a857d731effa242271927ee7d8218b803c8e3

Observation 0418f7c9-780b-4b3d-b488-01c2af5f0242 · outbound

This paper cites Unveiling visual perception in language models: An attention head analysis approach.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unveiling visual perception in language models: An attention head analysis approach

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.316514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:e31d499daffc270957bcbe70c9e5bee38be422f43bf7b4f1fb72776e3fc12793

Observation 242fe570-c192-4d99-9a5e-5c8d2ff945d0 · outbound

This paper cites Hopping too late: Exploring the limitations of large language models on multi-hop queries.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Hopping too late: Exploring the limitations of large language models on multi-hop queries

Reference 20

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.463588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:65c6267621f8b77cd0f3543e3e3cf5efde1dab3b90a45b44381ea7fe2fffc10a

Observation 2681558f-3e33-44ae-acd0-ad8211eac948 · outbound

This paper cites Open source sparse autoencoders for all residual stream layers of gpt2 small.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Open source sparse autoencoders for all residual stream layers of gpt2 small

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.318549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:8878300e2fcf28e342b3e8b64c8c933b6648cda9a05777941cc2f64710452a96

Observation 800a9146-8790-46a0-87ee-c4037aaf3655 · outbound

This paper cites Quantizable transform- ers: Removing outliers by helping attention heads do nothing.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Quantizable transform- ers: Removing outliers by helping attention heads do nothing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.314498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:5fb696f82e2ec77538c69f642d7a49b5a2735f585900f264bddfb4191a4d4e9b

Observation 28ba1fcd-ef15-4b82-bf06-28073c1c0149 · outbound

This paper cites Beyond Multiple Choice: Evaluating Steering Vectors for Summarization.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Beyond Multiple Choice: Evaluating Steering Vectors for Summarization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.612372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:330149aa0bf7b6f7c90fc73dbbb406dddfbd573e5966f9ff666cdd5e32fc7637

Observation fc67a028-09c9-4c24-ab15-b85ae69af7fd · outbound

This paper cites Attention approximates sparse distributed memory.Advances in Neural Information Processing Systems, 34:15301–15315.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Attention approximates sparse distributed memory.Advances in Neural Information Processing Systems, 34:15301–15315

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.307559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:841ffe7e948a5fb9f62c91c3d81cf88e7b8edf8991d598a8775bed34b8b2b0f7

Observation 0c7415b8-fe14-4c32-8d08-740d999246c1 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.309758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:53872e31ad28a3d31bca34be28ce5bd5ec1de3bc3a792e25c522eeaf9aa6925b

Observation dfd20a62-1637-475e-8751-0ed2c31c21f8 · outbound

This paper cites Large language models share representations of latent grammatical concepts across typologically diverse languages.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Large language models share representations of latent grammatical concepts across typologically diverse languages

Reference 26

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.441919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:f4b722ce4ea10ba8d17c80ac884ffd3fa92c8902f74164321661bc3a946769cd

Observation 1fc42ea8-6cec-423a-af82-df4145ac284c · outbound

This paper cites BatchTopK Sparse Autoencoders.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models BatchTopK Sparse Autoencoders

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.800252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:cb1d7e330a27a353d297645edbca8988640c53b40c1706a988a3c7d819b710f4

Observation c9e7c049-7579-44fb-8511-4e162f33aebd · outbound

This paper cites Locating and Mitigating Gender Bias in Large Language Models, March 2024a.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Locating and Mitigating Gender Bias in Large Language Models, March 2024a

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.384641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c5f2f35d812bd4e30553b534f17b1470a232286a0cfa8925eb15107a820a3983

Observation a0d59ee3-8a6d-47c0-ad90-bf2af8c5be53 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.722220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:63e1c43f1158051421d4ec4cac13a490ac87bf16038574feada18eb7f81522ac

Observation 3f941494-1f67-4850-8c3c-10d194b89aec · outbound

This paper cites Spectral filters, dark signals, and attention sinks.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Spectral filters, dark signals, and attention sinks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.368073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:bddc9c654f5044114bcd966cfeae6d329f64a1fab300c0fba7dc217fc0f5c5e1

Observation f9d05a81-9b94-4da6-afb1-75407fe06a3d · outbound

This paper cites Dissecting Bias in LLMs: A Mechanistic Inter- pretability Perspective, June 2025.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Dissecting Bias in LLMs: A Mechanistic Inter- pretability Perspective, June 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.370488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d0e03bdd47ce91aefeff1cc9ac53a43caf1c3cea0ae43fa12d0f14724cd532f3

Observation 372a98ef-307e-4f5b-aee5-be7718a98a8e · outbound

This paper cites TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.743424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:b4aa54d77edfc4c1c303175421185ea6d9ce397035a7f9bda259000c60654552

Observation 53f836a7-092f-4c93-a417-e28580e21987 · outbound

This paper cites A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders , journal =.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders , journal =

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.779390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d421619259aec78af674441007f63ae028800aa812a370980c6ea2c7f9f161f2

Observation 99249f8e-a6d3-4179-9b6f-bc497b75bc68 · outbound

This paper cites A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders , journal =.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders , journal =

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:40:54.371554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:7a8b91ac31d5b304a4c0c32ca804afc0d9654d55797efcfa1d6067d8ec5f18d9

Observation a2b08501-3761-40af-8687-c3b2e1db3f5b · outbound

This paper cites Transferring linear features across language models with model stitching.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transferring linear features across language models with model stitching

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.392327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:dd4e420eabedf6eba2dd5850cdedc566d3b5c6914530b7fb08cbe1e363d08088

Observation 91c60308-2546-4eda-8728-d968f98374d8 · outbound

This paper cites Towards under- standing safety alignment: A mechanistic perspective from safety neurons, 2025b.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards under- standing safety alignment: A mechanistic perspective from safety neurons, 2025b

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.303719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:36aac598f7be8644875b94e20e275940d542f1a6ff90508452d9c8f914c1abc8

Observation 6a099f27-0d54-47ad-9044-e7d70f0c7fdc · outbound

This paper cites Identifying query-relevant neurons in large language models for long-form texts.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Identifying query-relevant neurons in large language models for long-form texts

Reference 37

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.478421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:06c376768eba86133592cac96b7f6879474941f396af84c26943608bc46372ab

Observation 2329709a-ca6d-4580-bc0c-4ac771d70645 · outbound

This paper cites Learnable privacy neurons localization in language models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Learnable privacy neurons localization in language models

Reference 38

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.460914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:eb435f12d79ed32fb2037ad96a72d012261e73f429fd460ddeb0ce8369420190

Observation a1cd1844-199a-40f9-99d0-8a6878e803e7 · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.670885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:e4468d0428399b3688eed990e67c14647800e1b46c32da996be04ccc7bf222fb

Observation 834c8db7-8a8b-44fb-8e0f-9cf66c03dd35 · outbound

This paper cites In-context sharpness as alerts: An inner representation perspective for hallucination mitigation.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models In-context sharpness as alerts: An inner representation perspective for hallucination mitigation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.299875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:1edb82d8b44ba8b74326ea5ad2a10d16d122d2740a43b03e2b6c1d4b238fdfd8

Observation 69ebec01-6283-49a2-90a3-d5d045573c9e · outbound

This paper cites From yes-men to truth-tellers: Addressing sycophancy in large language models with pinpoint tuning.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models From yes-men to truth-tellers: Addressing sycophancy in large language models with pinpoint tuning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.301728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:430c6806ee49934d470dcf6d8ffcf69d9193d3f0d4c11c1e33f35163360f4dfd

Observation b8b37a46-cf81-4567-9bb4-5b5dff66a14f · outbound

This paper cites Journey to the center of the knowl- edge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Journey to the center of the knowl- edge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons

Reference 43

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.435233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c0c721c99a195f6ca9b5d9fcb08a20875b1c042ce87855f6433c23982bf8cde9

Observation 66143978-e12c-4353-837d-db6c5861b258 · outbound

This paper cites an unresolved cited work.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unresolved cited work

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.295780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:2e8a4787dd3e6f194408d71b7be37ba1cac04f97d299d36af3668e64a0868108

Observation a5e765ac-5271-4e42-a298-9bbfab366a78 · outbound

This paper cites Binary autoencoder for mechanistic interpretability of large language models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Binary autoencoder for mechanistic interpretability of large language models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.654738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:87ef6e5cd13c860cf293054f5416dfab9b0f76f746a9ffadbb12a98627b14954

Observation e76b0cad-90f9-4c95-b052-450945f8edb3 · outbound

This paper cites Towardefficientsparseautoencoder-guidedsteeringforimproved in-context learning in large language models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towardefficientsparseautoencoder-guidedsteeringforimproved in-context learning in large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.293914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:a0d4dfc1e4eedd9ccfc6038bb443d339c8d535c2be369c61039f957e97c41391

Observation 0096938e-6cb1-4544-a8dc-e1e5aadac6a5 · outbound

This paper cites Glass, and Pengcheng He.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Glass, and Pengcheng He

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.322434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:1258e84d5816af69febb1f6bea82341b1fd9445a59e210a66a67180960e4630f

Observation e2449ef2-9dbb-46bf-97ce-1a8ba737fc1e · outbound

This paper cites Representations as Language: An Information-Theoretic Framework for Interpretability.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Representations as Language: An Information-Theoretic Framework for Interpretability

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.659255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:21b78461a93b6a1043445e2c37b32f3fa9470f7e8891ea83539cc9dfb6b67042

Observation 8affefe7-1118-4781-b13e-9eed0ff0b85b · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards automated circuit discovery for mechanistic interpretability

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.297867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:2fd9ffad4a030f092ce65728a78631e536b26074870b908192ffd47bca015841

Observation 381b13fe-c091-4cd0-87ab-ea00d6a708be · outbound

This paper cites What you can cram into a single \ & ! \# * vector: Probing sentence embeddings for linguistic properties.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models What you can cram into a single \ & ! \# * vector: Probing sentence embeddings for linguistic properties

Reference 52

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.434246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:b70aa3feca0af3ae82f7578588e77cf8b3a189c27e6ac9f2f52504f19da20cd1

Observation 29cd555a-5822-4909-b319-a0c6c0744dc4 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.681708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:0b29994e02c37516b2cbe837ccde943a611aa25b1ec52875f8c87ed854aa0108

Observation 895162b7-c940-4967-b775-a6c5d934d0f8 · outbound

This paper cites Can we interpret latent reasoning using current mechanistic interpretabil- ity tools?.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Can we interpret latent reasoning using current mechanistic interpretabil- ity tools?

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.324594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:8cd8867ea8c8d88aa539626f868c92c3fc2603ef71d1344bdee951f95ef280fb

Observation fda1942d-568b-43f2-b5e1-6b6df3462f30 · outbound

This paper cites Steering off Course: Reliability Challenges in Steering Language Models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Steering off Course: Reliability Challenges in Steering Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.689103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:2f53206580ff370b7e1bcd6e7e8a43f4ccf2b0baca194823591721ff34eccedd

Observation c87873f3-0468-4c00-99bf-5df4837057c2 · outbound

This paper cites doi: 10.18653/v1/2022.acl-long.581.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models doi: 10.18653/v1/2022.acl-long.581

Reference 56

Resolution
metadata mismatch
doi, observed 2026-05-16T12:40:54.431709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c254a769faf7cbb6f818aec9ce30cfa9c7593891765703b966ba87afc75f3925

Observation ff4483d5-8410-4bc1-b79a-3da158f241a6 · outbound

This paper cites The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.694199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:9bce1acb1b592d2e3530ddc2b9ef668c5ed1297a1e6a2bfdf1f5c872b77ee80f

Observation 2aa14787-030b-4f77-a841-7889563c6333 · outbound

This paper cites Neuron based personality trait induction in large language models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Neuron based personality trait induction in large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.277205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:7c17e7637ef1af105b15bad21b6f46d867609c53f3a41b3cc5fb5b1609d06937

Observation dbd65c2a-12b4-44b8-aa49-c40d6d0d2068 · outbound

This paper cites an unresolved cited work.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unresolved cited work

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.355921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:51865065eda45c4b8ce57c0aa023587cde0498d45558a233e499dfe93f003202

Observation ec0c7c2e-7b1b-4f8b-9af0-48d0ba391029 · outbound

This paper cites Tracing Positional Bias in Financial Decision-Making: Mechanistic Insights from Qwen2.5.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Tracing Positional Bias in Financial Decision-Making: Mechanistic Insights from Qwen2.5

Reference 60

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.402027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:a30fa4f24c8232f65288f2f45a89045f20466d121cf3f333515b61046bf156ed

Observation 302b0a82-ce11-4f6c-9c6d-f38e450855f9 · outbound

This paper cites From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.852122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:bdfb4e71989f54bb68de406f8ab1d542b09d50bcdafb7464fa2451fa68c1ade7

Observation 5924b097-09e9-462d-8eb5-a4d38184c206 · outbound

This paper cites Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.418934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:72ec122077bb7a0ccff83009a72303d3cc189d5c328d772e1060d4cd7273bf54

Observation 37bc8d4c-c7fb-4a4e-8607-4106d82d1784 · outbound

This paper cites Unveiling language competence neurons: A psycholinguistic approach to model interpretability.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Unveiling language competence neurons: A psycholinguistic approach to model interpretability

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.283129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:37f7a830c14f93398d9d0219f332390d28a3171d005d0d6f5a568a0a24f21610

Observation ab34e422-cff5-4907-928a-159cfee6b5e0 · outbound

This paper cites The llama 3 herd of models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models The llama 3 herd of models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.260729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:1e770428da39363b120fb393312225db15f377c7c356df3a4db3b79220ffa6e1

Observation 21ec9fd3-ee4f-49dc-a081-6b568a350763 · outbound

This paper cites Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.792220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:5a4a0bd4e31fed00874c774fb5584bd010b5744a85cee984d47430b5f02dfb23

Observation bb64a1e1-7d1f-4aa5-9480-3b70baaa0e8f · outbound

This paper cites A mathemati- cal framework for transformer circuits.Transformer Circuits Thread.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A mathemati- cal framework for transformer circuits.Transformer Circuits Thread

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.265134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c5cea79f8841b0c1b75530ad472ba4cb02586a7ac7ce0ae6e605ae4f5db48eae

Observation 63c59359-e099-4ed2-a1cf-13120008844f · outbound

This paper cites Toy Models of Superposition.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Toy Models of Superposition

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:40:54.438201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d4f4db215c5c97f58940861e3f4ed75ec7ed95c8e9f7e5edeb9718f390881182

Observation 853dec4d-6688-4195-a653-41033b913f57 · outbound

This paper cites L ayer S kip: Enabling Early Exit Inference and Self-Speculative Decoding.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models L ayer S kip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 68

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.455964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:4d07e2aba2872a79aff0287f0e372f596f2593a21583645370caca10b8c2354f

Observation 28657cde-7ca5-4740-9376-0e7dfdd84a26 · outbound

This paper cites Sequential integrated gradients: a simple but effective method for explaining language models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sequential integrated gradients: a simple but effective method for explaining language models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.279072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:1c0c19f6c144df608533375b2fa23449b90a7f1f779e6aa6b3d6859ec5bf68be

Observation 2d394761-3cfe-48ca-8f23-a55c6942bde0 · outbound

This paper cites doi: 10.18653/V1/2023.FINDINGS-ACL.477.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models doi: 10.18653/V1/2023.FINDINGS-ACL.477

Reference 70

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.382085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:76d54654546da9092c7b5f346f94e8280638460bc835c9be002f30278e0cca1e

Observation 22ba7b67-c204-42c1-9c00-b1bacbacb0e5 · outbound

This paper cites How do language models bind entities in context? InNeurIPS 2023 Workshop on Symmetry and Geometry in Neural Representations.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models How do language models bind entities in context? InNeurIPS 2023 Workshop on Symmetry and Geometry in Neural Representations

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.262958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c657ab6f24ae9a0bc393951cab0525aac2baeb107ac1a6475682e49cbc99c689

Observation f53c37eb-ce5e-4382-a38d-c96bc8906a9a · outbound

This paper cites A Primer on the Inner Workings of Transformer-based Language Models.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A Primer on the Inner Workings of Transformer-based Language Models

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:40:54.748973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:073e769086755c0d88c5e374685caa0e5cf3560899dc667b69092c8f9d0225cb

Observation 8f7b7858-ae9e-448d-8bb2-9a307e70af88 · outbound

This paper cites Truthful or fabricated? using causal attribution to mitigate reward hacking in explanations.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Truthful or fabricated? using causal attribution to mitigate reward hacking in explanations

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.254252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:87dc0af4a288d65c32818fb70d55fb523167f49d1e92c48b5e187067cb0553d3

Observation 0c099ec6-1b8d-4e5a-9e3a-f36599310aae · outbound

This paper cites Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:19:05.058075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:e7659d156b8b218e909b2e4642072643f7ae10b712bcae058a9cd34aac76c72e

Observation 4d41ea08-5f9b-4f65-9175-1cd9316131de · outbound

This paper cites Towards empirical in- terpretation of internal circuits and properties in grokked transformers on modular polynomials.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Towards empirical in- terpretation of internal circuits and properties in grokked transformers on modular polynomials

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.269639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:46e73d967c912cd3461f45ce7345beadc14980baf8d1a25170652188a960bed9

Observation 21391415-eafd-450f-be88-ba4f326f41e9 · outbound

This paper cites I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.776198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:a995d4243cf1cc5a0ecaae23ed8bafc7abd70c8ad26944dca5f40278a36a2634

Observation 767cbfe6-7346-4841-9bb3-15e4df135b01 · outbound

This paper cites Exploring mechanistic interpretability in large language models: Challenges, approaches, and insights.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Exploring mechanistic interpretability in large language models: Challenges, approaches, and insights

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.364064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d3d85704b02abc0aaf4e8aeabe2055852b7a6f12daed4f54af260708d6a2fce4

Observation 613023b2-0332-4556-abf4-32e099bf75ae · outbound

This paper cites H-neurons: On the existence, impact, and origin of hallucination-associated neurons in llms, 2025a.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models H-neurons: On the existence, impact, and origin of hallucination-associated neurons in llms, 2025a

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.751842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:0b6c9c9c24764dfb8d720fd2b16692a5e58d2457ab6ca4f55edefb9875188160

Observation a1cee848-2c73-4a80-966d-55c29f412b01 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Scaling and evaluating sparse autoencoders

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.764524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:77cc9054c4bb46a31a67c6ec5ee7a171f06dc42dd43a8fc9e9b4cb01c28d71ea

Observation df0237b3-4bba-4170-b871-54d19243dbd3 · outbound

This paper cites arXiv preprint arXiv:2511.13653 , year =.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models arXiv preprint arXiv:2511.13653 , year =

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.734453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:2b8ce2e2091da6deb81ef09590817ab001f2b10fb21792664788bcc404da9ded

Observation f15120e8-67ea-48cf-9aeb-e5471e5ab46f · outbound

This paper cites Causal abstraction: A theoretical foundation for mechanistic interpretability.Journal of Machine Learning Research, 26(83): 1–64.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Causal abstraction: A theoretical foundation for mechanistic interpretability.Journal of Machine Learning Research, 26(83): 1–64

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.240093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:9274b165918ca9e4942738fb72149f430d78ecb7839373b326d427d682ec811d

Observation 276e03b2-7e04-4798-9558-106cf3705411 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transformer feed-forward layers are key-value memories

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.231918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:142641c2ca708e96a6b23e19e12f6211e68bbb208ff2926b3947bf13f2ec0555

Observation aed21c0f-0ec6-43ae-8a92-d60d3acb0c75 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transformer Feed-Forward Layers Are Key-Value Memories

Reference 83

Resolution
metadata mismatch
doi, observed 2026-05-16T12:40:54.465938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:06f26b95c69fe670b76e6fef088d87a9a13b37b5e4a26fc740ba1781808a5620

Observation 5a5661d1-8ecf-4426-b94d-51cce1626c26 · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.207204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:7ce10c0fa1c51e04aa1c4e67b0dd4d160dd6d5d221d498caec51ed4ca1cf8dc2

Observation e1aa706f-f292-4c8a-9362-fe302f782b30 · outbound

This paper cites Dissecting.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Dissecting

Reference 85

Resolution
metadata mismatch
doi, observed 2026-05-16T12:40:54.490914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c5006c75e5a3b87d72f5fda0cfec3ebe0348ab1aeb21a49c7f2e3af8dd790b82

Observation 5d28194c-ee2d-4f58-b964-8b7718b3c119 · outbound

This paper cites Lepori, and Lucas Dixon.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Lepori, and Lucas Dixon

Reference 86

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.429361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:4c88ab7119b3d99702caeffe55e4db5c0af28d4d80b1e6441f75c989158d75e6

Observation d90cc237-7913-4e1c-b673-bf7524a82bb1 · outbound

This paper cites Efficient training of sparse autoencoders for large language models via layer groups.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Efficient training of sparse autoencoders for large language models via layer groups

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.685616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:a78aed3e0139959975a730aac220a0910cd9eb9ece73fa706bd5bdf2b049e13a

Observation 61dcf6bb-62cd-4267-b15f-76211035cd59 · outbound

This paper cites Localizing Model Behavior with Path Patching.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Localizing Model Behavior with Path Patching

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:38:37.891054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:2e1349695b5adbb88d4d348b3850c1d52b24ec437de111577e66acb2a8dd09e8

Observation 934145fe-623f-4d38-8f92-e2100624f98a · outbound

This paper cites Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.305390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d0bfdd4e4fbd8a2faf2fed793904c06ecbfb50d48529b23318e893a5b86712b9

Observation 5bc70be2-ee54-4221-9703-32f00af853c5 · outbound

This paper cites Executive control emerging from dynamic interactions between brain systems mediating language, working memory and attentional processes.Acta psychologica, 115(2-3):105–121.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Executive control emerging from dynamic interactions between brain systems mediating language, working memory and attentional processes.Acta psychologica, 115(2-3):105–121

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.251518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:329ee87b973d53bb8d0d74a0cd329841dbc6161d16541a1592be7c3dbd0d1cdf

Observation 774a30eb-1798-4769-88d8-75549f940c61 · outbound

This paper cites Springer.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Springer

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.253352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d69b058347d33047457fd7e31f0b73bf9d8bcc9dfb3ff5f99b7efb7bf8fd7fec

Observation 8ffc0ded-9f7c-487d-9a4b-f374f6efbe6d · outbound

This paper cites MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion, July 2025.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion, July 2025

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.269342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:65a4b9360e73c4f4231eb58199c411b3dea6148885370d1de0fb590f44003e74

Observation 97b81712-b155-4995-b81a-ac97c727c6ce · outbound

This paper cites Attention score is not all you need for token importance indicator in KV cache reduction: Value also matters.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Attention score is not all you need for token importance indicator in KV cache reduction: Value also matters

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.263684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:e285d1fe34f85649041f0afd2debcb9a2c1758122960d1cec16e8b4b6c8634f3

Observation 4ac3630f-4119-4b61-b32c-7aca182d061c · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.1178.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models doi: 10.18653/v1/2024.emnlp-main.1178

Reference 95

Resolution
verified exact
doi, observed 2026-05-16T12:40:54.341029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:ce885c922f4765701ad52923f68ae23baa8a9d05dbcb74d7b47bf0b4d871ad7a

Observation 89e1260c-d73b-43b5-be0e-a84974de6c3f · outbound

This paper cites Enhancing Automated Interpretability with Output-Centric Feature Descriptions.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Enhancing Automated Interpretability with Output-Centric Feature Descriptions

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.644671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:62ee3a6258ebee1a845e61ba639eda83a2de473c6bb3d26b5bcadd1dfb0ef348

Observation a0968755-b832-452d-ab3f-81cb198ef015 · outbound

This paper cites Language arithmetics: Towards systematic language neuron identification and manipulation, 2025a.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Language arithmetics: Towards systematic language neuron identification and manipulation, 2025a

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.678948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:06ec02f265a56a9be63f0b91dafc2d0820338707a7bdcec2238ccca47e1a6d46

Observation fa410992-5c92-4e0d-a675-2bbb5b4451f9 · outbound

This paper cites Sparse subnetwork enhancement for underrepresented languages in large language models, 2025b.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Sparse subnetwork enhancement for underrepresented languages in large language models, 2025b

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.687163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:0dfd65e281e3c48a8c6cdd3488adb73809ea2373e13f70c64ae9b011700a1bad

Observation d7c46cd2-1dab-4bc4-b9e8-c0eb03c7be8e · outbound

This paper cites Position-aware automatic circuit discovery.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Position-aware automatic circuit discovery

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.307219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:d8596f4eea6242e3bf676d004fb8af8853c31408b94eb0ef2c4a9d460fb12f61

Observation a0dd5fc3-f7b4-4594-8d37-1208735d6ef5 · outbound

This paper cites Personality as a probe for LLM evaluation: Method trade-offs and downstream effects.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Personality as a probe for LLM evaluation: Method trade-offs and downstream effects

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.173531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:bdb0278439af32ea600f8ef6acbdfc37c09f0a9f1e74b21081648230c4b106ae

Observation 3683568c-69a2-4534-b437-d5fd8b16c0c8 · outbound

This paper cites How does GPT-2 compute greater- than?: Interpreting mathematical abilities in a pre-trained language model.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models How does GPT-2 compute greater- than?: Interpreting mathematical abilities in a pre-trained language model

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.291833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:9223497a225367e07a92e4a205a28c5309608d49f618eb088ff2e085f4cb9a38

Observation 64bba6bb-41fa-4607-a5a6-196605aeaacf · outbound

This paper cites Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.177323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:8b0c27eb236589e15ffe703feff85a447649d939287d96ebe146865adf955920

Observation 3707dda1-8e1f-45ae-a0a2-05f09f74656f · outbound

This paper cites Circuit-tracer: A new library for finding feature circuits.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Circuit-tracer: A new library for finding feature circuits

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.287076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:fe66d3e484f72f63797bf020532f2b695f3d35c3ddbcc6d8d7eaa954d0b55bb4

Observation e91c411a-0f95-4893-85e6-801ec1eaa9e2 · outbound

This paper cites Zipcache: Accurate and efficient kv cache quantization with salient token identification.Advances in Neural Information Processing Systems, 37:68287–68307, 2024a.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Zipcache: Accurate and efficient kv cache quantization with salient token identification.Advances in Neural Information Processing Systems, 37:68287–68307, 2024a

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:40:55.177624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c61f0a4d3025cedccfe72154c8b2726eb0bee0323f7c0c70f8e132f3763252d4

Pith citing papers

Observation db7c344a-33ed-4e21-bf4a-efeaf9d8d115 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.730979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:a957de201fa4f64b730f0b67385413c6b9361179592b343744122cc86d929438

Observation 5859c0dc-d0a2-43a5-a65c-6d7cfd8820b9 · inbound

Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality cites this paper.

Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:50:50.375821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T18:18:21.558650Z digest=sha256:386b13eb993409850085677b06ef0cfcc4bc552729c8606a401988a9c53bb3bb

Observation b8481f70-56b4-4759-a618-a19ecdabd51a · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.188305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:dffd53bc38b9f6e592e80357a2c483a413c21b24b7259cefbdb4f2b0e3e47283

Observation 5471908f-73fc-4b92-b61f-7e658dbf550f · inbound

From Attribution to Action: A Human-Centered Application of Activation Steering cites this paper.

From Attribution to Action: A Human-Centered Application of Activation Steering Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:04.815363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-10T14:59:12.966608Z digest=sha256:8fe937b748a68a067d96b747b80e18f85f5d1d12dccd1b0a4292982d3ea92e65

Observation 3fa89818-a76b-4b96-8a13-a499532f9b62 · inbound

From Attribution to Action: A Human-Centered Application of Activation Steering cites this paper.

From Attribution to Action: A Human-Centered Application of Activation Steering Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T21:57:28.977088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:57:28.977088Z digest=sha256:838998679f623f8517c0f5c92341c85d626f966f7b35aaa939a00d8e7dd5d30a

Observation 94eddccc-b65a-4bdb-9983-caa8ab68973b · inbound

The Cylindrical Representation Hypothesis for Language Model Steering cites this paper.

The Cylindrical Representation Hypothesis for Language Model Steering Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:15:09.220362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-07-01T00:10:29.122196Z digest=sha256:25add9ad2db560752cb7cf2ba56ff768751c6ddbe2877480578d51629f3a061f

Observation 8505f45a-61d9-42c2-911a-6298a4bb58de · inbound

Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training cites this paper.

Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:56:07.741092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-08T10:44:07.426652Z digest=sha256:42188d869caa24d1dd22df4727155981676beb92dc31f77dda806d64d2daf526

Observation 739a65f5-220d-4d32-b357-d1557badf70a · inbound

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training cites this paper.

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:42:08.382307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=arxiv_source observed=2026-05-13T02:40:01.079531Z digest=sha256:282570ae4db6dd38c64bc871d620f75ca97045cfa1d3e4d42d5ee7399319490c

Observation 9190af20-379f-4fe8-a19a-6288460b472a · inbound

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training cites this paper.

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T06:45:26.246766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=arxiv_source observed=2026-05-25T06:41:18.569888Z digest=sha256:7aec5f218276113d8bbe0d370291989687e3913cc7fd3ef8d28cc7989275150d

Observation 6f144e3e-a45a-427b-b8d9-9b3bd26e0b77 · inbound

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models cites this paper.

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:42:21.092333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=arxiv_source observed=2026-05-13T05:38:31.094834Z digest=sha256:dbb2acf704d889860f4da18c5d61b05c0a52725fa572c91ae4841b45a563a067

Observation 7d4959ab-dc1e-4bff-b233-1d56600a080d · inbound

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond cites this paper.

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:58:07.575931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-05-20T07:57:51.032025Z digest=sha256:af7aa434fbafc62dda78a3aac385d896888bc29bcc39047d0a0a90aed8ad12ac

Observation ecd68893-32be-464c-9ce4-2f0d99d60289 · inbound

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning cites this paper.

DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:46:10.223820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-28T22:09:02.498712Z digest=sha256:f4241e7d194f72405b393b9e8a5635711b60da834a0b2114724d1160fc9319f6

Observation e9fda900-ecda-4f1b-ada9-2b811b5cd39e · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 121

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:47.042635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:dc72f96160b157774ac5821a60693c0b3df3e4ea180a9c81a395895e73bdcb28

Observation 1df6662f-2293-49e3-8999-1b969d70c1be · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 121

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:a6690f849b433cfb8bb3b4fbc294e23fe00ac270d035c0d1e9ef1171d6d04207

Observation b7ef4b74-8caa-47b0-89ab-70c3e3e99175 · inbound

READER: Robust Evidence-based Authorship Decoding via Extracted Representations cites this paper.

READER: Robust Evidence-based Authorship Decoding via Extracted Representations Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:27:39.624645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-22T06:31:00.163083+00:00.

source=pdf_text observed=2026-06-27T13:19:48.681128Z digest=sha256:01cbbde4e337393997eb849e029b6d11f009483dbcd000341ded6ce6d1ae6544