Pith. sign in

Paper Citation Record · LEDGER

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

As of 11 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2607.07316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07316 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T14:58:58.363330Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact29
  • verified fuzzy21
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a8a6d68-bfe8-4406-9685-7ed44b353349 · outbound

This paper cites Unboxing the black box: Mechanistic interpretability for algorithmic understanding of neural networks,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Unboxing the black box: Mechanistic interpretability for algorithmic understanding of neural networks,

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.059376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:446b0bf9a7b14e016a437cc2cdea734d10c3306cd36b5dcb6ec5b69a0ccc3879

Observation 335bcf68-0073-4ba1-ab13-2998315e7442 · outbound

This paper cites arXiv preprint arXiv:2602.11180 , year=.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning arXiv preprint arXiv:2602.11180 , year=

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.054286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:cbdce793085ff3f0c7ed3fa9705b8cd4ef4989ba7243e3436967b08b996549df

Observation f799cb1e-c1e7-4365-b8c7-96e9e3dcabe0 · outbound

This paper cites A mathematical framework for transformer circuits,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning A mathematical framework for transformer circuits,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.959455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:2123dbeed25c46cfb5114461f9f7d5ccd58f7ce44aabc25591e9142550e0c147

Observation 91ef92df-e4ff-4abe-a16e-1f4a50db67f5 · outbound

This paper cites What Makes for a Good Saliency Map? Comparing Strategies for Evaluating Saliency Maps in Explainable AI (XAI).

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning What Makes for a Good Saliency Map? Comparing Strategies for Evaluating Saliency Maps in Explainable AI (XAI)

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:17.995144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:daed77e1bba0e140accdbe2d5f7d6f369860d49a04f0ae54b8e3f8f2da210221

Observation 81ffb238-3df7-4477-be95-8b9594f07424 · outbound

This paper cites A perspective on explainable artificial intelligence methods: Shap and lime,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning A perspective on explainable artificial intelligence methods: Shap and lime,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.953594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e1274e97aec41c1f5c3b6fc3038cf8d210d85e9574060f763dfb84b668fbaad2

Observation f6e53894-027e-4bc5-b666-56a00db6f6c8 · outbound

This paper cites Walker, Christos Bergeles, Kai Xu, and Dragos A.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Walker, Christos Bergeles, Kai Xu, and Dragos A

Reference 6

Resolution
malformed identifier
doi_truncated, observed 2026-07-09T15:06:17.890477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:0493d99abe2a975ea3b804e990f526c4d14f40b09b30f3a7e1c38f28d47a7703

Observation 3f44b898-5aec-45be-b593-d32a93f74ed2 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Understanding intermediate layers using linear classifier probes

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.075143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e39eb4fe8e0c02509c2db4f969d3b2ad70f8036f2c52fb550525ead87234cad4

Observation 2223e21b-8d10-4db0-a66d-e6e6c3710765 · outbound

This paper cites Visualizing Attention in Transformer-Based Language Representation Models.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Visualizing Attention in Transformer-Based Language Representation Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.063401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:4b445e9fdd48f89e3866073423072d2d689f5f09492cf5185ff63d5d3ea93a20

Observation 7d54d782-7e17-466b-9d7a-fe8e82495fd3 · outbound

This paper cites Estimating the attributable cost of physician burnout in the United States.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Estimating the attributable cost of physician burnout in the United States

Reference 9

Resolution
malformed identifier
doi_truncated, observed 2026-07-09T15:06:17.894409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:7b5c575314c330b0871177f873d290dc4b65e67cf70e29e62d35e8f13737a433

Observation 63e459a5-fdc2-47d6-a1b2-3e619a87c3e1 · outbound

This paper cites Scoping studies: towards a methodological framework,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Scoping studies: towards a methodological framework,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.949711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:2f921230a32d9487d4c5ff88fe1711805b4b0d18673b04072d9fa28c49e61fc0

Observation 133eccda-181d-4f01-ae74-3dc121766737 · outbound

This paper cites Circuit tracing: Revealing compu- tational graphs in language models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Circuit tracing: Revealing compu- tational graphs in language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.951469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:85cfa51970161d1c0189804f45c134cd89455adcf7635b1c48c3d105baf420f9

Observation 14bd7dad-d5fd-480c-9cea-ef3020aba4bc · outbound

This paper cites Open Problems in Mechanistic Interpretability.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Open Problems in Mechanistic Interpretability

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.032089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:3671eeadf280f006aa660d712995bf3637716fa4a3698c2c4c1f1a7d3453eca5

Observation a9f87277-4a1b-4803-a749-0ca9dfb11091 · outbound

This paper cites Transformerlens,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Transformerlens,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.963389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:ef5c0b2eebe6d58c9d670895adc91e2339a46c57a77ba396dc48205ca53df10b

Observation f50ba8a5-68c7-4a34-b109-163bc5f97de4 · outbound

This paper cites On the geometry and topology of representations: the manifolds of modular addition,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning On the geometry and topology of representations: the manifolds of modular addition,

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.082796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:7e73a9fefd79e4bf54ea728df0f8a3b559ee964b9ea4bd86dee37145c61afb81

Observation d43f212d-0589-46af-8853-0a3dde432882 · outbound

This paper cites What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.093603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e80037385c7c491fd1e12bcbbcee4bdcdb2b2ddb47eef83f2e0fbc3819e1d2c2

Observation 7b91a292-631d-441c-b89c-fdddee3ead41 · outbound

This paper cites In-context Learning and Induction Heads.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning In-context Learning and Induction Heads

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.090840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:71345db9a9455a8c4843640cfdfa22fcd115d8fc30cbce999c2f8838f47c32b7

Observation b329d57e-76b8-42b5-9bfb-857c5baae12b · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.040273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:d78ce514504f3faba9782726ff9e4a754254463add4a96e223cb37119897a800

Observation e5e570f7-f9a1-4573-b3be-e09adfb38e8b · outbound

This paper cites Investigating the Indirect Object Identification circuit in Mamba.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Investigating the Indirect Object Identification circuit in Mamba

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.085288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:015045a872bf9678c6fe9069d3b88502e95e8c6035d49a1d3214a77922eee68c

Observation b606e068-68d0-4af2-8390-d79275e856b1 · outbound

This paper cites Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.080095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:a3085204f3c4062102336f84a90416c150568da7a95d6c9a3a174020d0a75d4c

Observation 44d33f06-a2b5-4278-847a-5d8abde742ba · outbound

This paper cites Towards Automated Circuit Discovery for Mechanistic Interpretability.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Towards Automated Circuit Discovery for Mechanistic Interpretability

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.046309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:124c753c31124b587017de640a879d4db3b489a3ab15c08291ecc8efd218ed74

Observation 9d3d5443-63d6-4341-a8a0-04feb2f5101d · outbound

This paper cites Efficient automated circuit discovery in transformers using contextual decomposition,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Efficient automated circuit discovery in transformers using contextual decomposition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.975300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:99284a46c3152ea2eb3dbb9f3179bfb702028b6a7e63b2cb4c7afc7e9b6904b4

Observation 2883b690-0e48-42d1-bc40-8d59d882bfe8 · outbound

This paper cites Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.081558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:f1aa716d4f67d07fde7fff615079f4348e422e4a78b7516d735bc308a6af0be6

Observation f70448e4-0fe1-49ba-a2dd-6ffc4785b611 · outbound

This paper cites Pahq: Accelerating automated circuit discovery through mixed-precision inference optimization,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Pahq: Accelerating automated circuit discovery through mixed-precision inference optimization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.977689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:404320c17e3c854a51b68b7133c7b1296a74f2a88b08eb9e48827afd3b653e3a

Observation a4a04406-1fa7-43c5-b351-448841c485c0 · outbound

This paper cites Available: https://arxiv.org/abs/2510.23264 7.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Available: https://arxiv.org/abs/2510.23264 7

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.096417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:8eb18b04bdc3f2f031a7eb9ca59b07d5e4728297df9fcaf180603b64199f998f

Observation ca41d87d-bee1-4123-9402-a5533041be43 · outbound

This paper cites Reinforcement learning fine-tuning enhances activation intensity and diversity in the internal circuitry of llms,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Reinforcement learning fine-tuning enhances activation intensity and diversity in the internal circuitry of llms,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.043477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:103d0743b44f9635a05c0ee2fb3be848e866153c33c34aa25ef636d16087a476

Observation ac80940c-7da5-44f4-812f-40177e46e9ed · outbound

This paper cites Information flow routes: Automatically interpreting language models at scale,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Information flow routes: Automatically interpreting language models at scale,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.980416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c695b505d29f824e3593afd242b117e823130cd31e1911456862f2e056baf94f

Observation e3c36992-318a-4a23-a5a6-305eff92a5d2 · outbound

This paper cites A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:17.992693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:6d2cdb71c9139621bbcabfaaf558fab32851af14ca6b7ad94bd821eb04674004

Observation 64692f95-270b-4d37-9075-c10c0e586235 · outbound

This paper cites Emergent symbolic mechanisms support abstract reasoning in large language models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Emergent symbolic mechanisms support abstract reasoning in large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.993990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:05e00e35955c5ed0c111098069ceba9b6d845eb429ee975e52eeba8e35f7d4d0

Observation 9bcdee38-5b1e-4fad-80dc-ed56c40c4eea · outbound

This paper cites Improving Sparse Autoencoder with Dynamic Attention.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Improving Sparse Autoencoder with Dynamic Attention

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.009978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:174d0ba944da964205ce722b90af8e1105287d7eaeec4fa788e9c6daa404c01b

Observation a3808a72-2861-4e9f-8d45-01916f5cdf70 · outbound

This paper cites Mechanistic interpretability with sparse autoencoder neural operators,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Mechanistic interpretability with sparse autoencoder neural operators,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.001846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:33f5d33da2522315dc54b444e19f608c36521e6ca301697acce86813a468f8d7

Observation 2f2cf6b8-dff7-433f-aac9-8e9fd179ccd2 · outbound

This paper cites Mechanistic Interpretability with Sparse Autoencoder Neural Operators.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Mechanistic Interpretability with Sparse Autoencoder Neural Operators

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.077625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:ae276492d077aad51588960347375086be1031662a3d5fd59028a38bedbb1a37

Observation d28f71f1-9dcb-4d8f-8fb2-44acf8d5cd49 · outbound

This paper cites Toy Models of Superposition.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Toy Models of Superposition

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.072279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:bbeec0e9cbd696a929a0c7d9acf13081629591d92a488643f4ba544b40e66642

Observation 7d6f4be9-dff7-4329-b336-b34f5fc6d09c · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Towards monosemanticity: Decomposing language models with dictionary learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.970387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:581cef5a87d97523cf0725824172280edb2d5f98642e42c4c921792ccaa528ba

Observation f5b2eb1b-1fb1-44ad-950b-e4e237767793 · outbound

This paper cites Finding Neurons in a Haystack: Case Studies with Sparse Probing.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Finding Neurons in a Haystack: Case Studies with Sparse Probing

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.076196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:8c6e8d3f0d4e4ab58075dd5b36fe1d3460667e6bcbaa0876bcf87a1c5a6a78a1

Observation ffcb526e-c0fe-4bb6-9880-66796529f848 · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.972674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:a0f600b0f791fc96738d7d685dbc7d05f8887bd4aceaddd5525766c5679850c1

Observation 7ea73d2a-9bc0-4859-993c-1d5d69319310 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.005735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:4f6c410535d739d1e49b96dd97cab017e2dc32334acfd683ff22b332bfef552c

Observation 30c866d9-141b-4470-a276-ba5f721bdf87 · outbound

This paper cites arXiv preprint arXiv:2503.05613 , year=.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning arXiv preprint arXiv:2503.05613 , year=

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.088269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:29f055aaaae05003524d0bdd27ba5b983da6521c4a37613a4986d2023833753d

Observation 8db91b84-83df-4784-a41f-ca281954b333 · outbound

This paper cites Improving Dictionary Learning with Gated Sparse Autoencoders.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Improving Dictionary Learning with Gated Sparse Autoencoders

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.019153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:19ef3e0e9dc56414c45914b8de5abce95f257cf54f90eaf478cb8a939d07cdc2

Observation 235e1987-b44e-4e83-b0fb-de12c9d54521 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Scaling and evaluating sparse autoencoders

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.086817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:8ad38ae42d8d07883fb525514a5d8f6861e521dc7793f0441c9df64796cd104b

Observation e8fddea9-8b35-4737-a383-eada0b901c16 · outbound

This paper cites BatchTopK Sparse Autoencoders.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning BatchTopK Sparse Autoencoders

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.048729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:540621cae22dbbe4f9a0ec7fc07d8584ac2bdd7be795e68dabac7b27597a3889

Observation 077a9b76-8336-44f2-9a37-1fc4b87c3bad · outbound

This paper cites Audenaert, K.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Audenaert, K

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.028400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:aa7f70120216eebb9ccf47133483121972bb80ca35b87989ce0a88c059afe4ce

Observation 53703399-7927-475c-99bc-6b63e14c41e3 · outbound

This paper cites Activation oracles: Training and evaluating llms as general-purpose activation explainers,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Activation oracles: Training and evaluating llms as general-purpose activation explainers,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.003836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:690c14038ebda5994cff971bbfc6896f050392799213968658f271275ffc3884

Observation 1fd9a259-9392-4d76-a039-b745e8456d07 · outbound

This paper cites Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.056758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:f8170244d2dd98cad0a14ef7860ea01a94a0ae581811d799ed117ccc364e3cb3

Observation 142f55d9-02e5-485e-a854-5890cf86d591 · outbound

This paper cites Sparse feature circuits: Discovering and editing interpretable causal graphs in language models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Sparse feature circuits: Discovering and editing interpretable causal graphs in language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.006000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e03f4937c0907520e186efc57fd6eb88004b9dc93dddb5d5b1147848588db51c

Observation 592881ed-6b90-4ad2-ada9-af85ad615057 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:17.997850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:3dc9aebda88e3fbd3273b15acaae8b7720b1ff67359c5ae6626ab488d19f0d31

Observation c53ab872-313b-4f33-803c-bd9ff9ba7b49 · outbound

This paper cites Evaluating explanation faithfulness in toy models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Evaluating explanation faithfulness in toy models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.007862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:243e1bd21577084e8da4b8c3b819d8acada1ee9e772018c04d25a82033c15f53

Observation 70c3af1d-2434-48eb-bc8a-6922bae50f2c · outbound

This paper cites Transcoders Find Interpretable LLM Feature Circuits.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Transcoders Find Interpretable LLM Feature Circuits

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.057814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:1ee7c66daf100dfb19c7720eeb5e68e05fc385d46f6a52ff7814133267825b1d

Observation 7e639826-0ce8-46bb-8c5a-55e0af2be5ec · outbound

This paper cites Circuit insights: Towards inter- pretability beyond activations,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Circuit insights: Towards inter- pretability beyond activations,

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.079064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:b587320f59fad7f54c8586908fcf94acf627373592e02a34b6846d26655110cf

Observation 9b6031d9-f633-4147-83a1-b963273e6dff · outbound

This paper cites Viacheslav Sinii, Nikita Balagansky, Gleb Gerasimov, Daniil Laptev, Yaroslav Aksenov, Vadim Kurochkin, Alexey Gorbatovski, Boris Shaposhnikov, and Daniil Gavrilov.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Viacheslav Sinii, Nikita Balagansky, Gleb Gerasimov, Daniil Laptev, Yaroslav Aksenov, Vadim Kurochkin, Alexey Gorbatovski, Boris Shaposhnikov, and Daniil Gavrilov

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.034965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:8cd9a38466ca618886f3333510838f059bf589d92a8c4516f8d1023151f4f84f

Observation b257c578-4d8b-4dab-ac70-1eeae03666d3 · outbound

This paper cites Model whisper: Steering vectors unlock large language models’ potential in test- time,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Model whisper: Steering vectors unlock large language models’ potential in test- time,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.999781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:9564bccc6e1255f8f1c6805af8efebedb61463902c95d7347eb941a017e135cb

Observation 8d6fc394-2615-41fe-afdb-9dc0ff8bb575 · outbound

This paper cites Toward a Flexible Framework for Linear Representation Hypothesis Using Maximum Likelihood Estimation.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Toward a Flexible Framework for Linear Representation Hypothesis Using Maximum Likelihood Estimation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.000935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:81212a44ac2dad5b5b4cbbb944cc2893316f5c84a59c9feb06a67a9c435a9cd6

Observation 6c51d80d-f7d2-4adf-b113-02c320bf61b9 · outbound

This paper cites Representation Engineering for Large-Language Models: Survey and Research Challenges.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Representation Engineering for Large-Language Models: Survey and Research Challenges

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.061917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:22e01e112d5f726dac8b457a4365accdca93af9992f6c6b86771d3ec5ca7073d

Observation 9306888a-afae-4172-baed-357ce49becde · outbound

This paper cites Steering Language Models With Activation Engineering.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Steering Language Models With Activation Engineering

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.064625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:9ebc4cfdb6f66d054e5ce1e8661ded2db7af46ceba41728b0a282f0e7038e48a

Observation 798d7223-9156-4c23-aab9-b87ae5d1a9cc · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Steering Llama 2 via Contrastive Activation Addition

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.071191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:08eafe153b9dc387f51fa9c6c91c2127385d167c4d1fac444be1ca01e1d76111

Observation 1d3be7df-a21f-4b83-8e1a-6bac0e8e64e8 · outbound

This paper cites Emotion concepts and their function in a large language model,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Emotion concepts and their function in a large language model,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.009978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:b60917c9e85f1665b1b8fb9aa0eb5e929d8caebbbd6bb26b06e83a2833817317

Observation f33bc7f5-1f00-42d3-bfaf-b0c109102fe9 · outbound

This paper cites Interpretable Steering of Large Language Models with Feature Guided Activation Additions.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.052467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:749ff0a655d8c8cc66c2b26dc477c67e4f39107a8f8799e70b5497b72f6c9585

Observation 7e71e554-704c-43c3-ad31-04c5b6c3f812 · outbound

This paper cites Dspa: Dynamic sae steering for data-efficient preference alignment,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Dspa: Dynamic sae steering for data-efficient preference alignment,

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.060488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:319ba497a3e89d883e9d69afcf4307972b202a0f78854e0fefb15c40f77f0378

Observation 600aa25f-de6e-493b-b00a-67fc8b8c2852 · outbound

This paper cites Fold-se: An efficient rule-based machine learning algorithm with scalable explainability,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Fold-se: An efficient rule-based machine learning algorithm with scalable explainability,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.991860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:199170fad9a82b3e5e8bfe23edd11abf3d775d4f5d1dce421ce82832174dedd6

Observation 1b470a6b-a012-4abd-b922-09b8af20cea8 · outbound

This paper cites FOLD-SE: An Efficient Rule-based Machine Learning Algorithm with Scalable Explainability.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning FOLD-SE: An Efficient Rule-based Machine Learning Algorithm with Scalable Explainability

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.069562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:95bd11f79128eba9416edb74b16dafd33e2e5ffb3cad107ca54dea93e78f87bc

Observation a0950a2e-d21d-4f96-88f7-4358eef3d4cf · outbound

This paper cites Nesyfold: Neurosym- bolic framework for interpretable image classification,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Nesyfold: Neurosym- bolic framework for interpretable image classification,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.997440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:dec1aae3a9788014e4fedf02734f51670e17d0fa421a5fdac67e6318a915c545

Observation 72c5a146-39e8-49b7-bed7-8fd5ec05e676 · outbound

This paper cites NeSyFOLD: Neurosymbolic Framework for Interpretable Image Classification.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning NeSyFOLD: Neurosymbolic Framework for Interpretable Image Classification

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.073728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:7b0ce6b40a92b61d1f143417f19d4b93593b4ab486a1ecea5e058e14016ec795

Observation 706718ba-0b62-490b-8cd4-66c287373325 · outbound

This paper cites Theory and Practice of Logic Programming 25(4), 722–738 (2025).

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Theory and Practice of Logic Programming 25(4), 722–738 (2025)

Reference 62

Resolution
metadata mismatch
doi, observed 2026-07-09T15:06:17.896973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:11b0bc198630fc82ddfbb83ec8d4dd4e66d47383ca076f5c575944c794fb3175

Observation 43abf867-60dd-4f7a-b873-3fdf51929ebd · outbound

This paper cites Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.988371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:63c1161a1a6ae2c875fb545c74ae4187870084482e71ea6387fdea6395915bf1

Observation 3c352f67-2928-490d-8700-39183cbbbcdb · outbound

This paper cites Llms can’t plan, but can help planning in llm-modulo frame- works,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Llms can’t plan, but can help planning in llm-modulo frame- works,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.984081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:69e26fc7dc98262abad0a510a5c2702da3d415ad858b9a9a7126ef6b3396445d

Observation 0350b262-9dc1-4971-a7fb-b351476dfebc · outbound

This paper cites Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.019911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:b0900e4312594ddccd54a1a69b073bc92735733b5164e426fe06ef515906d7ac

Pith citing papers

No inbound Pith citation observations are available.