Pith. sign in

Paper Citation Record · LEDGER

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

As of 22 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2607.07316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07316 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T14:58:58.363330Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact29
  • verified fuzzy21
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a8a6d68-bfe8-4406-9685-7ed44b353349 · outbound

This paper cites Unboxing the black box: Mechanistic interpretability for algorithmic understanding of neural networks,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Unboxing the black box: Mechanistic interpretability for algorithmic understanding of neural networks,

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.059376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:7c8cf7c7269031331f5a95d77d232a1f10f200e4840284af33b680ca7c8f8ab5

Observation 335bcf68-0073-4ba1-ab13-2998315e7442 · outbound

This paper cites arXiv preprint arXiv:2602.11180 , year=.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning arXiv preprint arXiv:2602.11180 , year=

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.054286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:bc57748c825590d372b289cbcb5a70cdb20c9c124357349b56c0bf973d4481c2

Observation f799cb1e-c1e7-4365-b8c7-96e9e3dcabe0 · outbound

This paper cites A mathematical framework for transformer circuits,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning A mathematical framework for transformer circuits,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.959455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:1f565be236d492182456e19462072ae7190678be09dd8b9e700f89a671a2b8a1

Observation 91ef92df-e4ff-4abe-a16e-1f4a50db67f5 · outbound

This paper cites What Makes for a Good Saliency Map? Comparing Strategies for Evaluating Saliency Maps in Explainable AI (XAI).

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning What Makes for a Good Saliency Map? Comparing Strategies for Evaluating Saliency Maps in Explainable AI (XAI)

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:17.995144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:832f3a787546ccb5313385abbbbf20e4485f5c1ffdbb1368b035a89d7242580f

Observation 81ffb238-3df7-4477-be95-8b9594f07424 · outbound

This paper cites A perspective on explainable artificial intelligence methods: Shap and lime,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning A perspective on explainable artificial intelligence methods: Shap and lime,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.953594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e9009afaf776c6807688baa6d3e15ec46eab0859559336952ad53e33e319555a

Observation f6e53894-027e-4bc5-b666-56a00db6f6c8 · outbound

This paper cites Walker, Christos Bergeles, Kai Xu, and Dragos A.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Walker, Christos Bergeles, Kai Xu, and Dragos A

Reference 6

Resolution
malformed identifier
doi_truncated, observed 2026-07-09T15:06:17.890477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c497a44b32bd507d009241fb7545a0027a70c5bdfb830d98d9087e6f396de668

Observation 3f44b898-5aec-45be-b593-d32a93f74ed2 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Understanding intermediate layers using linear classifier probes

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.075143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:a8f037257863a86789bc5ce29e4a153b69c8da744c5a0800013167872cd651bd

Observation 2223e21b-8d10-4db0-a66d-e6e6c3710765 · outbound

This paper cites Visualizing Attention in Transformer-Based Language Representation Models.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Visualizing Attention in Transformer-Based Language Representation Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.063401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:7adfc4633b97e8f028a026facf7a7308a57bc700515fb6d6a1068faaec90d70e

Observation 7d54d782-7e17-466b-9d7a-fe8e82495fd3 · outbound

This paper cites Estimating the attributable cost of physician burnout in the United States.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Estimating the attributable cost of physician burnout in the United States

Reference 9

Resolution
malformed identifier
doi_truncated, observed 2026-07-09T15:06:17.894409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:8cbc6924d5fe93c9f273d84a15484835efaf7b924d229818500ea51e4fb882e4

Observation 63e459a5-fdc2-47d6-a1b2-3e619a87c3e1 · outbound

This paper cites Scoping studies: towards a methodological framework,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Scoping studies: towards a methodological framework,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.949711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:dd429a2d3018e5c0f7396e5dc902b9ccae52e2cab9cb5e441b27e36d3d2d3860

Observation 133eccda-181d-4f01-ae74-3dc121766737 · outbound

This paper cites Circuit tracing: Revealing compu- tational graphs in language models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Circuit tracing: Revealing compu- tational graphs in language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.951469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c344efef6702703fd362ba3dfd629b43729a154adc2b0b8a07008f5bd07b4a96

Observation 14bd7dad-d5fd-480c-9cea-ef3020aba4bc · outbound

This paper cites Open Problems in Mechanistic Interpretability.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Open Problems in Mechanistic Interpretability

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.032089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c75e2ba60cdeb698de4511bfbbc435f5d49fea8c79cda7c365a4da48ad8efe19

Observation a9f87277-4a1b-4803-a749-0ca9dfb11091 · outbound

This paper cites Transformerlens,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Transformerlens,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.963389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:1a309681d38cc06c951913cb177eb705eea0acba54b4dbaacd1e9d55dbaa6756

Observation f50ba8a5-68c7-4a34-b109-163bc5f97de4 · outbound

This paper cites On the geometry and topology of representations: the manifolds of modular addition,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning On the geometry and topology of representations: the manifolds of modular addition,

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.082796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:4d0f5535631c42bfcc90e70580c2763035f4f41bac535e6b61a7c28f6535e7fc

Observation d43f212d-0589-46af-8853-0a3dde432882 · outbound

This paper cites What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.093603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:9715c3c780c4f12317cd5ecc942b9ba548105a0e5ad10b2f9e8673950f61fbbd

Observation 7b91a292-631d-441c-b89c-fdddee3ead41 · outbound

This paper cites In-context Learning and Induction Heads.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning In-context Learning and Induction Heads

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.090840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c9d030417faf5b0c94dc05d1093e731f4486a62d591b140b542ee4f4ec166f12

Observation b329d57e-76b8-42b5-9bfb-857c5baae12b · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.040273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:ed9106259de33665344e1ca9ef98e3363f394ed1b70b0b19aed9b266a7337fa2

Observation e5e570f7-f9a1-4573-b3be-e09adfb38e8b · outbound

This paper cites Investigating the Indirect Object Identification circuit in Mamba.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Investigating the Indirect Object Identification circuit in Mamba

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.085288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:0214a7ae2ae212ef16e337147a2cbde5c7d83a10936aa0b97c79440c474168d2

Observation b606e068-68d0-4af2-8390-d79275e856b1 · outbound

This paper cites Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.080095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:15088190ac04935d36544dc0bdd1613b05abb17ec8ad9aaba90ac878812e4ceb

Observation 44d33f06-a2b5-4278-847a-5d8abde742ba · outbound

This paper cites Towards Automated Circuit Discovery for Mechanistic Interpretability.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Towards Automated Circuit Discovery for Mechanistic Interpretability

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.046309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c098e1fa7ee52dcb474b31c2e026d1eefe3d8d60c0c92da38a7af5d0878608df

Observation 9d3d5443-63d6-4341-a8a0-04feb2f5101d · outbound

This paper cites Efficient automated circuit discovery in transformers using contextual decomposition,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Efficient automated circuit discovery in transformers using contextual decomposition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.975300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:dc5def00519a2fb74c330e8d7dfc94a67747336083d928a29f5adfcc41fc76f9

Observation 2883b690-0e48-42d1-bc40-8d59d882bfe8 · outbound

This paper cites Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.081558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:a6c83bcb52cde7c74c65aa7343ab7d12e0631236d723422c30ffc760f0f5fdc5

Observation f70448e4-0fe1-49ba-a2dd-6ffc4785b611 · outbound

This paper cites Pahq: Accelerating automated circuit discovery through mixed-precision inference optimization,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Pahq: Accelerating automated circuit discovery through mixed-precision inference optimization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.977689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:ecf2ad0d665e6af0e8a9bce1f939aa731d12fc4dccbe6414050794d4b5344cf3

Observation a4a04406-1fa7-43c5-b351-448841c485c0 · outbound

This paper cites Available: https://arxiv.org/abs/2510.23264 7.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Available: https://arxiv.org/abs/2510.23264 7

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.096417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:b10566ad6d1475997c84dcba0c94ac6f33eeb936a06e09c2d52efacdcd615289

Observation ca41d87d-bee1-4123-9402-a5533041be43 · outbound

This paper cites Reinforcement learning fine-tuning enhances activation intensity and diversity in the internal circuitry of llms,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Reinforcement learning fine-tuning enhances activation intensity and diversity in the internal circuitry of llms,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.043477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:d9f8191fd4b9c62eb34696b8b5de0fa0b63b763cf4418cf68340e1020fc6c422

Observation ac80940c-7da5-44f4-812f-40177e46e9ed · outbound

This paper cites Information flow routes: Automatically interpreting language models at scale,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Information flow routes: Automatically interpreting language models at scale,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.980416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:d1674db2c7e3c67a67e4ae1a4d7c36168684f489020e3f191597b1a90a40cc52

Observation e3c36992-318a-4a23-a5a6-305eff92a5d2 · outbound

This paper cites A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:17.992693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:1e908f1aab380d9b89622b7c3b2eb48e3ab23f2bbaba4f75601fc27d631b5a1a

Observation 64692f95-270b-4d37-9075-c10c0e586235 · outbound

This paper cites Emergent symbolic mechanisms support abstract reasoning in large language models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Emergent symbolic mechanisms support abstract reasoning in large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.993990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:cec29ea70edb58e8e4fe05b80e4708a5c9580cd78084a1637ecb25e7bdc9e9f0

Observation 9bcdee38-5b1e-4fad-80dc-ed56c40c4eea · outbound

This paper cites Improving Sparse Autoencoder with Dynamic Attention.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Improving Sparse Autoencoder with Dynamic Attention

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.009978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:42b150b2232c699a83decb2c235fc52f8e9bae733cfcab66ba3800907fc70399

Observation a3808a72-2861-4e9f-8d45-01916f5cdf70 · outbound

This paper cites Mechanistic interpretability with sparse autoencoder neural operators,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Mechanistic interpretability with sparse autoencoder neural operators,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.001846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:29eaa89aca29e939915de44a05a9b009ee43cd6fbb3aeff8495124427c4b6cd6

Observation 2f2cf6b8-dff7-433f-aac9-8e9fd179ccd2 · outbound

This paper cites Mechanistic Interpretability with Sparse Autoencoder Neural Operators.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Mechanistic Interpretability with Sparse Autoencoder Neural Operators

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.077625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:1314910c2b111f329db8e5e8e216608cd3726fd06d8026b8c6b9e87c0c4cee76

Observation d28f71f1-9dcb-4d8f-8fb2-44acf8d5cd49 · outbound

This paper cites Toy Models of Superposition.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Toy Models of Superposition

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.072279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:d01bef728b7405dd7dbf168327a6169869bb4dd9089d0d4bea3502fe8be1905b

Observation 7d6f4be9-dff7-4329-b336-b34f5fc6d09c · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Towards monosemanticity: Decomposing language models with dictionary learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.970387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:8d2b9f7a90d1269d440751664438a9a4dd72f9ae6bb5db4eb03e537c1293011a

Observation f5b2eb1b-1fb1-44ad-950b-e4e237767793 · outbound

This paper cites Finding Neurons in a Haystack: Case Studies with Sparse Probing.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Finding Neurons in a Haystack: Case Studies with Sparse Probing

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.076196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e0779abb3ba746e3c2a908ef5550393fdfa8d994abbafe6bd5be891705775540

Observation ffcb526e-c0fe-4bb6-9880-66796529f848 · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.972674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:54ef54b12ff593045b104c29e92d2d3db9d2f899495dfc035186bb87012fafc2

Observation 7ea73d2a-9bc0-4859-993c-1d5d69319310 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.005735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:229cc4d5cd4ec3f4ab6f907fc38f1d6b73a18f2287cebe557289b90c9beac467

Observation 30c866d9-141b-4470-a276-ba5f721bdf87 · outbound

This paper cites arXiv preprint arXiv:2503.05613 , year=.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning arXiv preprint arXiv:2503.05613 , year=

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.088269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:fa6585d777daef593a90e8acac58437228c20886fd489a38546f2f326bc1a73b

Observation 8db91b84-83df-4784-a41f-ca281954b333 · outbound

This paper cites Improving Dictionary Learning with Gated Sparse Autoencoders.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Improving Dictionary Learning with Gated Sparse Autoencoders

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.019153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:979461df0a687907eb45e1d69dc7539343b7498419e39d268ecd119aa1a0d74b

Observation 235e1987-b44e-4e83-b0fb-de12c9d54521 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Scaling and evaluating sparse autoencoders

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.086817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:08e75238564b7407903ecd7ff16e39680dd75b62b8ffa646e75b3f849a320bb0

Observation e8fddea9-8b35-4737-a383-eada0b901c16 · outbound

This paper cites BatchTopK Sparse Autoencoders.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning BatchTopK Sparse Autoencoders

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.048729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c26b37eb0887fb4fecc0b45b3d09f2e93c317a65273318fcc09208777eaede16

Observation 077a9b76-8336-44f2-9a37-1fc4b87c3bad · outbound

This paper cites Audenaert, K.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Audenaert, K

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.028400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:1be127a20c02a639a846dd7fe1a14de9e3e024546af239299d0937371d72f2af

Observation 53703399-7927-475c-99bc-6b63e14c41e3 · outbound

This paper cites Activation oracles: Training and evaluating llms as general-purpose activation explainers,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Activation oracles: Training and evaluating llms as general-purpose activation explainers,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.003836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:4910b174a70d658422d6d8bab762bde71470e652a1acdeb8009bd553b9ac9d3a

Observation 1fd9a259-9392-4d76-a039-b745e8456d07 · outbound

This paper cites Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.056758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:88dc4fff021cf943e55dd7ea9da6cff25c403aca7959096da634b8daa7d485b9

Observation 142f55d9-02e5-485e-a854-5890cf86d591 · outbound

This paper cites Sparse feature circuits: Discovering and editing interpretable causal graphs in language models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Sparse feature circuits: Discovering and editing interpretable causal graphs in language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.006000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:f128a1bd3f6fa622b1d181496857349a1c278a2bcb2b02893bc2696cb85fafbd

Observation 592881ed-6b90-4ad2-ada9-af85ad615057 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:17.997850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:3a93cc7da0befd75126b47cf8eafac17a21330904a6dd53309488c3b6b0dfe03

Observation c53ab872-313b-4f33-803c-bd9ff9ba7b49 · outbound

This paper cites Evaluating explanation faithfulness in toy models,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Evaluating explanation faithfulness in toy models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.007862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:4a77bc5d2cb2cdadc21086aa3b37cdf79fddb1cffcee0f4ff8a8339b8f234605

Observation 70c3af1d-2434-48eb-bc8a-6922bae50f2c · outbound

This paper cites Transcoders Find Interpretable LLM Feature Circuits.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Transcoders Find Interpretable LLM Feature Circuits

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.057814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e3bb54c50ce9cae45395f0382a8eeba51bf4c17362d8419ba72911cd3c7a792e

Observation 7e639826-0ce8-46bb-8c5a-55e0af2be5ec · outbound

This paper cites Circuit insights: Towards inter- pretability beyond activations,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Circuit insights: Towards inter- pretability beyond activations,

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.079064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:00b9afe1e97d881435c50ca4c5f80b2166e5007f35e840c2e05ae56e4d162794

Observation 9b6031d9-f633-4147-83a1-b963273e6dff · outbound

This paper cites Viacheslav Sinii, Nikita Balagansky, Gleb Gerasimov, Daniil Laptev, Yaroslav Aksenov, Vadim Kurochkin, Alexey Gorbatovski, Boris Shaposhnikov, and Daniil Gavrilov.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Viacheslav Sinii, Nikita Balagansky, Gleb Gerasimov, Daniil Laptev, Yaroslav Aksenov, Vadim Kurochkin, Alexey Gorbatovski, Boris Shaposhnikov, and Daniil Gavrilov

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T15:06:18.034965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:392f15b4c36dd798c7d3812f9bdb9655cf49c5a5002f091508102b6c7d2d8513

Observation b257c578-4d8b-4dab-ac70-1eeae03666d3 · outbound

This paper cites Model whisper: Steering vectors unlock large language models’ potential in test- time,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Model whisper: Steering vectors unlock large language models’ potential in test- time,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.999781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:9b482eecb9b198ec501ca4ac551a3ffa3eb4481f32c5647321ba4a69747b743a

Observation 8d6fc394-2615-41fe-afdb-9dc0ff8bb575 · outbound

This paper cites Toward a Flexible Framework for Linear Representation Hypothesis Using Maximum Likelihood Estimation.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Toward a Flexible Framework for Linear Representation Hypothesis Using Maximum Likelihood Estimation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.000935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:0f1a035598e1562dc14bb940db9de35215270a8f2ad404b5ce371447f46253c2

Observation 6c51d80d-f7d2-4adf-b113-02c320bf61b9 · outbound

This paper cites Representation Engineering for Large-Language Models: Survey and Research Challenges.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Representation Engineering for Large-Language Models: Survey and Research Challenges

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.061917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:f66125400b4e2adf7a4aae701cb73f40b169b30b4fe767ee660910520ac1d12c

Observation 9306888a-afae-4172-baed-357ce49becde · outbound

This paper cites Steering Language Models With Activation Engineering.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Steering Language Models With Activation Engineering

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.064625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:240e574a3eb6141762d80cce7a10e7972a16a77e0be4b0e3e0439b2f3c0fa240

Observation 798d7223-9156-4c23-aab9-b87ae5d1a9cc · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Steering Llama 2 via Contrastive Activation Addition

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.071191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:192f19a1f5cdf0aa6f9585e90739ee9836161d7804927ce6ca8583e55805069f

Observation 1d3be7df-a21f-4b83-8e1a-6bac0e8e64e8 · outbound

This paper cites Emotion concepts and their function in a large language model,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Emotion concepts and their function in a large language model,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:19.009978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:6e52dc4a68e1b38b683f82fcd15273c2e17ea5631831f010d720e52df0ea09f1

Observation f33bc7f5-1f00-42d3-bfaf-b0c109102fe9 · outbound

This paper cites Interpretable Steering of Large Language Models with Feature Guided Activation Additions.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Interpretable Steering of Large Language Models with Feature Guided Activation Additions

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.052467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:62bd54b62b4dd0e914073fb7665b6fb3de482cb68a5ddb204822984ff6aab2bf

Observation 7e71e554-704c-43c3-ad31-04c5b6c3f812 · outbound

This paper cites Dspa: Dynamic sae steering for data-efficient preference alignment,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Dspa: Dynamic sae steering for data-efficient preference alignment,

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-09T15:06:18.060488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:d0f176b5074c72655c7f3469190189bcc8aa7450f0fb600e0f89708fa24c09af

Observation 600aa25f-de6e-493b-b00a-67fc8b8c2852 · outbound

This paper cites Fold-se: An efficient rule-based machine learning algorithm with scalable explainability,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Fold-se: An efficient rule-based machine learning algorithm with scalable explainability,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.991860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:f8bff955d447bbe945dcebce54aaf0700dd6aef0cfc53930bde2ac9684dd9e63

Observation 1b470a6b-a012-4abd-b922-09b8af20cea8 · outbound

This paper cites FOLD-SE: An Efficient Rule-based Machine Learning Algorithm with Scalable Explainability.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning FOLD-SE: An Efficient Rule-based Machine Learning Algorithm with Scalable Explainability

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.069562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:8450bfb7f3f7274d59c06bfc113e56335e0e75669d49f83450c7de427971e509

Observation a0950a2e-d21d-4f96-88f7-4358eef3d4cf · outbound

This paper cites Nesyfold: Neurosym- bolic framework for interpretable image classification,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Nesyfold: Neurosym- bolic framework for interpretable image classification,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.997440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:9bbfbb6236bb30ada2086b32743c9e4fb53e28bf15949ae396b6cd623be860c9

Observation 72c5a146-39e8-49b7-bed7-8fd5ec05e676 · outbound

This paper cites NeSyFOLD: Neurosymbolic Framework for Interpretable Image Classification.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning NeSyFOLD: Neurosymbolic Framework for Interpretable Image Classification

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T15:06:18.073728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e615d323c3a5bc9c66b60d76ff9306801a480c28d242847fe54cca546f052042

Observation 706718ba-0b62-490b-8cd4-66c287373325 · outbound

This paper cites Theory and Practice of Logic Programming 25(4), 722–738 (2025).

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Theory and Practice of Logic Programming 25(4), 722–738 (2025)

Reference 62

Resolution
metadata mismatch
doi, observed 2026-07-09T15:06:17.896973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:c6147d752865b7c618084409414d6d375c346de4640a2311c43b7d9aca23bc9c

Observation 43abf867-60dd-4f7a-b873-3fdf51929ebd · outbound

This paper cites Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.988371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:655bd9c1105b380b789e16e9be511a9b4538b553a94add54b9525acd0dc0c56e

Observation 3c352f67-2928-490d-8700-39183cbbbcdb · outbound

This paper cites Llms can’t plan, but can help planning in llm-modulo frame- works,.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Llms can’t plan, but can help planning in llm-modulo frame- works,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T15:06:18.984081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:e135ce958064f9a6b43004adc0b1782dd257597d012240c4f09b1ab6ea96ee69

Observation 0350b262-9dc1-4971-a7fb-b351476dfebc · outbound

This paper cites Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.019911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:cd3a811c30d52c54136d2e5588433aa4415e0609b552d7c29dbd44866e971db4

Pith citing papers

No inbound Pith citation observations are available.