Pith. sign in

Paper Citation Record · LEDGER

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

As of 21 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2505.20254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20254 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:03.819990Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-04T17:01:39.362411Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:09:58.298491Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d985f89e-c3d7-481f-bcb8-a67ba9aeba66 · outbound

This paper cites SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:58.487711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:58.487711Z digest=sha256:69d745b6ee145c450a04f1d1722739c82b507d822c725c3c6f79b9799f11f27f

Observation 84fc0e70-98a5-443c-a3cd-cbfa08c00f8f · outbound

This paper cites an unresolved cited work.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:03:08.646727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:02:58.567689Z digest=sha256:fdecf94817601d0dade801975918a3c160106bfa6e85b9a87395e22933457f6c

Observation 645c8277-9cca-4b46-8479-21efa3b5726c · outbound

This paper cites New algorithms for learning incoherent and overcomplete dictionaries.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs New algorithms for learning incoherent and overcomplete dictionaries

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:58.726604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:58.726604Z digest=sha256:22cca34e1a1413cc7b17d3da6a6c2570ebfe6ad18037f2231557d8a1e6bc7a1f

Observation 34c35eff-2e77-4f2d-a7b9-d27f17667d93 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Pythia: A suite for analyzing large language models across training and scaling

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:08.494837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:02:58.750981Z digest=sha256:dd3b8945e6e83cf91ffb5e5168859dee7bc07e1b6ed4345f126ef7aaaaebe0ca

Observation d388b3e2-e670-49b7-8ca1-79094ac30591 · outbound

This paper cites Turner, Cem Anil, Carson Denison, Amanda Askell, et al.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Turner, Cem Anil, Carson Denison, Amanda Askell, et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:08.359800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:02:58.823522Z digest=sha256:0b049c812601ec1d588da18e2c1b6504f33877b9ffe5d6729cb969b12b869dd2

Observation e0457bd2-f38e-4f23-a0f5-9e841a08dcc5 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Towards monosemanticity: Decomposing language models with dictionary learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:08.219243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:02:58.921893Z digest=sha256:82a50f7079a13af91173691df8b4cbd373a13883cc9b749de820b7576241ca7c

Observation 6a38c3d4-af46-42a5-87b0-f690c11b3c21 · outbound

This paper cites BatchTopK Sparse Autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs BatchTopK Sparse Autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:58.976794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:58.976794Z digest=sha256:7cec44603c98ff7340e385aaafbc9c4d10dd469a0192c89f9c534a66035688dc

Observation 4f48c707-9311-4ae1-a416-49bbd095833e · outbound

This paper cites Learning Multi-Level Features with Matryoshka Sparse Autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Learning Multi-Level Features with Matryoshka Sparse Autoencoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.051713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.051713Z digest=sha256:057d7d92241ec403d7e10c3e8f9c21d9b11cf1344b4418c2e63ad7db8ddedf7c

Observation d04aa17b-e3fd-4de2-a6cf-85e93322aba3 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.154435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.154435Z digest=sha256:84774141875a15d1f4d894fe912c0d663d5f0b478d47836471ca9dcf4f7fb3aa

Observation 1363a7c9-9a08-4aba-8d9a-1e8ee822cd3a · outbound

This paper cites A is for absorption: Studying feature splitting and absorption in sparse autoencoders.arXiv preprint arXiv:2409.14507, 2024.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs A is for absorption: Studying feature splitting and absorption in sparse autoencoders.arXiv preprint arXiv:2409.14507, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.217821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.217821Z digest=sha256:6ff91d54b76e6d955b380179581fc656642f0dcee30c5049506a754c01b325f9

Observation 052532a1-9c86-440b-96ad-92209798469f · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.277009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.277009Z digest=sha256:bd52ea2673f358b3ccbf70d2bdea2c2cd4234b2ca24362a0a6e0459a89e5e9c5

Observation eb97890a-4497-4757-91c6-f1b8787df341 · outbound

This paper cites Optimally sparse representation in general (nonorthogonal) dictionaries via l1 minimization.Proceedings of the National Academy of Sciences, 100(5):2197– 2202, 2003.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Optimally sparse representation in general (nonorthogonal) dictionaries via l1 minimization.Proceedings of the National Academy of Sciences, 100(5):2197– 2202, 2003

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:08.092039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:02:59.383652Z digest=sha256:9611f60a85ada049d2d9279bf173eeb4a5b6581ed33cf27d9a49db4914f4e81f

Observation ea1036f8-040f-4298-a3da-7ae2ee0a1cdf · outbound

This paper cites Toy Models of Superposition.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Toy Models of Superposition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.477693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.477693Z digest=sha256:db36661aa805bb201cfd8a309ca469bd2564f76226e2c99a64ae91854aa7fddb

Observation 6e73a26d-6fa2-4fbd-b629-66a9067e20c9 · outbound

This paper cites A mathematical framework for transformer circuits.Transformer Circuits Thread, 1(1):12, 2021.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs A mathematical framework for transformer circuits.Transformer Circuits Thread, 1(1):12, 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.553114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.553114Z digest=sha256:cf2b0f5a05b6b7f9fe4d2a9db95d745215926a9134b58da59b56a8f07e8e468e

Observation db898c7e-37ed-42a2-aed4-64927b95be1a · outbound

This paper cites Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.630371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.630371Z digest=sha256:74b5cfa9d3408756c19399bc07eed032e1bc259ee907d03700da44dce8cb9030

Observation 17262248-3e93-4e91-99ee-b4c16cfb830b · outbound

This paper cites Scientific inference with interpretable machine learning: Analyzing models to learn about real-world phenomena.Minds and Machines, 34(3):32, 2024.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Scientific inference with interpretable machine learning: Analyzing models to learn about real-world phenomena.Minds and Machines, 34(3):32, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:07.963257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:02:59.700647Z digest=sha256:4817d55fb48ffa135f02883006047a6808467743dfb48490f1dfc7856d081fc0

Observation 9a340735-39cf-4e46-b67a-61471657af26 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.781474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.781474Z digest=sha256:446b8a96997d9c252ffd96cced538d8061eb22bd0450cb758983996a5681ea3f

Observation 44650b81-0291-42d5-9ccd-10c19e3d95ab · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Scaling and evaluating sparse autoencoders

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:59.886828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:59.886828Z digest=sha256:c609fc680abfeec78b5b1865c65ed9ab4e16cbfaa40ed766c548b90f2f5c180d

Observation 16e0a001-c085-469c-9db9-6eaa35db7c86 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Scaling and evaluating sparse autoencoders

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:07.848712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:00.040407Z digest=sha256:7e5d0bf758526935f07d8054b81464403aac20b778c6c94a5e09fe1c4646d8f3

Observation ed653d77-400e-476f-b68f-b3871cd5a005 · outbound

This paper cites Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.147887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.147887Z digest=sha256:8029dd1e4546701951570a2527f1159eb8fd19290cb686b0d672cc066066dc12

Observation 711f7380-ba74-4da7-866d-6b0f7556bf10 · outbound

This paper cites SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.224128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.224128Z digest=sha256:cdea479a8fb5dd7da25a3bac38f6a2a5458f97c832f8085d5bbeec6461383c8c

Observation 3dc34d7f-6197-46c2-b0e1-71ca5554b36e · outbound

This paper cites When can dictionary learning uniquely recover sparse data from subsamples?IEEE Transactions on Information Theory, 61(11):6290–6297, 2015.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs When can dictionary learning uniquely recover sparse data from subsamples?IEEE Transactions on Information Theory, 61(11):6290–6297, 2015

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:07.725224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:00.335412Z digest=sha256:56e4011673f08f5e48f5380aa0b6a3ed23c0e6709f7ed88a639f4efa5ea5bdcf

Observation 1e0a86e9-fb72-408a-aca9-8bd88cdc2b1d · outbound

This paper cites Projecting assumptions: The duality between sparse autoencoders and concept geometry.arXiv preprint arXiv:2503.01822, 2025.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Projecting assumptions: The duality between sparse autoencoders and concept geometry.arXiv preprint arXiv:2503.01822, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.464472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.464472Z digest=sha256:7bc04386660c9f85b3bd633910b0f91d5a1404ace2c77545016f9c2f8ac3f5e6

Observation 30ea7564-3ac3-4918-8c04-284def90e3d4 · outbound

This paper cites Independent component analysis: algorithms and applications.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Independent component analysis: algorithms and applications

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.545532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.545532Z digest=sha256:d3b9fee5f839879883f1fe7ff58333ae38ac377c0e9164043ba3dc30c725d619

Observation 40145398-b7ba-4508-9d24-4cac2ad7d40a · outbound

This paper cites Identifiable steering via sparse autoencoding of multi-concept shifts.arXiv preprint arXiv:2502.12179, 2025.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Identifiable steering via sparse autoencoding of multi-concept shifts.arXiv preprint arXiv:2502.12179, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.599280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.599280Z digest=sha256:0d01e073632f4aebd98866407b34f10def43cb463adcd9e25b495c8b1410d091

Observation c34b9b99-f69f-4e55-acd7-16470e758eac · outbound

This paper cites SAEBench: A Comprehensive Benchmark for Sparse Autoencoders.https://www.neuronpedia.org/sae-bench/info, 2024.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders.https://www.neuronpedia.org/sae-bench/info, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:07.558514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:00.650036Z digest=sha256:eb97db4c2ae032358d5fb8738a9cdde1ad7cd6406ee6ca9fd71100ee6bd55200

Observation d5a3890e-9477-4eb5-adb4-83caf2541240 · outbound

This paper cites Measuring progress in dictionary learning for language model interpretability with board game models.Advances in Neural Information Processing Systems, 37:83091–83118, 2024.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Measuring progress in dictionary learning for language model interpretability with board game models.Advances in Neural Information Processing Systems, 37:83091–83118, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:07.323869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:00.752206Z digest=sha256:3ee3cc79483acd716134e8d7e31ab51f68c1e055141752014cafd819ee8c567e

Observation 9472c785-10ea-4cc4-bee1-c5ab6b941e6d · outbound

This paper cites Sparse Autoencoders Do Not Find Canonical Units of Analysis.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Autoencoders Do Not Find Canonical Units of Analysis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.843301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.843301Z digest=sha256:7322265ce0a036bb54ce34bb01c91141be8823234c5cd35741a2d70425e41038

Observation 5afc856f-5def-4f83-a651-10d75b848acb · outbound

This paper cites an unresolved cited work.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.948695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.948695Z digest=sha256:d06ad74afbb78f21af950083388fdde6f3134b66d701ef878cc968bc83d8c411

Observation a6d3d15b-1428-4035-a894-0823aaa09cf9 · outbound

This paper cites The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.Queue, 16(3):31–57, 2018.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.Queue, 16(3):31–57, 2018

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.063367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.063367Z digest=sha256:b65f1a78f4fe35501fdad89e977711f6e4acfa95f85bf596d04fda7f684a7c10

Observation 9e74e2b7-c5fd-44be-881c-5fd1e2935ffc · outbound

This paper cites Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.213348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.213348Z digest=sha256:4c0496a322b690ec5e0e58ea39ce98efbb19dd9accbd6e6a1a64e5ab09bd19ce

Observation 4ccd8b59-f9fb-46b4-bddd-424074957b94 · outbound

This paper cites Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.307047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.307047Z digest=sha256:f2b17a75330faf991f33163b118b297e2fe16de6f8b89072a38504f829711ca8

Observation c2c6a64f-b920-4ca4-ab4c-4408c1615757 · outbound

This paper cites Dictionary learning.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Dictionary learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:07.152154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:01.452767Z digest=sha256:c2dd048651d374ff3a897c96e44270753b9d4900c349272a330293a2f5bc5b99

Observation 2aa0ffdd-ad11-4ee0-a36f-eb6a806197e4 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.502160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.502160Z digest=sha256:fa23e0f24f1c9245869dd425894875cb28893ad35058a978b20252a7d86c0cfc

Observation 70443d27-4dcd-47a8-af9a-664c3dc8acec · outbound

This paper cites Sparse feature circuits: Discovering and editing interpretable causal graphs in language models.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse feature circuits: Discovering and editing interpretable causal graphs in language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:06.970419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:01.565340Z digest=sha256:da1740273931134b43c714437d2d47b05e360d2640370346babc523fbb6f513c

Observation 611694cd-002a-4a4c-9c56-819eca673eed · outbound

This paper cites Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.725857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.725857Z digest=sha256:ce10a2cd1496609a864400357a6f4776206207c3e9fb38915e0a2db074e64376

Observation 5aa98ec1-5520-44bb-8dfb-b821c908129a · outbound

This paper cites SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.830004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.830004Z digest=sha256:635cecf85f137b9e345b4f0c9b7305284577931b75d36654ecf01b31ac92a29c

Observation 8971053d-2e39-4d5b-9596-5976a58bb55b · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Steering Language Model Refusal with Sparse Autoencoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.923666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.923666Z digest=sha256:f93f937a9aea7c93cfc138028924ca705401c32970796de049287670488c69b3

Observation c0744883-b547-4b5c-b50f-e89ef3fdaee1 · outbound

This paper cites Mechanistic interpretability, variables, and the importance of interpretable bases.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Mechanistic interpretability, variables, and the importance of interpretable bases

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:06.824381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:02.049323Z digest=sha256:e6eaf043532d9fff15d24339d81539205bd5b1fd74705cc93590207fd06b886a

Observation f6606495-3294-41fe-b978-60cefaafdf67 · outbound

This paper cites Zoom in: An introduction to circuits.Distill, 5(3):e00024–001, 2020.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Zoom in: An introduction to circuits.Distill, 5(3):e00024–001, 2020

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.169916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.169916Z digest=sha256:e767da0ad6d9b367d79f9e45374e6988081eeb10d169d0863f5cd61b53a224ba

Observation 3db0a845-75cf-492c-8210-cf42f735372d · outbound

This paper cites The building blocks of interpretability.Distill, 3(3):e10, 2018.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The building blocks of interpretability.Distill, 3(3):e10, 2018

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:06.551552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:02.289336Z digest=sha256:0cd2dbba738716ff2e39d1e73c4b7db3a9bb2eac53722af7a9a66118df680706

Observation 80efe505-6c96-40c4-806b-a586fdc29537 · outbound

This paper cites Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.401440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.401440Z digest=sha256:0d73b93f747ddf532b2530be10b5176dd5c64cb37790a07629b03bded764b4f9

Observation 962b0a1f-79ea-4677-a29a-dddb71067f29 · outbound

This paper cites Sparse autoencoders learn monosemantic features in vision-language models.arXiv preprint arXiv:2504.02821, 2025.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse autoencoders learn monosemantic features in vision-language models.arXiv preprint arXiv:2504.02821, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.554813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.554813Z digest=sha256:143c9a484fa8e02fd5ae4bc3f7f7c2c557a052e193dae4d2ee887cec00b234fe

Observation 0174c757-c952-4a4b-86e9-ca37bda1d101 · outbound

This paper cites Sparse Autoencoders Trained on the Same Data Learn Different Features.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Autoencoders Trained on the Same Data Learn Different Features

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.656998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.656998Z digest=sha256:286e277759df51d5df5e972d4048e7cef00c9596fc67ea0e447a2a63ac3bf820

Observation 69d5d226-67de-45f7-aa85-cd7999a7ea30 · outbound

This paper cites Automatically Interpreting Millions of Features in Large Language Models.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Automatically Interpreting Millions of Features in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.798784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.798784Z digest=sha256:2c8c4b24778882335f5c19498d7c9132115df601d3678471463120cb6c8ce53e

Observation 9157edce-d3e0-47ee-afc8-fd9ec6246fb9 · outbound

This paper cites Improving Dictionary Learning with Gated Sparse Autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Improving Dictionary Learning with Gated Sparse Autoencoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.963645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.963645Z digest=sha256:a0ed3a699d9041b09ee2dee4f97fca2ff1d0b884b891263b776f593a22c71f1b

Observation 49a5bc79-cd46-4749-b6f5-eef6fb7db0f6 · outbound

This paper cites Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.074303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:03.074303Z digest=sha256:5609f4d6c35e414341c273f830bd867cdd26837d4112e75bb9f30667f406ea49

Observation c62a6bcc-f559-49e7-8447-343d0259e375 · outbound

This paper cites Global identifiability of overcomplete dictionary learning via l1 and volume minimization.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Global identifiability of overcomplete dictionary learning via l1 and volume minimization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:06.409527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:03.174968Z digest=sha256:8ce5a175e00f12083ab5814edb1220fd10a8dc5ab5cc600c7d721fa5d3eca86f

Observation c4230207-d7dc-463a-b104-9ce6c21900af · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Gemma: Open Models Based on Gemini Research and Technology

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.259155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:03.259155Z digest=sha256:42751239801d48784f895a2a36bdcd78828448df2536d39d5192ec4b19192870

Observation 34708252-3bf4-45cc-8e99-1421f135b8af · outbound

This paper cites Identifiability of overcomplete independent component analysis.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Identifiability of overcomplete independent component analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.323826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:03.323826Z digest=sha256:885b794098ef862878a2ff9b52f7748fd811100e929a42dbaf459e8685240596

Observation 6f2f630b-77d6-4bae-8949-39cd1ddfbc23 · outbound

This paper cites Only the most frequent clusters show consistent reproducibility, indicating severe capacity limitations where the dictionary should prioritize only the dominant clusters.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Only the most frequent clusters show consistent reproducibility, indicating severe capacity limitations where the dictionary should prioritize only the dominant clusters

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:05.923848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:03.575402Z digest=sha256:df567a465e14f1ec38a661fe7f05da640f175c8b93978f0aa934fe425fcf82a6

Observation 9d6bfcb5-31d0-4e1f-84e2-1d5cb285e950 · outbound

This paper cites an unresolved cited work.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:03:05.688187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:03.648467Z digest=sha256:c806d552f5d932fdd28b30f2df421315055ae61c4152815f8d6bf20d7a5c5c3f

Observation 05aabd42-d69e-481d-9abc-b6c47954fd7c · outbound

This paper cites The substantial increase in capacity 28 Figure 31:Two-phase model with dictionary size.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The substantial increase in capacity 28 Figure 31:Two-phase model with dictionary size

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:05.440082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:03.719640Z digest=sha256:90eaaae67b8a2c97c83ef4deebbcb843a6653e5cddf406a863afffa2d82665c7

Observation 475cb85b-8abe-441b-a0b3-9d13000a2038 · outbound

This paper cites Figure 30:Two-phase model with dictionary size.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Figure 30:Two-phase model with dictionary size

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:06.246169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:03.436592Z digest=sha256:5279317c4b1eb6d0a0c02746c363c9e86e0c908ef8cf1c00cd67268301db1c4c

Observation c5fb5432-f866-431a-b510-6ae0accdb7c2 · outbound

This paper cites Figures 29 through 32 demonstrate how dictionary size affects feature reproducibility across the activation frequency spectrum.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Figures 29 through 32 demonstrate how dictionary size affects feature reproducibility across the activation frequency spectrum

Reference 160

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:06.087258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:03.496179Z digest=sha256:78ae6d10352ff30a54c234d0b68eed902952ba086262ad6a48eb30d81f35e34d

Observation 8cfb17c6-1cd2-4de3-bf46-92a29dd0e9ad · outbound

This paper cites the same.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs the same

Reference 1000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:05.279790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:03:03.819990Z digest=sha256:fc5cc8a0c08adf356e85939372f330f9012cfc8cd461675d5331809db51f2958

Pith citing papers

Observation 29fe38af-348b-4e31-95a9-11839d546684 · inbound

Graph-Regularized Sparse Autoencoders for LLM Safety Steering cites this paper.

Graph-Regularized Sparse Autoencoders for LLM Safety Steering Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:40:28.916617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T18:39:02.687301Z digest=sha256:e10b89bbf5ab3c4fd00da366807af955ba201f4d55de8a148b0595f2922c5a31

Observation bab755f3-6fa3-47e9-87e4-fef0bc161a45 · inbound

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing cites this paper.

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:58.300413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T17:01:39.362411Z digest=sha256:ef5e65765ca0c22159434d0d4add5b18e526f4939e52c984f26824604c5f02a5