Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:03.819990Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2505.20254.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:03.819990Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-04T17:01:39.362411Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T17:09:58.298491Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d985f89e-c3d7-481f-bcb8-a67ba9aeba66 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84fc0e70-98a5-443c-a3cd-cbfa08c00f8f · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 645c8277-9cca-4b46-8479-21efa3b5726c · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs New algorithms for learning incoherent and overcomplete dictionaries
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c35eff-2e77-4f2d-a7b9-d27f17667d93 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Pythia: A suite for analyzing large language models across training and scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d388b3e2-e670-49b7-8ca1-79094ac30591 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Turner, Cem Anil, Carson Denison, Amanda Askell, et al
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e0457bd2-f38e-4f23-a0f5-9e841a08dcc5 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Towards monosemanticity: Decomposing language models with dictionary learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6a38c3d4-af46-42a5-87b0-f690c11b3c21 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs BatchTopK Sparse Autoencoders
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f48c707-9311-4ae1-a416-49bbd095833e · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Learning Multi-Level Features with Matryoshka Sparse Autoencoders
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04aa17b-e3fd-4de2-a6cf-85e93322aba3 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Improving Steering Vectors by Targeting Sparse Autoencoder Features
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1363a7c9-9a08-4aba-8d9a-1e8ee822cd3a · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs A is for absorption: Studying feature splitting and absorption in sparse autoencoders.arXiv preprint arXiv:2409.14507, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 052532a1-9c86-440b-96ad-92209798469f · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb97890a-4497-4757-91c6-f1b8787df341 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Optimally sparse representation in general (nonorthogonal) dictionaries via l1 minimization.Proceedings of the National Academy of Sciences, 100(5):2197– 2202, 2003
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ea1036f8-040f-4298-a3da-7ae2ee0a1cdf · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Toy Models of Superposition
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e73a26d-6fa2-4fbd-b629-66a9067e20c9 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs A mathematical framework for transformer circuits.Transformer Circuits Thread, 1(1):12, 2021
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db898c7e-37ed-42a2-aed4-64927b95be1a · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17262248-3e93-4e91-99ee-b4c16cfb830b · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Scientific inference with interpretable machine learning: Analyzing models to learn about real-world phenomena.Minds and Machines, 34(3):32, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9a340735-39cf-4e46-b67a-61471657af26 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44650b81-0291-42d5-9ccd-10c19e3d95ab · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Scaling and evaluating sparse autoencoders
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e0a001-c085-469c-9db9-6eaa35db7c86 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Scaling and evaluating sparse autoencoders
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ed653d77-400e-476f-b68f-b3871cd5a005 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 711f7380-ba74-4da7-866d-6b0f7556bf10 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc34d7f-6197-46c2-b0e1-71ca5554b36e · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs When can dictionary learning uniquely recover sparse data from subsamples?IEEE Transactions on Information Theory, 61(11):6290–6297, 2015
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1e0a86e9-fb72-408a-aca9-8bd88cdc2b1d · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Projecting assumptions: The duality between sparse autoencoders and concept geometry.arXiv preprint arXiv:2503.01822, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ea7564-3ac3-4918-8c04-284def90e3d4 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Independent component analysis: algorithms and applications
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40145398-b7ba-4508-9d24-4cac2ad7d40a · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Identifiable steering via sparse autoencoding of multi-concept shifts.arXiv preprint arXiv:2502.12179, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c34b9b99-f69f-4e55-acd7-16470e758eac · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders.https://www.neuronpedia.org/sae-bench/info, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d5a3890e-9477-4eb5-adb4-83caf2541240 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Measuring progress in dictionary learning for language model interpretability with board game models.Advances in Neural Information Processing Systems, 37:83091–83118, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9472c785-10ea-4cc4-bee1-c5ab6b941e6d · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5afc856f-5def-4f83-a651-10d75b848acb · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d3d15b-1428-4035-a894-0823aaa09cf9 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.Queue, 16(3):31–57, 2018
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e74e2b7-c5fd-44be-881c-5fd1e2935ffc · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ccd8b59-f9fb-46b4-bddd-424074957b94 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c6a64f-b920-4ca4-ab4c-4408c1615757 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Dictionary learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2aa0ffdd-ad11-4ee0-a36f-eb6a806197e4 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70443d27-4dcd-47a8-af9a-664c3dc8acec · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 611694cd-002a-4a4c-9c56-819eca673eed · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa98ec1-5520-44bb-8dfb-b821c908129a · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8971053d-2e39-4d5b-9596-5976a58bb55b · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Steering Language Model Refusal with Sparse Autoencoders
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0744883-b547-4b5c-b50f-e89ef3fdaee1 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Mechanistic interpretability, variables, and the importance of interpretable bases
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f6606495-3294-41fe-b978-60cefaafdf67 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Zoom in: An introduction to circuits.Distill, 5(3):e00024–001, 2020
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db0a845-75cf-492c-8210-cf42f735372d · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The building blocks of interpretability.Distill, 3(3):e10, 2018
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 80efe505-6c96-40c4-806b-a586fdc29537 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962b0a1f-79ea-4677-a29a-dddb71067f29 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse autoencoders learn monosemantic features in vision-language models.arXiv preprint arXiv:2504.02821, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0174c757-c952-4a4b-86e9-ca37bda1d101 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Autoencoders Trained on the Same Data Learn Different Features
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d5d226-67de-45f7-aa85-cd7999a7ea30 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Automatically Interpreting Millions of Features in Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9157edce-d3e0-47ee-afc8-fd9ec6246fb9 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Improving Dictionary Learning with Gated Sparse Autoencoders
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a5bc79-cd46-4749-b6f5-eef6fb7db0f6 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c62a6bcc-f559-49e7-8447-343d0259e375 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Global identifiability of overcomplete dictionary learning via l1 and volume minimization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c4230207-d7dc-463a-b104-9ce6c21900af · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Gemma: Open Models Based on Gemini Research and Technology
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34708252-3bf4-45cc-8e99-1421f135b8af · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Identifiability of overcomplete independent component analysis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2f630b-77d6-4bae-8949-39cd1ddfbc23 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Only the most frequent clusters show consistent reproducibility, indicating severe capacity limitations where the dictionary should prioritize only the dominant clusters
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9d6bfcb5-31d0-4e1f-84e2-1d5cb285e950 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 05aabd42-d69e-481d-9abc-b6c47954fd7c · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs The substantial increase in capacity 28 Figure 31:Two-phase model with dictionary size
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 475cb85b-8abe-441b-a0b3-9d13000a2038 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Figure 30:Two-phase model with dictionary size
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c5fb5432-f866-431a-b510-6ae0accdb7c2 · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Figures 29 through 32 demonstrate how dictionary size affects feature reproducibility across the activation frequency spectrum
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8cfb17c6-1cd2-4de3-bf46-92a29dd0e9ad · outbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs the same
Reference 1000
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 29fe38af-348b-4e31-95a9-11839d546684 · inbound
Graph-Regularized Sparse Autoencoders for LLM Safety Steering Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bab755f3-6fa3-47e9-87e4-fef0bc161a45 · inbound
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.