Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:24:53.899236Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 15 inbound Pith citation observations for arXiv:2501.18823.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:24:53.899236Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:51:54.930894Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:07:28.632106Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ff980a68-7ac6-4f52-b33f-861499eda96b · outbound
Transcoders Beat Sparse Autoencoders for Interpretability write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ef0f91-5384-4edd-ae4b-0689fe697a6a · outbound
Transcoders Beat Sparse Autoencoders for Interpretability write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1deecc2a-75a9-4a21-8d97-7fc1af5964f5 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Linear algebraic structure of word senses, with applications to polysemy
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75f0885-0c9d-4757-85d2-b0dd53aa682d · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ceac99-7411-4570-8fd5-1c8e73cd3969 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Mechanistic Permutability: Match Features Across Layers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c92d59d-ecd3-48c4-bba3-55e9d361e7ac · outbound
Transcoders Beat Sparse Autoencoders for Interpretability and Baraniuk, R
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 171141ac-5dd8-443d-9911-fbc3b5c474a8 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183bc9a1-993c-46aa-a87f-aec5fa3beb1d · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Language models can explain neurons in language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 34ecd6ab-fc1a-4ff6-9fd1-d4953946937e · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Interpreting Neural Networks through the Polytope Lens
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2286ff-1297-4bc4-b0e3-70a2126e340c · outbound
Transcoders Beat Sparse Autoencoders for Interpretability E., Hume, T., Carter, S., Henighan, T., and Olah, C
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5a83ef4a-c769-4058-bbf5-cf57a4232c51 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability L., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Tamkin, A., Nguyen, K., McLean, B., Burke, J
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c0e40d2d-ca77-4bc1-be06-adb47144d086 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Learning multi-level features with matryoshka saes
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 32562f45-8c26-4456-af7b-cfbeb6f2ac8f · outbound
Transcoders Beat Sparse Autoencoders for Interpretability BatchTopK Sparse Autoencoders
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389ab1e5-e18c-4de9-867d-e47d9a8d8a5a · outbound
Transcoders Beat Sparse Autoencoders for Interpretability A is for absorption: Studying feature splitting and absorption in sparse autoencoders
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8eeb20a-a837-4c55-896a-7ecb026afc01 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Redpajama: an open dataset for training large language models, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 69061fed-e49d-4c2b-8b98-dd09645abde2 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Towards Automated Circuit Discovery for Mechanistic Interpretability
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395cc7e1-7f8b-4705-9253-755d0aa90007 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability The Llama 3 Herd of Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e814619-2107-4595-9062-a38af8b9a3cd · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Transcoders Find Interpretable LLM Feature Circuits
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df94c201-6434-452c-b0bd-d3f2697fb5e2 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Toy Models of Superposition
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1633fff6-25b8-475b-910f-d95bf28d11e6 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Decomposing The Dark Matter of Sparse Autoencoders
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ecd7d15-425d-4c92-8168-8b4d661dbcc7 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92cfc6f-07db-40e1-af73-6b5443be6c4d · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Scaling and evaluating sparse autoencoders
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe171f0-b4e3-422f-9576-ec5e485081a1 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Openwebtext corpus
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a12753-7b4d-4ee9-b5d2-14881b336d8d · outbound
Transcoders Beat Sparse Autoencoders for Interpretability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ebfe189-4b58-4945-9c96-e72c608cf0cc · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Finding Neurons in a Haystack: Case Studies with Sparse Probing
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b284469-f521-48a7-8ea5-57e1dedfdef2 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Universal Neurons in GPT2 Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 818139f1-5ca3-4a22-a4e3-edec26335ba1 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability and Templeton, A
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0ce159be-dc41-4fcf-896d-9e2740c0b022 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Open source automated interpretability for sparse autoencoder features, July 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d99addce-95f8-44be-894f-8742ecf48bff · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Saebench: A comprehensive benchmark for sparse autoencoders, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5b7092ab-42a5-4337-b731-bfb30cf0c284 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Adam: A Method for Stochastic Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 837accc0-d372-44ad-95a2-468329773b9b · outbound
Transcoders Beat Sparse Autoencoders for Interpretability dictionary\_learning repository, 2023
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 641a3294-e646-4bd9-8b15-5d1d54510444 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3aeeeab-a92a-4579-8e2e-1c08ed9acdea · outbound
Transcoders Beat Sparse Autoencoders for Interpretability and Black, S
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebddf325-94da-4cc4-b676-e1c83342243a · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Matryoshka sparse autoencoders
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c73c176d-a85c-43b6-aed6-04752cfb4e1c · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Zoom in: An introduction to circuits
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3227a6-362c-4f27-9081-ab790f79be4d · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Automatically Interpreting Millions of Features in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5345284-d47d-47ee-8b94-e6238ff5ea68 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e6aca4-03a4-4a3a-89ba-c2b5f79cbd5b · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0c6696-fcc1-4806-8547-36b3c200c9f0 · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Gemma 2: Improving Open Language Models at a Practical Size
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2122e0c-0eea-4d37-860a-88768c3e5d5c · outbound
Transcoders Beat Sparse Autoencoders for Interpretability Predicting future activations
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d504b34f-91c6-413d-a3b8-897d9553d8ef · outbound
Transcoders Beat Sparse Autoencoders for Interpretability L., McDougall, C., MacDiarmid, M., Freeman, C
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d946496c-4574-4620-ba8c-415dd49173f8 · inbound
Capacity Matters: a Proof-of-Concept for Transformer Memorization on Real-World Data Transcoders Beat Sparse Autoencoders for Interpretability
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf2d350c-cb39-4c0f-a65c-3cbc13eb9773 · inbound
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Transcoders Beat Sparse Autoencoders for Interpretability
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc1aede-4cea-4018-8c24-e849deda137a · inbound
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Transcoders Beat Sparse Autoencoders for Interpretability
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c7f318a9-708d-40f5-bdcf-ec8586895f68 · inbound
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Transcoders Beat Sparse Autoencoders for Interpretability
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb10d04b-b467-4a0c-9798-ec2865a3e301 · inbound
Improving Robustness In Sparse Autoencoders via Masked Regularization Transcoders Beat Sparse Autoencoders for Interpretability
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3d5a7cdd-308c-4eb7-aceb-ab156224698d · inbound
WriteSAE: Sparse Autoencoders for Recurrent State Transcoders Beat Sparse Autoencoders for Interpretability
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e9270133-f940-483f-bae7-05cf498133f4 · inbound
WriteSAE: Sparse Autoencoders for Recurrent State Transcoders Beat Sparse Autoencoders for Interpretability
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3141c122-7565-4733-93fa-82bf691bff82 · inbound
WriteSAE: Sparse Autoencoders for Recurrent State Transcoders Beat Sparse Autoencoders for Interpretability
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ee2ed23b-0ccf-442a-8ca6-883847cd3c92 · inbound
WriteSAE: Sparse Autoencoders for Recurrent State Transcoders Beat Sparse Autoencoders for Interpretability
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9c655d3f-e870-4ddb-81b2-5819a0ea2fad · inbound
Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift Transcoders Beat Sparse Autoencoders for Interpretability
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 12c2ecd1-63ed-4ae3-a766-263a3282c8b5 · inbound
Interactions Between Crosscoder Features: A Compact Proofs Perspective Transcoders Beat Sparse Autoencoders for Interpretability
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1f9592c2-90e0-4c90-8058-7b157fe71aae · inbound
Transcoders for Investigating Deception in Language Models Transcoders Beat Sparse Autoencoders for Interpretability
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c532033-01e8-40d0-9379-caa8dc15910d · inbound
Verbalizable Representations Form a Global Workspace in Language Models Transcoders Beat Sparse Autoencoders for Interpretability
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543bdd20-a41a-4afd-a618-ce2f082335bd · inbound
Evading Chain-of-Thought Monitoring Through Model Poisoning Transcoders Beat Sparse Autoencoders for Interpretability
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39137c1-d6e6-4e27-838f-4d3f4bb6b6ac · inbound
Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability Transcoders Beat Sparse Autoencoders for Interpretability
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.