Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2410.13928.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:21:24.356298Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 93ba4555-f892-4879-b9d5-5d231f5d255a · inbound
Interpretable Company Similarity with Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf2712e3-c2a5-4c81-826a-ad5abe3ff685 · inbound
Interpretable Steering of Large Language Models with Feature Guided Activation Additions Automatically Interpreting Millions of Features in Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490207d3-85f9-47f7-8811-19206fda2a07 · inbound
Propositional Interpretability in Artificial Intelligence Automatically Interpreting Millions of Features in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b5e109b-dc60-434f-b6d1-9e38a1c46758 · inbound
Sparse Autoencoders Trained on the Same Data Learn Different Features Automatically Interpreting Millions of Features in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ae2cc8-7ddf-476a-8d9e-5e36b9d259a3 · inbound
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3227a6-362c-4f27-9081-ab790f79be4d · inbound
Transcoders Beat Sparse Autoencoders for Interpretability Automatically Interpreting Millions of Features in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a4bfb8-1bce-41ae-bf01-96cc4d3259b4 · inbound
Partially Rewriting a Transformer in Natural Language Automatically Interpreting Millions of Features in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7018977c-9838-49b4-badc-ba1620c36b4b · inbound
Converting MLPs into Polynomials in Closed Form Automatically Interpreting Millions of Features in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404a1caf-9e0e-48b1-b0bc-f3deb3b9ee22 · inbound
Breaking Down Bias: On The Limits of Generalizable Pruning Strategies Automatically Interpreting Millions of Features in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec0bdb8c-202d-4452-97e6-e8969bf8b555 · inbound
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models Automatically Interpreting Millions of Features in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d5d226-67de-45f7-aa85-cd7999a7ea30 · inbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Automatically Interpreting Millions of Features in Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dced67a-930b-4899-906d-9d518eb3e32c · inbound
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy Automatically Interpreting Millions of Features in Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e6b6711-37ba-499f-b663-95b831dca041 · inbound
Evaluating SAE interpretability without explanations Automatically Interpreting Millions of Features in Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a332b2-4bc4-4fd7-a769-e59f5658fa1c · inbound
Insights into a radiology-specialised multimodal large language model with sparse autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc944e05-a04d-4239-905f-0ffb17c18ff6 · inbound
Teach Old SAEs New Domain Tricks with Boosting Automatically Interpreting Millions of Features in Large Language Models
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a8fa38-3718-4b12-9db2-5754e64ce507 · inbound
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38cada16-d398-4e83-9f60-d806293f9950 · inbound
Distribution-Aware Feature Selection for SAEs Automatically Interpreting Millions of Features in Large Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51200a8a-a44b-4b9a-a1cf-f3c2243dc8dd · inbound
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework Automatically Interpreting Millions of Features in Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8f99d151-6988-49f1-b0a0-644c1b031df4 · inbound
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Automatically Interpreting Millions of Features in Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 80cfe4ec-b4c5-455a-9f9b-ca5f4da4545e · inbound
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Automatically Interpreting Millions of Features in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fa6dbd-3e7a-44e4-9b18-20d62b4e15f5 · inbound
Prototype Transformer: Towards Language Model Architectures Interpretable by Design Automatically Interpreting Millions of Features in Large Language Models
Reference 1995
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24ca6d44-11fb-4f92-bdf2-1bd1d35942ab · inbound
Visual Persuasion: What Influences Decisions of Vision-Language Models? Automatically Interpreting Millions of Features in Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0a86bbe-392a-4f0e-b54e-70c86ce35603 · inbound
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs Automatically Interpreting Millions of Features in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99504919-6994-4918-bf36-9c0050b1a81a · inbound
MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents Automatically Interpreting Millions of Features in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 18442c76-67d5-4f2f-b0b5-2d67361ad47b · inbound
LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering Automatically Interpreting Millions of Features in Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a18cc7-8459-40cc-9acb-5163d2c730bb · inbound
Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Automatically Interpreting Millions of Features in Large Language Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e7e69111-eaf2-43e7-a12a-c31c0c074c66 · inbound
From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features Automatically Interpreting Millions of Features in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c22a0357-cea8-4493-bb17-9119b92e3f35 · inbound
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9408900b-b8ff-4f83-aad1-f656f7fbcb0e · inbound
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 92412ab3-e893-4d21-a272-7edffd59c421 · inbound
Domain Restriction via Multi SAE Layer Transitions Automatically Interpreting Millions of Features in Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b619e30-0b39-4ad3-92ea-65ee8030c09b · inbound
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Automatically Interpreting Millions of Features in Large Language Models
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 90c19cd7-8b28-4b13-9d10-29df29899b30 · inbound
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features Automatically Interpreting Millions of Features in Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2023cc4a-ebf5-4238-84e6-b6d728fb9fbd · inbound
Why Retrieval-Augmented Generation Fails: A Graph Perspective Automatically Interpreting Millions of Features in Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4472d2ca-8d2f-417b-ba76-a821d63b5696 · inbound
The Rate-Distortion-Polysemanticity Tradeoff in SAEs Automatically Interpreting Millions of Features in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 54eeb67b-f683-4846-b5c7-6f591391cb73 · inbound
Are Sparse Autoencoder Benchmarks Reliable? Automatically Interpreting Millions of Features in Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e599bae1-869d-471e-af08-ac21f317a024 · inbound
Features have life history. And we should care Automatically Interpreting Millions of Features in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ffd95b3-ce31-4606-902a-cf1f63dc9dc1 · inbound
SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization Automatically Interpreting Millions of Features in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 59891a17-6ca1-43c0-a706-ff00c4581768 · inbound
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aca94aaf-78fb-4915-a67b-705996dd12b4 · inbound
ICA Lens: Interpreting Language Models Without Training Another Dictionary Automatically Interpreting Millions of Features in Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b94cef44-8a3d-49b8-9a8b-b2329d44269a · inbound
Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Automatically Interpreting Millions of Features in Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be03c80f-db8d-48c6-9e78-80c0d70d54a6 · inbound
Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f5c018c0-08fe-4128-a372-79e995f9b0fa · inbound
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? Automatically Interpreting Millions of Features in Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e6cb9511-ab47-4020-be0c-38adefadde2f · inbound
Training, Reading, and Editing Legible Transformers Automatically Interpreting Millions of Features in Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378236f7-3528-43ed-b6df-9fe19f1c222f · inbound
Verbalizable Representations Form a Global Workspace in Language Models Automatically Interpreting Millions of Features in Large Language Models
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f466a6-23af-49c0-9269-a828189cd789 · inbound
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels Automatically Interpreting Millions of Features in Large Language Models
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.