Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2403.00745.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:44:55.426570Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation cb6c96a8-3327-4d2b-8554-8db5b559bc76 · inbound
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a9fc4b30-04b2-4085-9193-6378f20f8b78 · inbound
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3476cbbb-24e0-4dfd-9492-ed5e5f5c9ffd · inbound
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc3b048-ff07-46d0-b532-91eb721ae6a2 · inbound
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1feee041-aced-46ba-90f2-5eef86974bed · inbound
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61541782-adb1-407b-872e-f46ea5afb7cc · inbound
Language Models Use Trigonometry to Do Addition AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 178e618b-1e7b-4d39-bdf7-12e5b6a543dc · inbound
Circuit Stability Characterizes Language Model Generalization AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1215c80c-f397-4304-89cd-3a2e21d794a6 · inbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c63bfe-0596-4ab6-bba2-962ba17708a8 · inbound
Steering Conceptual Bias via Transformer Latent-Subspace Activation AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730a5abb-3bfe-4ef2-9158-8b517b1b6833 · inbound
Stochastic Parameter Decomposition AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 985b5627-8aaa-47aa-9bd3-5ba6ef2d94fe · inbound
Can Interpretation Predict Behavior on Unseen Data? AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fe8a39-8c9d-4395-a5d8-e911148a90ba · inbound
From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e5023d-d217-4135-b3f4-dfb4296ce340 · inbound
Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b7172ed3-f045-4d23-963d-bba7affc696b · inbound
Weight Patching: Toward Source-Level Mechanistic Localization in LLMs AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef1a1761-b1f8-49dd-859b-44b974575a16 · inbound
Cell-Based Representation of Relational Binding in Language Models AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation af8e2676-9e65-4045-80f5-1b3dee73b568 · inbound
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 60b445b2-3634-42d4-8bd6-3db03aa0f56d · inbound
How LLMs Are Persuaded: A Few Attention Heads, Rerouted AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a01933e-9258-4752-82aa-db9a0d792ccd · inbound
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ed325fa1-5b59-4798-bcc5-af82bdf3da05 · inbound
WriteSAE: Sparse Autoencoders for Recurrent State AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9b3f5625-0118-460a-a276-7614e5b53e05 · inbound
WriteSAE: Sparse Autoencoders for Recurrent State AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3b120abc-9cfa-44c0-b74b-86b5886ac5a6 · inbound
Polymorphism Is Rotation: Operational Mechanistic Interpretability from a Two-Layer Transformer to Pythia-70m AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4fbca113-87b9-49e8-b0ef-13689a794349 · inbound
Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 20f6d7b7-97b2-4622-8e0b-86c688ee4fb0 · inbound
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dfbe8c23-2036-40e9-bb70-a153aefae165 · inbound
Subliminal Learning is a LoRA Artifact AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 798b7c1b-5a32-4a70-8b18-838af695c2a5 · inbound
Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 65472b2b-5d9b-42d9-890d-f62ae5d12c07 · inbound
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2997d91a-32dc-4312-91a4-75a6063e6f3c · inbound
When Attribution Patching Lies: Diagnosis and a Second-Order Correction AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a03226b5-68db-4d60-8ba7-6a5d3d399e43 · inbound
Localizing Anchoring Pathways in Language Models AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 313d37f0-6013-4bd7-abb9-11a3ac5410bc · inbound
Beyond Importance: Interchange-Sobol Sensitivity Reveals Task-Specific Content Channels in Transformer Components AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf7af405-a3e5-4ff6-b25e-517c8c9f69e7 · inbound
Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 92b12914-ff66-4388-b3c1-3292c675f2f1 · inbound
Verbalizable Representations Form a Global Workspace in Language Models AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529e8a8c-b88a-4f1f-ba12-4ab4cd46af65 · inbound
Circuit Claims Depend on What Is Extracted and How It Is Compared AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ffb5c0-1972-4a93-9f01-27fb72ec9e53 · inbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74cab79d-a110-4087-b37f-9a423274db9a · inbound
Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b10829-4f1c-4001-9a7c-2183dbf21c1c · inbound
Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.