Pith. sign in

Paper Citation Record · LEDGER

Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2303.02536.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.02536 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:32:36.785476Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2fc06f0-e7a3-46f3-afb6-47d45d3628ec · inbound

Localizing Model Behavior with Path Patching cites this paper.

Localizing Model Behavior with Path Patching Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:38:37.793826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T19:38:37.751487Z digest=sha256:f347ba54f7f6ce35ef365a7bddd68c7e808bbc1e8bd5400a5dbe97f352aee719

Observation 8488324c-11e8-4be5-9705-375d81cb4875 · inbound

Towards Best Practices of Activation Patching in Language Models: Metrics and Methods cites this paper.

Towards Best Practices of Activation Patching in Language Models: Metrics and Methods Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:56:11.181983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T11:56:11.053897Z digest=sha256:bb4bd44be167e3a13f1d2f7dfee0aad9a4d68d72dd5fbb7a5d7094538aae16ba

Observation 024c1d07-87f5-4b2e-9bc5-4ecb81d25b8b · inbound

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition cites this paper.

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:55:02.030913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:55:02.030913Z digest=sha256:52929c2bdd9370b173d26407d5e8f702b87fc9376481c6192185c127438449c2

Observation 401c55ea-6fe8-4735-bc5c-fac8ed930e13 · inbound

What is a Number, That a Large Language Model May Know It? cites this paper.

What is a Number, That a Large Language Model May Know It? Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T15:05:03.727012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:05:03.727012Z digest=sha256:beb17c975a0060feb84a7486575e12a20f7fd021d50323c840c9815df3bbe65f

Observation 11ac7f8a-d936-4d62-9859-6d8e98dfdabb · inbound

Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models cites this paper.

Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:32:36.785476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:32:36.785476Z digest=sha256:d82fe4227bb6077dc143acb4f0a7ed8bce7f3c380d56a7aac82c599ecf76f7b6

Observation 68bbd057-2773-4fa5-a82f-ded588030d21 · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.463921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.463921Z digest=sha256:ffd970ada7251b2d1cd99fb830cece9ef561352660d7cfc0a485874315973c67

Observation 2a1e138e-84c0-4cc9-bea9-3ed10f1220bd · inbound

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control cites this paper.

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:32:49.534799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T02:31:07.932802Z digest=sha256:a7542b4280899cfed105ebef2bd9cb495916152921a087a55267c100e262ac69

Observation 816d3d25-ddd9-4c28-bbf9-050b72c14ae2 · inbound

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs cites this paper.

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.593672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T08:34:14.310656Z digest=sha256:d78d670dea287456a329b966d7cbb8fd6e9d7ff43d5ab4fbcb818d5ab46eb535

Observation ca0342c1-706c-46aa-b8d3-f0f377628d56 · inbound

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes cites this paper.

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:10.260228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T11:42:16.090162Z digest=sha256:bec56c301ad13ab5d7784a36c74e5b3bc0941fdb067a1cc55ac43b0f0f03e9ca

Observation 3ca91b54-05de-4028-90bd-1be1c889cbd0 · inbound

From Mechanistic to Compositional Interpretability cites this paper.

From Mechanistic to Compositional Interpretability Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 148

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:46:18.305880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T02:42:26.173782Z digest=sha256:fad2ab104f799a9d71c1c33d03830fc85bc24f3176ebd3694d0fa12b5c2d59b2

Observation dc656d93-f134-406a-bc8d-c3e8ffb42b30 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:54.794369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:580a107906bc8a9b95eb6ec0266017992d4729b888ef266387d70698014fdbd7

Observation 66f400a5-dc44-4b54-8dde-ab0f3ec44b19 · inbound

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models cites this paper.

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T19:02:31.790237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:02:31.790237Z digest=sha256:524dceb75bad6e6cdfe7cc10047ae2ec6d40d16db1783918a0dac3e94f187524

Observation 9d75ec05-917e-428b-b71f-ce40c9a25fc6 · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:16.527795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:16.527795Z digest=sha256:0bd1af1d9cada780e055af9f5f1a96f1c161ca6dc4644f62f62633030a28feaa

Observation 4d4552c1-0b11-4981-8e23-faf7b4de599b · inbound

LAWFUL: Law-Aligned Witness for Faithful Use of Latents cites this paper.

LAWFUL: Law-Aligned Witness for Faithful Use of Latents Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:47:34.695081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:47:34.695081Z digest=sha256:2d96d69f6b996bafca0dad2f6f55aeb19805ec85d3ae4724f0a2d621de5a7238