Pith. sign in

Paper Citation Record · LEDGER

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2301.04709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.04709 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:44.215403Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

10
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f2e48289-ffce-497a-9a7b-aa5db8fe4d98 · inbound

Localizing Model Behavior with Path Patching cites this paper.

Localizing Model Behavior with Path Patching Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:38:37.803414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T19:38:37.751487Z digest=sha256:fb3ddf7f8126575a03a09f48b9cacc871c52a3d2dcf6109f534b4f56c504cbcd

Observation ac85d2ad-2034-4509-b7ff-a108d493aa27 · inbound

Linear Representations of Sentiment in Large Language Models cites this paper.

Linear Representations of Sentiment in Large Language Models Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:43:06.462170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-15T12:43:06.361170Z digest=sha256:4bca151fe8540584be5ff60ba96566c743a0c32b7fe19f09e988fffab9b8ac25

Observation 072600d1-7c4a-4cc7-98f8-874fb6884ae6 · inbound

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models cites this paper.

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:15:10.703650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T13:15:10.632115Z digest=sha256:be4bbb58d5506332968f4f99e581c24cf7221ff8d17796c6ec79b7a7d581aeb8

Observation e44f166a-e856-466f-9106-7dacbddeacb6 · inbound

Factored space models: Towards causality between levels of abstraction cites this paper.

Factored space models: Towards causality between levels of abstraction Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:28:45.301776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:28:45.301776Z digest=sha256:8b9e44e421431249924fa24749aa6c85894b809f854a9ba77df88a8da685c670

Observation 07fe9ccb-2b20-411f-ba91-3906b2393cb6 · inbound

Removing Spurious Correlation from Neural Network Interpretations cites this paper.

Removing Spurious Correlation from Neural Network Interpretations Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:03:10.792600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:03:10.792600Z digest=sha256:71a1947198a9d3facc8bec14c00a24d742c6d8eb635ee601ff066af06bc80d89

Observation a94c6dd8-759c-4b98-b19c-1e391c6f6b8e · inbound

What is causal about causal models and representations? cites this paper.

What is causal about causal models and representations? Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T20:38:36.470902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:38:36.470902Z digest=sha256:e1e7852e7b6bdd53a6a21088b2c9ff6f3ec96bf4632e4fd41d46e68c93478c41

Observation 0d7cd494-b1de-4b9b-b4bd-199414ac0a7c · inbound

MIB: A Mechanistic Interpretability Benchmark cites this paper.

MIB: A Mechanistic Interpretability Benchmark Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.215403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.215403Z digest=sha256:16567a60063352c9d5962115e2dcc8a4623a3c79ff9c1bc9bcf0266dd9069510

Observation 22aea982-0987-44f7-ab5b-592ababe55dd · inbound

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii cites this paper.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.733991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.733991Z digest=sha256:e65e5de154d8073bc6fb51ac8f16087be3fbe818189e927e0d6ef869a3571bc6

Observation 831539e5-aef9-43d1-9f5a-1fa53e94a015 · inbound

$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks cites this paper.

$K$-MSHC: Unmasking Minimally Sufficient Head Circuits in Large Language Models with Experiments on Syntactic Classification Tasks Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:57.678412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:57.678412Z digest=sha256:161ef8c4620d409285f91a113f5d0bd89bb8841d03b4be6e868eac5ef108d784

Observation 60e2b358-386b-40d1-b04e-197bf857f57f · inbound

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors cites this paper.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.583205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.583205Z digest=sha256:0f05f9545993b69d92f3364e2f816bad5e090c04707806e03ba39136eb6ab306

Observation 8240af5a-2aed-4c27-985e-c4e89b35704f · inbound

Explaining Neural Networks with Reasons cites this paper.

Explaining Neural Networks with Reasons Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:40.162144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:40.162144Z digest=sha256:28d280ffb54077f184ef5dbda50fa96cabb1b308e9602c42759f462521ea66b6

Observation ed653d77-400e-476f-b68f-b3871cd5a005 · inbound

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs cites this paper.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:00.147887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:00.147887Z digest=sha256:f54a747b9d0e1d3e03d7ba86ac39e08a374c30a65a3004c94e28d4f4b09c1861

Observation 60b65955-83ca-44f0-98ca-bf7a822f87cd · inbound

How Do Transformers Learn Variable Binding in Symbolic Programs? cites this paper.

How Do Transformers Learn Variable Binding in Symbolic Programs? Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:07.811798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:07.811798Z digest=sha256:aa39603a3b5e4c20c0387e9a8a1239058efb48c75510ab6edeb14b98866ae535

Observation a99928f2-fe36-4677-b909-72e3eae049fc · inbound

Identifying a Circuit for Verb Conjugation in GPT-2 cites this paper.

Identifying a Circuit for Verb Conjugation in GPT-2 Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:46.623506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:16:46.623506Z digest=sha256:6d31652cb48b6574b3b8dc3dda454eebdcb8c86747455ea5808615a04d5924ed

Observation 88fa77e4-1605-455c-a09d-0b79149d539c · inbound

Identifiability in Causal Abstractions: A Hierarchy of Criteria cites this paper.

Identifiability in Causal Abstractions: A Hierarchy of Criteria Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:22:28.725891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:22:28.725891Z digest=sha256:03294ba19e2797ae15c8689433e35d22b1b1135772221f2d11df10fa85fc7609

Observation 0e43d9ef-db7d-46cd-aa37-b05d7851210e · inbound

Can Interpretation Predict Behavior on Unseen Data? cites this paper.

Can Interpretation Predict Behavior on Unseen Data? Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:28.638962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:28.638962Z digest=sha256:d94f16872d4fa516818b4bfbd0227a88578763072cfb7ee0209aad142b7a368b

Observation 4ecdeea4-ca1e-4230-8825-1855bc0ca1db · inbound

Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies cites this paper.

Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:51:37.348803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:51:37.348803Z digest=sha256:8476d18311bb4dc8d72a7b738e0e6aa7d39d5532a08f2cb53c49ba6eb01f76d5

Observation d7dee443-348d-4baa-98ad-9c9fe9f2a572 · inbound

How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding cites this paper.

How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:12.926237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:21:12.926237Z digest=sha256:0f978325a49a295dc76649d3c9daeca1b13138e731c3ebac69c22a077c91492e

Observation 03fbbe87-c027-441e-a4b8-17876f48c024 · inbound

LLMs Should Not Yet Be Credited with Decision Explanation cites this paper.

LLMs Should Not Yet Be Credited with Decision Explanation Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:06:33.196776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T18:41:24.552939Z digest=sha256:5a35858019b79fbaa5051a97004acca0df86abd3e5366cb3ec1f292f6088589f

Observation 38219848-7677-43dd-aed7-9a0cbec2407f · inbound

Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction cites this paper.

Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:36.290694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T19:12:30.627633Z digest=sha256:0b0e7d15b36235db6b2ef6a93c8cf4d950bd7ee31d94ad67d07df6fd757a8afa

Observation 9395ec40-e26a-4326-91d3-68fa7c649b7c · inbound

From Mechanistic to Compositional Interpretability cites this paper.

From Mechanistic to Compositional Interpretability Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.264696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T02:42:26.173782Z digest=sha256:1f85a73bb6b6543650d229cce68725a12441bcc8bbf5ab023875423223118b21

Observation dcc8bb27-776b-4289-895d-4364afb81e32 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:54.151029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:01bd6e531e92236e81cfcb8040f123a74e6fc1bed8f8da37c7cd7c9cccdb04d2

Observation 540a2dd8-b6c1-4a46-b7d8-b72425f6549b · inbound

From Weight Perturbation to Feature Attribution for Explaining Fully Connected Neural Networks cites this paper.

From Weight Perturbation to Feature Attribution for Explaining Fully Connected Neural Networks Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:07:41.081855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T16:06:05.508610Z digest=sha256:df9fa8a6e2d1937964c009f1961ad274d42a19e18e4eced17819b74e49309fd2

Observation e0173868-2849-46b3-9fed-ba9a2a8aec7b · inbound

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express cites this paper.

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:13:59.318138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:13:32.378862Z digest=sha256:32a046aa8f85314cf5388c6e1c79b97bf4704d7d7cac92dd3c848f723367f630

Observation 82492bf6-f21a-4999-941f-563b442d697d · inbound

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English cites this paper.

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.895670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T18:44:23.296625Z digest=sha256:224dc27a67da432a606c5f9354cbb48844e57c20d3a37b1fdcbdc569c793069d

Observation 9dbb2565-b89d-4eed-9a79-7e1331e76973 · inbound

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English cites this paper.

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:14:20.387389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:14:20.387389Z digest=sha256:4c03828e8ed1ba9eeea5f85155e19d147fc6fe5f72cc5461e7ee31ea5383c106

Observation c934caf3-ffe0-40ac-961f-61d6a5a5394a · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.137704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:babebb28bb6ed78ebeace9e26e5e31863ca35c7abe7553a3e68e843dd37a2f2c

Observation 7b896aa9-f0cd-4ff4-8612-394cf17fddc8 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:ed9433c5d4f2ae3ff5a21ed01a5add01f010fd516ecb274de9af642bc89adb78

Observation f2455939-37e6-42cb-8548-afb76926aa3b · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.287226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:3aced7f2d7543773050caad39b2f8cf4c6a542ee4768106084f542b7600d65b9

Observation 6204ff38-c9a2-4847-8cf6-6787df2f3870 · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.074500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:60c6bee0d4ca7e299bd262bc32fe069e553579b8675a80c4f09ddde8f6de31a5

Observation 201b7313-9d69-4b6e-b126-58c5af926390 · inbound

XtrAIn: Training-Guided Occlusion for Feature Attribution cites this paper.

XtrAIn: Training-Guided Occlusion for Feature Attribution Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.382726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T13:37:35.691503Z digest=sha256:cdfa02d33b0bb0dfe6c41f91a198d535d854d9ee81e0cd4dc983d5d9a125f2b4

Observation 832bb100-a667-4f0d-adbb-bf5ed7e9d6c0 · inbound

Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks cites this paper.

Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:30.217596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T17:49:42.069456Z digest=sha256:de4375a991100c71af843708ea4251ceef1e22e86c62eae1dd3927b93acb5ff5

Observation c6fe8413-40de-4ea3-aecc-856b9dbde6f4 · inbound

Steering Vision-Language Models with Joint Sparse Autoencoders cites this paper.

Steering Vision-Language Models with Joint Sparse Autoencoders Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:07.779400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T20:56:57.716246Z digest=sha256:294de4edf521411b5ae0d94f9cceef94c62325f4d756e23b49b10b4f77edd4bd

Observation 51b25163-ad74-4382-87a1-03216aaa2ef6 · inbound

Safety from Honesty in a Disinterested AI Predictor cites this paper.

Safety from Honesty in a Disinterested AI Predictor Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.196264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T06:53:19.737408Z digest=sha256:31ee36f873d67703961b196e9e4a2a7afc9b29695ebe89d2b4eb608a88a94161

Observation 890e0878-d980-4dd5-9b5e-32e7b8e7c48c · inbound

Safety from Honesty in a Disinterested AI Predictor cites this paper.

Safety from Honesty in a Disinterested AI Predictor Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T17:02:19.189814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T17:02:19.189814Z digest=sha256:9f9b865730783a5e9f21c82e2d03900b5d6e4037327c578751a786c65257d6ec

Observation 891db3d4-d0ba-45a8-b03d-7375297297fd · inbound

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects cites this paper.

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:03:55.219054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:03:55.219054Z digest=sha256:a7d54cb353378bd9f968ee333159e740be3f99dbb317ec0c8e7744bac7d5f617