Pith. sign in

Paper Citation Record · LEDGER

Open Problems in Mechanistic Interpretability

As of 10 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 100 inbound Pith citation observations for arXiv:2501.16496.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16496 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 177 of 177 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 102 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:55:02.232925Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact8
  • verified fuzzy45
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch7

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c923f1b3-d1fb-44db-bf56-0754c2c8918b · outbound

This paper cites a is b” fail to learn “b is a.

Open Problems in Mechanistic Interpretability a is b” fail to learn “b is a

Reference 1

Resolution
verified exact
doi, observed 2026-05-14T18:30:06.027122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:d40b9ec61595041733ff99ebb90d3090063a6edd85ca4b573acdddef3c2b017a

Observation 7354f512-55a5-4d67-9115-d2a6c561d750 · outbound

This paper cites https://distill.pub/2019/activation-atlas.

Open Problems in Mechanistic Interpretability https://distill.pub/2019/activation-atlas

Reference 2

Resolution
verified exact
doi, observed 2026-05-14T18:30:05.975539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:966e7e8ccfb6a6290e04edce725a4ec4b49389a19be95960b9abd563d27e03f1

Observation 4fadc288-2e14-48ae-a8f7-9d68faa121d4 · outbound

This paper cites Transactions of the Association for Computational Linguistics9, 160–175 (2021).

Open Problems in Mechanistic Interpretability Transactions of the Association for Computational Linguistics9, 160–175 (2021)

Reference 3

Resolution
metadata mismatch
doi, observed 2026-05-14T18:30:05.981712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:14723b39d40855a2b9f94402579e3291ad0d7f553243c0c968bc3c80a8d00988

Observation ce5c2ddc-2cd2-446f-aac3-3ad238d49f20 · outbound

This paper cites In: Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP, RepE- val@ACL 2016, Berlin, Germany, August 2016.

Open Problems in Mechanistic Interpretability In: Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP, RepE- val@ACL 2016, Berlin, Germany, August 2016

Reference 4

Resolution
verified exact
doi, observed 2026-05-14T18:30:05.986398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:a5c5b4d46be212af2681b7be6add7af35d7db387789f23f12c2c0cea2ed341dd

Observation 71ac7d6e-65b9-481e-b224-208d9e84bcbd · outbound

This paper cites News from Generative Artificial Intelligence Is Believed Less.

Open Problems in Mechanistic Interpretability News from Generative Artificial Intelligence Is Believed Less

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:05.992913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:3f4a2148a4a02d1dda921156ea6ce90ed2c16899b995c116b6c33efc50582891

Observation bc772c23-b7a3-4894-8ca1-5a6433b0f8ab · outbound

This paper cites A Structural Probe for Finding Syntax in Word Representations.

Open Problems in Mechanistic Interpretability A Structural Probe for Finding Syntax in Word Representations

Reference 6

Resolution
metadata mismatch
doi, observed 2026-05-14T18:30:05.997680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:f43b016d317321ab0e26eb147df344c8e4c2bf2d04d49b4f02666113b1bbe501

Observation 180db348-902e-4fbc-9ed3-0f0f732c27d9 · outbound

This paper cites Backward Lens: Pro- jecting Language Model Gradients into the Vocabulary Space.

Open Problems in Mechanistic Interpretability Backward Lens: Pro- jecting Language Model Gradients into the Vocabulary Space

Reference 7

Resolution
verified exact
doi, observed 2026-05-14T18:30:06.002546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:b2cf37139f42ff0a19e8c0aabec265db61cb31983fb74f3211fb8bacc11d47fd

Observation 876af1f0-0ae5-4f45-a4d1-626a50d64d18 · outbound

This paper cites Understanding Deep Image Representations by Inverting Them.

Open Problems in Mechanistic Interpretability Understanding Deep Image Representations by Inverting Them

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:30:06.009807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:87268ff639fda7cb7ca3a8becf37fd58018dec5ec04ffc6732cef85355f9dcac

Observation e6c5ae18-3944-4aa7-b2f4-d55cf0c6692f · outbound

This paper cites URL https://distill.

Open Problems in Mechanistic Interpretability URL https://distill

Reference 9

Resolution
verified exact
doi, observed 2026-05-14T18:30:06.016148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:6346c3399a3691914704898ac680b460735a3b8045e657903e3fc32bed4001ee

Observation 7861a7dd-40ac-4ba7-b676-e12f1258afe4 · outbound

This paper cites Does Transformer Interpretability Transfer to RNNs?.

Open Problems in Mechanistic Interpretability Does Transformer Interpretability Transfer to RNNs?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.053221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:ad6a34626458295f370a7799b07c7aadb5211c4aa7c33a3aca118ae9b76c9ae3

Observation 20d8ff8f-dcdc-44d5-b6fe-d6e676ea4ca5 · outbound

This paper cites Weight-based Decomposition: A Case for Bilinear MLPs.

Open Problems in Mechanistic Interpretability Weight-based Decomposition: A Case for Bilinear MLPs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.059231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:1ef51fe30cf2806c2d80c95a6bfad32ca9f9f79a8c836542a95407bd16ddf5bd

Observation 733c560a-3e06-47fb-b145-26325ebdae61 · outbound

This paper cites Jorge Pérez, Javier Marinkovi ´c, and Pablo Barceló.

Open Problems in Mechanistic Interpretability Jorge Pérez, Javier Marinkovi ´c, and Pablo Barceló

Reference 12

Resolution
metadata mismatch
doi, observed 2026-05-14T18:30:06.031452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:4606b99786e82d200879d91f86d67231b0927d70be7542058138b04212f681f9

Observation 70f8a027-b816-4dcb-b739-657179aa7612 · outbound

This paper cites Dynamic models of large-scale brain activity.

Open Problems in Mechanistic Interpretability Dynamic models of large-scale brain activity

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-05-14T18:30:06.036048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:61946333dd1b885cbd7a34e415e3d7df2dfc93ce5fbc1c28fe298dd59b30fab7

Observation ed55669f-b1a4-48aa-8426-aafe99b3f707 · outbound

This paper cites why should i trust you?.

Open Problems in Mechanistic Interpretability why should i trust you?

Reference 14

Resolution
verified exact
doi, observed 2026-05-14T18:30:06.040799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:f5210db5d9f03f1a6731712feef74f71e6837f20499cdfa6ce1103da739eee32

Observation 070cb7c8-af62-4070-8a51-0c1eda90f730 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Open Problems in Mechanistic Interpretability Gemini: A Family of Highly Capable Multimodal Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:30:06.046857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:10952f82dc0aae5625f7216e57153e8d856e30fd7eca37cbb0a26e04ce2dba43

Observation f996f9eb-11eb-4278-843d-81c8ac638337 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Open Problems in Mechanistic Interpretability Representation Engineering: A Top-Down Approach to AI Transparency

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:30:06.022244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:9cdd712b2c74ccd191231dc3640c0565df345b0cec8512d8277dff41b282a4ce

Observation 54952ac5-de83-4b8c-a063-a89db4407a75 · outbound

This paper cites What isomorphism or what approximation of a neural network (or parts of it) is the best way to express it for the purposes of interpreting it? b.

Open Problems in Mechanistic Interpretability What isomorphism or what approximation of a neural network (or parts of it) is the best way to express it for the purposes of interpreting it? b

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.280695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:99acb3fce348184a731a7687635529ced5ea4870027484d2eb2411ad8d079ef7

Observation da960d65-d177-4de4-9691-917c2db17768 · outbound

This paper cites To what extent do models encode concepts linearly in their representations? b.

Open Problems in Mechanistic Interpretability To what extent do models encode concepts linearly in their representations? b

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.285689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:7c8a975256c60b866c2bc0bb920b4ba57d7d56bb0a18960747b6cb02695a40db

Observation 7142ca1c-dfb2-4d16-a8de-fde20845a135 · outbound

This paper cites Can we fully determine the causes of feature superposition and polysemanticity within neural networks? b.

Open Problems in Mechanistic Interpretability Can we fully determine the causes of feature superposition and polysemanticity within neural networks? b

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.290636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:6e9498dca400a480b2ebb9549ab157673afe2791c5829854421f05d6046ef954

Observation d58e9275-d820-483a-9da2-2f3143118ea8 · outbound

This paper cites What lies in SDL reconstruction errors? Will the errors converge to zero with methodological progress? b.

Open Problems in Mechanistic Interpretability What lies in SDL reconstruction errors? Will the errors converge to zero with methodological progress? b

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.296069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:4c4acb1518ad040974dc1f40a1c99528116a9594a36e32cb90b80bc3cbcd4aa8

Observation a2319578-15f4-401b-bac8-e2c4337e0c99 · outbound

This paper cites How can we identify the underlying functional structure of networks (which defines why acti- vations are located in particular geometric arrangements in activation space)? 75 b.

Open Problems in Mechanistic Interpretability How can we identify the underlying functional structure of networks (which defines why acti- vations are located in particular geometric arrangements in activation space)? 75 b

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.300906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:caad7b973b362f7820993a99b630cef0a22403fdca74c2aee0045a95c204c0dd

Observation 8d551ff8-d059-43de-b77e-ffff049af662 · outbound

This paper cites Can we distinguish parts of networks that underlie generalization from parts that underlie memorization? b.

Open Problems in Mechanistic Interpretability Can we distinguish parts of networks that underlie generalization from parts that underlie memorization? b

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.305547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:ad102de632100d48821a47dd80ee8a2924df8d57b1179373bda0fc6ac29d8033

Observation 36d6b044-49c1-4eb6-80f3-1e6deba5f818 · outbound

This paper cites Interpretability training: Can we train networks that are interpretable by default at low per- formance cost? b.

Open Problems in Mechanistic Interpretability Interpretability training: Can we train networks that are interpretable by default at low per- formance cost? b

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.310467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:db21b3edda939cbade2cac3cf2eb3f040998a436a200eb6c008679f98801863f

Observation ff490481-0010-4bed-bd44-30fa7397fec4 · outbound

This paper cites How can we avoid imposing human bias to explanations? b.

Open Problems in Mechanistic Interpretability How can we avoid imposing human bias to explanations? b

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.315126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:2f27968ef8b7e66209cdd2a720bb8593908c94c3596a05022b79fdcab6e1a503

Observation 9e7a2ef4-9d63-44c9-a977-cedc8afe1475 · outbound

This paper cites How can we develop attribution methods that capture higher-order effects beyond first-order approximations of model behavior? b.

Open Problems in Mechanistic Interpretability How can we develop attribution methods that capture higher-order effects beyond first-order approximations of model behavior? b

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.319803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:9147f99e1815db2fde332c27d69195c14bae9c3399e7131e0d3b5cae7a461ed5

Observation f48a12fc-9794-471d-9c2b-792212f06c56 · outbound

This paper cites Hydra effect.

Open Problems in Mechanistic Interpretability Hydra effect

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.323908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:175cbe9d7b4f3a70982590449e0c277b6f2d6d2b1df736e9bc30700ba29bdb6a

Observation fd6a03bd-378f-4e78-89f1-3f6e4a18d578 · outbound

This paper cites Can we improve on methodologies for evaluating hypotheses through their predictive power on activations of network components? 76 b.

Open Problems in Mechanistic Interpretability Can we improve on methodologies for evaluating hypotheses through their predictive power on activations of network components? 76 b

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.329061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:55853bcc485c33957668779c5fdbc9f46e170666aa6870adaba5ecbd0c907f37

Observation dd7c3def-47e8-4f2e-bb90-b3d64c780245 · outbound

This paper cites model organisms.

Open Problems in Mechanistic Interpretability model organisms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.333830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:712855c155eca6afd2577fb67651a846a1abd9af6316d87aa14490bc40e4fb85

Observation 14e69a69-563f-4137-8d83-523996405b29 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.338193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:4bbe68223654488074d8b068662efc1fcb48a98857879b7e6a211c375d5af32c

Observation f398f711-81c7-4a32-8a75-d7607e710b0d · outbound

This paper cites stress tests.

Open Problems in Mechanistic Interpretability stress tests

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.342427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:cb2d58391aa2401cc1575c3a212d213987b3ef07de05e6c0074888eeb718e2fa

Observation 5755e2ed-f3ec-4928-9301-fafb589f72b5 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.346926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:e12621db6e55f1573215c90813b2220501ab6cfad8c0f6c4e036e85ac6aeb4aa

Observation 3d09aa42-1993-4b2d-b114-1e85c75fbec5 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.063177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:bbc704cb2192d0d1b1abe2fa0e80d67cbfca5d0c7db2d50a71a27fb3c63e42ea

Observation 78a8fd19-d816-4ed1-b2d3-11a009d9c3d4 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.067203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:7c8e1e55b8d1bb9491cb619c92a9116d190a5b98abdc880b7d5cfbc897a7637c

Observation 922da718-ba11-4302-a220-746b459bc974 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.070676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:1a5ec77fa2629480b81101d0cec8de1da336ca0092dfcda40505578173d8df15

Observation 11234bc9-0d64-4d71-9d7f-4beef967e464 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.074272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:d3760ec867af1bcebe41596f99c7582ad05f64a2a201acb1abac9370053ecd16

Observation b50fc385-f7fe-48be-8db6-451b216e4dff · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.078001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:1b820fa7c969a86df77a0d8640f9c3791ad432a2ef71c58decf763f36abf5fbd

Observation 637940f9-a503-4482-90cb-3e5c80dc51a3 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.082145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:ef5acc9f1d19a0130dce11a0a693adf874847fa85b59d011befa368fedcc28ff

Observation aa50ed94-92f6-427e-8162-4156ae230790 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.086866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:c379bc8b61a652fe0c479b5758ee7aafcde6431d875e39e023d01f0574a875a1

Observation 8bf28823-4860-44f1-84e3-35126cae0eb5 · outbound

This paper cites Through automating the generation and testing of arbitrary hypotheses? b.

Open Problems in Mechanistic Interpretability Through automating the generation and testing of arbitrary hypotheses? b

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.091151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:69bf4bc511f0d07eb645873be27840b63ae21779e8b05bdf8afb5a4f8867defe

Observation 0ded2bb9-d480-4386-aca4-edd1ff675b13 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.095248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:6260c873a41d1bcacaa7f5d94a0a64fb75914280f46b1fb4e39c79d252bfd61f

Observation 52823d66-ce85-4cc0-b820-d9af4f453ed1 · outbound

This paper cites Conceptual interpretability research? b.

Open Problems in Mechanistic Interpretability Conceptual interpretability research? b

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.099447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:565d5c8564718160cd61a7c544d18bb3f80abac93179395710c562dd0716c90b

Observation c3a19264-694e-4037-a5b4-474e804a300f · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.103730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:5b6edde4980f22b1bb5c5e5c2e5a6cb4f43d90b72a3a47b067ffd7d424dc40ac

Observation 9ecbcb60-0529-4e4c-a37d-a64ba61aba85 · outbound

This paper cites white box.

Open Problems in Mechanistic Interpretability white box

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.108001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:07ea7e3911a0b885b2fdbd236ef00f53a41bbf81be48cf2a6e146ba8f0f798a1

Observation 31489b22-e6cb-461c-a85a-8a970a972393 · outbound

This paper cites Can we use interpretability insights to make red-teaming more efficient than current methods? b.

Open Problems in Mechanistic Interpretability Can we use interpretability insights to make red-teaming more efficient than current methods? b

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.112199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:fb8b5d58486ddeaad27a0b29a4c28a9f2e1820f2205a52a40eb6be1198f782d6

Observation f13ced3f-1e9e-4445-be0b-8dc8c21ce751 · outbound

This paper cites Can we get mechanistic anomaly detection to work? b.

Open Problems in Mechanistic Interpretability Can we get mechanistic anomaly detection to work? b

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.116859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:0f7ec05848e8f4c4cbf3ab23a123112a0df8c92843effeb1559b3e19417b3331

Observation 0f10da81-f714-4635-b378-ae44210a6ada · outbound

This paper cites How can we make activation steering more precise and reduce its side effects? b.

Open Problems in Mechanistic Interpretability How can we make activation steering more precise and reduce its side effects? b

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.121134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:02a3a278481895086c12166622823329d92f560ad813390532872a414d98a33f

Observation 8b6d10ce-c5c2-4d7a-bf3e-aa3687922b73 · outbound

This paper cites Will carving the network at its true joints help us improve on model unlearning and editing? b.

Open Problems in Mechanistic Interpretability Will carving the network at its true joints help us improve on model unlearning and editing? b

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.125782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:4e9a8a488970f6c5eafea0060693d848d7ce9cf09eeffafe8493bcfc188dfbc8

Observation e6a5ebc1-2fab-4852-b5b2-bd4e527cb3d3 · outbound

This paper cites Can we make finetuning more sample-efficient by targeting specific parameters? b.

Open Problems in Mechanistic Interpretability Can we make finetuning more sample-efficient by targeting specific parameters? b

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.130380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:7bca0469a00abf9736a600d9c527e77484ea66a0616cbde9f707dc333070ddd2

Observation f5db5b02-b6a8-48cf-b942-2c704909d5a9 · outbound

This paper cites values” or “goals.

Open Problems in Mechanistic Interpretability values” or “goals

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.134966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:d83df6cf38d1afd1f0773fba242ff7e6e9797f64cbb7269b6c8f9b361f4ae153

Observation 2a194e14-6273-4935-a009-8d463633ad6d · outbound

This paper cites Can current toy model verification approaches scale to frontier systems? b.

Open Problems in Mechanistic Interpretability Can current toy model verification approaches scale to frontier systems? b

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.139686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:cb63aba31455eb8323378a964106ebd759803adb6bab3cb3e5bff5d1e299191c

Observation fa1f6919-6fdf-4a9a-89e5-a397efe4dd9a · outbound

This paper cites enumerative safety.

Open Problems in Mechanistic Interpretability enumerative safety

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.144570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:cb3eb75c9eebccca63b2d8db84cfeb1b300eedc3f727c3e2c75961306783303b

Observation 728fdd09-3a9b-42b8-a986-10b2fb668575 · outbound

This paper cites Can we identify early signatures that predict emergent capabilities? b.

Open Problems in Mechanistic Interpretability Can we identify early signatures that predict emergent capabilities? b

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.149503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:d33d7e8d4943140ec89e2445ff2ac7902299de3da48a01d5e52503241013a484

Observation 6b570a64-7e75-431d-9ed9-27766a82bba9 · outbound

This paper cites How do specific training examples influence the development of model mechanisms? b.

Open Problems in Mechanistic Interpretability How do specific training examples influence the development of model mechanisms? b

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.154069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:5e1ecbc04a87506ea8c3410a1b07ff64a70962e2e27b08260a4a1a6b8ad6d50a

Observation 13bb5b87-ec8d-4a56-9fca-436aca87e015 · outbound

This paper cites How can we identify capabilities that could be ‘unlocked’ through prompting or finetuning? b.

Open Problems in Mechanistic Interpretability How can we identify capabilities that could be ‘unlocked’ through prompting or finetuning? b

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.159176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:1031fcc3fa2fd8eb1b019dbd392ad8c405aed507ae052950297c04eb196e18c4

Observation 37a55358-fc99-42e2-a397-27f2d166fddb · outbound

This paper cites How can we identify skippable computations without affecting outputs? b.

Open Problems in Mechanistic Interpretability How can we identify skippable computations without affecting outputs? b

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.164006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:3c0c90dabed5fe4c63aeb5701d0b889818e35111de9f47af17ba9d09d2194092

Observation 8b60cb08-2fab-496f-8d13-6a95b5940da0 · outbound

This paper cites Can we better select training data by understanding example influence? b.

Open Problems in Mechanistic Interpretability Can we better select training data by understanding example influence? b

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.168949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:dbc1cf34a09739ad3e83f356cebbe28805e013afcf76ae13c5d4a0826146ad02

Observation c7249d5b-baab-4110-aa85-0493b5c0c1b0 · outbound

This paper cites Can we design better inductive biases based on mechanistic insights? b.

Open Problems in Mechanistic Interpretability Can we design better inductive biases based on mechanistic insights? b

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.173860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:6643ee9707df30bf67f31eca9f95f6855dbd12c765b853d0382a08ec316295ef

Observation f4d30f0b-e153-498a-b3d3-3ade9cc32070 · outbound

This paper cites How can we extract novel patterns and predictors that models have found? b.

Open Problems in Mechanistic Interpretability How can we extract novel patterns and predictors that models have found? b

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.178610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:76ec9741bc4be883f7fd5f419fdeb94fe61e82173e2a57f2760ef886bfccaf45

Observation c8b34bf7-027d-47e8-bd4a-a014a73240a9 · outbound

This paper cites How can we detect when models have found genuinely novel patterns? b.

Open Problems in Mechanistic Interpretability How can we detect when models have found genuinely novel patterns? b

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.183194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:0096aa0447ba0e8103f9906844ec0e2d0a65668d512691914e231b2ee2dab327

Observation 33032b8d-fdb0-466b-b15d-49507245f6df · outbound

This paper cites Do current interpretability methods (SDL, circuit analysis) transfer to SSMs? Or, like the transition from CNNs to transformers, are new approaches necessary? b.

Open Problems in Mechanistic Interpretability Do current interpretability methods (SDL, circuit analysis) transfer to SSMs? Or, like the transition from CNNs to transformers, are new approaches necessary? b

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.188174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:2f163c9e62d2fd4f80c73faeaaabb14c4034a7af72500c34eaec8c3e7e717d5b

Observation 69a1577f-00cd-4909-9ec1-89990f4a40b1 · outbound

This paper cites universality hypothesis.

Open Problems in Mechanistic Interpretability universality hypothesis

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.192701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:e5e99f8d1ca4ebf9007817f1223d8e77195847e75153cb2dd808032b706285ae

Observation 4e14c264-7a9e-4bce-8376-98e1ac11da50 · outbound

This paper cites How can we prepare for interpreting novel architectures? b.

Open Problems in Mechanistic Interpretability How can we prepare for interpreting novel architectures? b

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.197493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:287571fc4c6a05094b84ef98853e987ce6ff6988fedddac95d3a30a80facd712

Observation c22ca0a1-8991-4a95-9a91-e92da7086776 · outbound

This paper cites How can we visualize model internals in an intuitive way? b.

Open Problems in Mechanistic Interpretability How can we visualize model internals in an intuitive way? b

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.202163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:75ce9d78e9c903d3ea8a255d032951d33e00bedac080ba1ee9b9fde1490df162

Observation 82fb5570-d563-41fe-88ef-ba5c209fc2a3 · outbound

This paper cites How can we help auditors find potential failure modes directly? b.

Open Problems in Mechanistic Interpretability How can we help auditors find potential failure modes directly? b

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.206840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:34285de25da945cca458cc89b4db30780b7cd9449860230707ba8115fd511184

Observation f436bcc4-edee-4330-beb8-e88787400224 · outbound

This paper cites How can transparency features help users calibrate trust? b.

Open Problems in Mechanistic Interpretability How can transparency features help users calibrate trust? b

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.211425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:a9361efe6d930f2b54cb62fbc109735a19a5d3a2496438aeda0529fa4f03bda7

Observation 0d343568-e0b4-45fa-8cf9-4d0c8cdc4f94 · outbound

This paper cites Can we identify specific mechanisms that caused AI failures? b.

Open Problems in Mechanistic Interpretability Can we identify specific mechanisms that caused AI failures? b

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.216371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:222f99627c4437ee4e15db1dbc77126bbe4a4920142e3b2aeaa1fd439f3aef0e

Observation 82df408c-34da-491c-b04a-a1ac5813ebbf · outbound

This paper cites Can we identify mechanisms responsible for specific dangerous capabilities? b.

Open Problems in Mechanistic Interpretability Can we identify mechanisms responsible for specific dangerous capabilities? b

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.220973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:189adf6eb6879a560417682dc40e98a70d82081a25a1d11953adf8993148b64f

Observation d8058ef9-d43e-4b2b-9631-ea40fe3ceee4 · outbound

This paper cites How can we trace decision mechanisms to explain model outputs? b.

Open Problems in Mechanistic Interpretability How can we trace decision mechanisms to explain model outputs? b

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.225663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:bab9d79aaacad009a8234e640e1c462db46a65922f01e10363ef1f5ea851e4be

Observation bc0c733c-5cbb-4c36-b777-d2b2dcbec6c3 · outbound

This paper cites How can we use interpretability to improve capability elicitation? b.

Open Problems in Mechanistic Interpretability How can we use interpretability to improve capability elicitation? b

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.230699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:559c22299bf88d5c78ccba1de1ba1ab0a3c2355990a5c2e91896f973842a0953

Observation f2b53f9b-7b37-4726-9e0c-6494fb436bb7 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.235543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:55482ab41edd3c60f3f73a952cc89f582bc611b0ab32e0e6c0ef2077d898b93e

Observation 206c09bb-3a8c-48d7-94ba-404764a8ce0c · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.241536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:79b0ed31c241abb161e64e19de1b37173822dcc09599b1a292ceadd47ba12846

Observation 3809b625-05f9-48e9-8ca7-ce838f0fa2e8 · outbound

This paper cites Can we use interpretability to construct reliable test-time monitors to detect AI incidents? b.

Open Problems in Mechanistic Interpretability Can we use interpretability to construct reliable test-time monitors to detect AI incidents? b

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.246817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:2d0e31290dd9558b9b4d72c9e0bd65ba897c487130137820705392a8277db1bd

Observation ac916f48-a562-4c95-8c4e-9a524698ab0a · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.251798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:362ac588d12a8b692346f8ee39036eab97fa32a84f4173a0f9f530bee496a489

Observation 10336af0-81e7-4135-b1ee-ddffd0bd86e8 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.258796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:a7a1c42c23208ebf05d14a478639bb3e9c3571450f5316746f953d603daa5bb8

Observation 874cc514-7cfe-4dde-a966-281ce07afc48 · outbound

This paper cites an unresolved cited work.

Open Problems in Mechanistic Interpretability Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-14T18:30:06.264148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:ea4e61c2766d7fdb86eaf50e7481ff39525425ad5097d63aa6f91883dbc9df5b

Observation 0ef76cd7-ad95-446a-9abc-d601c6cc535e · outbound

This paper cites What are the goals of the field? b.

Open Problems in Mechanistic Interpretability What are the goals of the field? b

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.269607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:04646fc14fd43b5ce901f97cfd3eabacb78ee386db1645c22145cf22c9f8720e

Observation 0594a943-fa8a-4b6e-9744-2a116cc5c8cb · outbound

This paper cites How can we communicate the results of our research such that the risk of their misuse is minimized? 82.

Open Problems in Mechanistic Interpretability How can we communicate the results of our research such that the risk of their misuse is minimized? 82

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T18:30:06.275564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:30:05.914095Z digest=sha256:ae8aa96b454e448892032f4ae835433c83d3d7855437239f194de44df28b4b21

Pith citing papers

Observation 5d1fdc47-bc03-46b5-b430-4f90cb9c1ea1 · inbound

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition cites this paper.

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition Open Problems in Mechanistic Interpretability

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T14:55:02.232925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:55:02.232925Z digest=sha256:876aa897810bd27f4c8d7bf0e2a2fbbb055a80d2b68a304fa2d17711496323a9

Observation 75c4e53e-9a54-450a-a2fb-d33a00ac8c4b · inbound

Low-Rank Adapting Models for Sparse Autoencoders cites this paper.

Low-Rank Adapting Models for Sparse Autoencoders Open Problems in Mechanistic Interpretability

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T20:18:31.107491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:18:31.107491Z digest=sha256:287aac954554519878c89cedfdf171cb5fab8c99c0d05cf74fff8e6df6c8db02

Observation 190ff932-4df3-4855-b830-d3c10bd7b4d1 · inbound

Position: Scaling LLM Agents Requires Asymptotic Analysis with LLM Primitives cites this paper.

Position: Scaling LLM Agents Requires Asymptotic Analysis with LLM Primitives Open Problems in Mechanistic Interpretability

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T11:29:21.292467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:29:21.292467Z digest=sha256:0e3a59ed3d4bb57b1c7d8715f331135964f5f80a25f185ccef775a8db0d5f822

Observation 9da5522c-3365-4982-b9da-49247b4c1bf2 · inbound

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation cites this paper.

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation Open Problems in Mechanistic Interpretability

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:55.034232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:15:55.034232Z digest=sha256:45e11aa6827dade6fb336631f58a3c4c334ffe1cea63640ab7cbc4ad08971ed1

Observation c5ab5823-34e2-4239-b5eb-d2abb76398a7 · inbound

(How) Can Transformers Predict Pseudo-Random Numbers? cites this paper.

(How) Can Transformers Predict Pseudo-Random Numbers? Open Problems in Mechanistic Interpretability

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:14.806991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:14.806991Z digest=sha256:a87fca3c4279b95a0c85e78521cbec334881b124a6e69d79d91cc7e7a3b56c19

Observation 17d25d3c-4e7f-4c85-a8bf-1e781629354c · inbound

Explaining Neural Networks with Reasons cites this paper.

Explaining Neural Networks with Reasons Open Problems in Mechanistic Interpretability

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:42.738291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:42.738291Z digest=sha256:344d7915ce902e5e734f792f5b794bb4d66a7cbcbfa7e8f8f55ec8a7416bda60

Observation b71ef92c-986b-492d-bf4c-c50c5a7a7587 · inbound

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering cites this paper.

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering Open Problems in Mechanistic Interpretability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:21.825428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:30:21.825428Z digest=sha256:fc8122ab3e664ad70bb42c7fe5e891aa5877239a39ee228f0176f40d1cb5d555

Observation 3da134b9-22d9-4cf2-a558-34b803115ff3 · inbound

Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models cites this paper.

Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models Open Problems in Mechanistic Interpretability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:43.200839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:43.200839Z digest=sha256:7c384ea702532545f6c260529457c8ecc70ed30cc2325a87c12961290f4aa0a0

Observation 30b1ff91-1129-4565-af17-7658331d90a5 · inbound

Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks cites this paper.

Uncovering a Universal Abstract Algorithm for Modular Addition in Neural Networks Open Problems in Mechanistic Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:54.459356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:54.459356Z digest=sha256:f9b956c7ff2711f9968276d49d56a6af05783c8e0e361f7cdd097c623805fb70

Observation 4591e58e-693c-4a8d-b7f6-dc5599d4d472 · inbound

How Do Transformers Learn Variable Binding in Symbolic Programs? cites this paper.

How Do Transformers Learn Variable Binding in Symbolic Programs? Open Problems in Mechanistic Interpretability

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:50:09.978542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:50:09.978542Z digest=sha256:641c64d4993f74d57f37510121f86a3c4813a92d722046daacbb9ad67421f736

Observation 44b7442f-32c0-4d6c-94e5-e38f4619f4be · inbound

Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling cites this paper.

Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling Open Problems in Mechanistic Interpretability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:41:57.279860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:41:57.279860Z digest=sha256:3de629e3158a8b2d9c696c060aed9dfb7461fd0a8c9a64284fd1a60649038126

Observation e2f85038-48b3-4789-9cad-ea89136df48a · inbound

Sparsification and Reconstruction from the Perspective of Representation Geometry cites this paper.

Sparsification and Reconstruction from the Perspective of Representation Geometry Open Problems in Mechanistic Interpretability

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:11:09.340402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:11:09.340402Z digest=sha256:05ccf991dcab85b991f59ea6d79232cbb5930eb4e0036540cfed7e83c1d41e55

Observation 919f7776-77bc-4d98-9e36-159147c403cf · inbound

Zero-Shot Vision Encoder Grafting via LLM Surrogates cites this paper.

Zero-Shot Vision Encoder Grafting via LLM Surrogates Open Problems in Mechanistic Interpretability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.455799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.455799Z digest=sha256:ee77f24348e726711bb5d75a24f5c9eb2aa6fad52bb6bb15a090460a0bcd5e81

Observation a4fccfa1-52b1-4557-94cb-cab9a1e8d023 · inbound

Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval cites this paper.

Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval Open Problems in Mechanistic Interpretability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:00.227957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:00.227957Z digest=sha256:ba7a2c90ce333fbd24de9f2e0cd3a0938a59da9b00a7045101cb74373373b739

Observation 9e8a247d-628d-4423-89f8-812a38e492f9 · inbound

Risks of AI-driven product development and strategies for their mitigation cites this paper.

Risks of AI-driven product development and strategies for their mitigation Open Problems in Mechanistic Interpretability

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:58.250420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:58.250420Z digest=sha256:9997bc68ee832219a60dd40ba1c58f56f4da7bd52dad7ccaa8ed209856e74d23

Observation 8c918ca0-37ad-4364-b348-98a8cb82979e · inbound

Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey cites this paper.

Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey Open Problems in Mechanistic Interpretability

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:20.241685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:46:20.241685Z digest=sha256:afe793826d68cdeaa199a22893f0266a5eb7cb8ddca875936b0c58b86e15b380

Observation 29a830c7-73c5-455f-b948-c73a26e39bb3 · inbound

On the Fundamental Impossibility of Hallucination Control in Large Language Models cites this paper.

On the Fundamental Impossibility of Hallucination Control in Large Language Models Open Problems in Mechanistic Interpretability

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:54.785925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:54.785925Z digest=sha256:239854bdbda15af9f917cee7d44bf8f2d99dd9a8116063515c560af2e7dbacc2

Observation 47020588-307b-4faa-a301-657d98bd9649 · inbound

InverseScope: Scalable Activation Inversion for Interpreting Large Language Models cites this paper.

InverseScope: Scalable Activation Inversion for Interpreting Large Language Models Open Problems in Mechanistic Interpretability

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:34.860460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:34.860460Z digest=sha256:f9c30203cce7b2606de3c8e4faedf1415c16b3f8160deb2d0ea7046bf03e690c

Observation 54ecaa4c-e899-43b5-925b-20aa394b7bd5 · inbound

How Visual Representations Map to Language Feature Space in Multimodal LLMs cites this paper.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Open Problems in Mechanistic Interpretability

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.977614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.977614Z digest=sha256:5e67fc1af94fe11722d3c1eaf6e220577965c306fb0a7046cbc805f3307dd547

Observation 5584384c-eafb-41cf-9f8c-642c945093dd · inbound

Because we have LLMs, we Can and Should Pursue Agentic Interpretability cites this paper.

Because we have LLMs, we Can and Should Pursue Agentic Interpretability Open Problems in Mechanistic Interpretability

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:21.495743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:21.495743Z digest=sha256:ec75fe862585164e9deeef492b7d5c48fbe91bdfbe1924e7a18eaccf7e2e45a1

Observation 40e474c2-869f-4a60-8bac-dba478f0581b · inbound

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) cites this paper.

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) Open Problems in Mechanistic Interpretability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:57.320407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:57.320407Z digest=sha256:54c28c59bd083585b3b0c405b02f43c5ff9b552855637d2ad45e1e71220a80b4

Observation aedf96a0-5269-409e-a0c7-9c25cfb7b7f0 · inbound

Mechanistic Interpretability Needs Philosophy cites this paper.

Mechanistic Interpretability Needs Philosophy Open Problems in Mechanistic Interpretability

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:50:47.297873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T23:49:19.683025Z digest=sha256:b7214c49ca1376d9c90bb13a4f9e73a892a201059b37606d9ac4c85d54c07d01

Observation c51c0ec3-ae89-4f7d-9ea1-328ace5d5dc0 · inbound

Stochastic Parameter Decomposition cites this paper.

Stochastic Parameter Decomposition Open Problems in Mechanistic Interpretability

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:38.733150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:48:38.733150Z digest=sha256:4949da650f078431991b95ffad37dbe2eda907aaa7fea9a52e3002d11d625720

Observation 1dd42544-8ad2-47ff-8e75-92947e5ea738 · inbound

Position: Use Sparse Autoencoders to Discover Unknowns cites this paper.

Position: Use Sparse Autoencoders to Discover Unknowns Open Problems in Mechanistic Interpretability

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:33.494081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:33.494081Z digest=sha256:59598817108e42b8be24a0ca2646606f7519548af2e403d1f59e13be3d94ed94

Observation cb795f77-cfc4-436b-8d93-9af6e3b3d8d9 · inbound

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence cites this paper.

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence Open Problems in Mechanistic Interpretability

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:15.531365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:15.531365Z digest=sha256:c03e541f1b47aff871afb2ec765d3e430253267c680effc7d92f55b86908b0d4

Observation 414bdef2-d59a-4fa1-b464-57ccf9deb578 · inbound

CytoSAE: Interpretable Cell Embeddings for Hematology cites this paper.

CytoSAE: Interpretable Cell Embeddings for Hematology Open Problems in Mechanistic Interpretability

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:54.104626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:54.104626Z digest=sha256:2b67f09dfb849115d202ccd44b435ea4ead6358bbe704e28e0551645cd0674d1

Observation 65d59515-c219-40fd-b934-1b5d17cd9996 · inbound

Task complexity shapes internal representations and robustness in neural networks cites this paper.

Task complexity shapes internal representations and robustness in neural networks Open Problems in Mechanistic Interpretability

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:51:54.745952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T23:50:09.749179Z digest=sha256:637f2d20ad3ddf613300adb3b509ed8677e08cdd671fabb64547ef87fac4ef35

Observation 1127dc0b-0c0f-49ac-b1f8-90151723e9e8 · inbound

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns cites this paper.

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns Open Problems in Mechanistic Interpretability

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T18:42:02.956496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:42:02.956496Z digest=sha256:3f9fbef0a0bb9e54a01a228a525c529a20c2cebb3894123c6165512f76fe6388

Observation 486e437e-19bd-4800-bbea-32c5ed1fca99 · inbound

SATORI: Static Test Oracle Generation for REST APIs cites this paper.

SATORI: Static Test Oracle Generation for REST APIs Open Problems in Mechanistic Interpretability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:48.632505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:26:48.632505Z digest=sha256:b69e857b9d32e251165347788353a52e68b07ae28c13197bc341f0b22fa1bf4f

Observation a00a0626-d2ba-46e1-a2c7-2d14dcd321c4 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Open Problems in Mechanistic Interpretability

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.332785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.332785Z digest=sha256:bc1dfe4fada6fec2ab4bf0e6806f79560d1cd9840b97185f1a118726b1b5b1e3

Observation 7aa04eb6-b3bc-42af-97a7-eff4940835a2 · inbound

Challenges in Understanding Modality Conflict in Vision-Language Models cites this paper.

Challenges in Understanding Modality Conflict in Vision-Language Models Open Problems in Mechanistic Interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T11:26:17.847288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:26:17.847288Z digest=sha256:dfe18352b55824b4a695e3b77b45794cbca2c0d1bf727e7b3d4d7a3e056a01f9

Observation 9feca2dc-b196-4eed-918f-dcdc12b219cf · inbound

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World cites this paper.

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World Open Problems in Mechanistic Interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:46.324925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:46.324925Z digest=sha256:2b5756d09c17e1b390e9d0c6ef201461f1a4571f7e5fe28c17ec3379ef41ea1e

Observation 46dd83e5-81b7-4785-8c23-1c4a5d7ea856 · inbound

We Need a New Ethics for a World of AI Agents cites this paper.

We Need a New Ethics for a World of AI Agents Open Problems in Mechanistic Interpretability

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:01:02.649733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:01:02.649733Z digest=sha256:de49cced5c6eb9e60cb6b98299a949b35129497998b384584a01b07e1b64f015

Observation 0dd62860-9d42-4914-b4ac-67e3bdb9673a · inbound

Do Activation Verbalization Methods Convey Privileged Information? cites this paper.

Do Activation Verbalization Methods Convey Privileged Information? Open Problems in Mechanistic Interpretability

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T15:42:42.156296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T15:41:46.771905Z digest=sha256:ac9ca08d891764dc0bc632520eb8ea20c5c93e5de7d548093cc291ef3428276d

Observation 3973880d-2d93-49e7-8241-a9134b7803b1 · inbound

ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack cites this paper.

ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack Open Problems in Mechanistic Interpretability

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T13:06:23.426950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:54d3667c15d7f4d2164a5f99bdf3738aa244902f305180a9e892b8b7a88ca7ac

Observation fe01834c-621b-426b-ad2f-964f424cf967 · inbound

Feature Identification via the Empirical NTK cites this paper.

Feature Identification via the Empirical NTK Open Problems in Mechanistic Interpretability

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:01:16.935370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:01:14.230977Z digest=sha256:196063d153a73b74b0fb18c4ee69582dc1d1fba901eb79b9fe4843d810f7e4aa

Observation afa98fb5-78dd-4682-82bf-0106ea6ebe34 · inbound

Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling cites this paper.

Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling Open Problems in Mechanistic Interpretability

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:41:16.005714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T10:39:32.910190Z digest=sha256:507be4d8af3dfe09f9529c6c096f423eb56e301ddc130172d36a3cfb21ec072d

Observation 8be092f4-4e0d-48bb-813f-c95844dc4543 · inbound

Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary cites this paper.

Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary Open Problems in Mechanistic Interpretability

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T13:24:53.325174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:22:37.107679Z digest=sha256:5e3aec5bb61a3792472cfe9622356016ecd0c7e57940daf692d2c93e4b3e984d

Observation 6894c222-62fd-4f54-8a11-30e74d357a9e · inbound

Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary cites this paper.

Decision Potential Surface: A Theoretical and Practical Approximation of Large Language Model Decision Boundary Open Problems in Mechanistic Interpretability

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T13:24:53.065282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T13:22:37.107679Z digest=sha256:c522f74a1dc11e9505f246dc02dbdc87acd9c348c7590d19f30e1589a33483c6

Observation 49e5798c-ace3-4197-a9c7-2e6038316587 · inbound

Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations cites this paper.

Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations Open Problems in Mechanistic Interpretability

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:54.672241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:54.672241Z digest=sha256:a67edf0ee8bc5695ea1e0a0ff0f1d649686227c3eceaf88ffb30f9a0da996eb2

Observation 69af4bca-e6f9-4c18-a824-b6177d61e9a6 · inbound

Some economics of artificial superintelligence cites this paper.

Some economics of artificial superintelligence Open Problems in Mechanistic Interpretability

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:18:42.819258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:18:42.819258Z digest=sha256:bc148bdadcd60f000986409f376b5a624c01b2ed906729c9d55d6e542baf34a7

Observation aca47b11-2890-4de7-8d0e-993852cd1ac5 · inbound

Internal Deployment in the AI Act cites this paper.

Internal Deployment in the AI Act Open Problems in Mechanistic Interpretability

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T17:54:18.538252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T17:51:47.841707Z digest=sha256:2277cf8d9b3f7fd7b78a7a9737e1270d5cb8a62f9726204acdebec4af0699296

Observation d610a1c0-cf66-4de5-b807-533ef22ec407 · inbound

Singular Vectors of Attention Heads Align with Features cites this paper.

Singular Vectors of Attention Heads Align with Features Open Problems in Mechanistic Interpretability

Reference 1972

Resolution
unresolved
no resolver link, observed 2026-08-02T23:38:14.568953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:38:14.568953Z digest=sha256:145a2bb4b43fbad9d6e4cd1062a5991f88641a7dea01e29d5e11774982d173a3

Observation 90b20a3c-93e6-4d89-888e-f2b6e3d23fb4 · inbound

Transformers converge to invariant algorithmic cores cites this paper.

Transformers converge to invariant algorithmic cores Open Problems in Mechanistic Interpretability

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T20:46:29.019024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:46:29.019024Z digest=sha256:f852f7e9b42b4ab2dcaadef3f23f28d72e78085c15fffe33b2b32903cc735024

Observation 327b16b1-7168-42ea-a292-e4ebbb6c1b41 · inbound

Steering at the Source: Style Modulation Heads for Robust Persona Control cites this paper.

Steering at the Source: Style Modulation Heads for Robust Persona Control Open Problems in Mechanistic Interpretability

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:21:26.170990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:21:26.170990Z digest=sha256:6b11a7e75f822499f8470e197fe902de54b6ae262b5c50199afb0c25f975aa9d

Observation 50faadc7-4624-4cf7-9ec9-c6c7d9942b90 · inbound

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space cites this paper.

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space Open Problems in Mechanistic Interpretability

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:39:50.311059Z digest=sha256:fc94665639a40877ea728d253c924cfdf77fb9af6726810904549c4098b74bf8

Observation d2337bd9-2222-4d26-8495-b65ad970e9af · inbound

The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts cites this paper.

The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts Open Problems in Mechanistic Interpretability

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:25:57.273965Z digest=sha256:a38a2f2493b72b7fff6a6ebf7957ed4b813c46cb4f4ddcec6ba5b111c92f6773

Observation 99127db3-ebd8-4dd4-8844-f2479e024303 · inbound

The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts cites this paper.

The Linear Centroids Hypothesis: Features as Directions Learned by Local Experts Open Problems in Mechanistic Interpretability

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:12:07.373414Z digest=sha256:02151bdd5a1cc300eee967bcbcea7a4609783206ebb3d2df0ed7c64c2b29265d

Observation 9c48241c-24be-43f4-b282-22e72fd9de63 · inbound

Functional Similarity Metric for Neural Networks: Overcoming Parametric Ambiguity via Activation Region Analysis cites this paper.

Functional Similarity Metric for Neural Networks: Overcoming Parametric Ambiguity via Activation Region Analysis Open Problems in Mechanistic Interpretability

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T17:51:29.832500Z digest=sha256:752e332819b2225154d65bb2cadf80ce6d6c8fb656dbf1742e5b57f9befe87b3

Observation fdad4d2b-6a09-4abb-a7e6-0ed2e1d0bc0c · inbound

Diverse Dictionary Learning cites this paper.

Diverse Dictionary Learning Open Problems in Mechanistic Interpretability

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:39:26.132003Z digest=sha256:4821a144653cee1f8b3c2e5a9a9c4b72cff995d1740b2a5c0f6e6dcfc7c14ae4

Observation 602ab21c-a12c-45c1-bf27-2de8195f388b · inbound

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety cites this paper.

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety Open Problems in Mechanistic Interpretability

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T04:17:43.880661Z digest=sha256:e15c5b9fe4b47129dfcf666b27fce25bfc3e202561c342113e6cf6d156756ad2

Observation a98b13c9-7ceb-4c25-8ef4-53b34d117048 · inbound

Understanding the Mechanism of Altruism in Large Language Models cites this paper.

Understanding the Mechanism of Altruism in Large Language Models Open Problems in Mechanistic Interpretability

Reference 230

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T01:36:50.329664Z digest=sha256:c7bfb52ee54f4efa97f2024d7816304d12e213e83a7b8cee22971dcd1628cb4a

Observation f898cbd0-13be-4004-9f07-cb2d2b90f160 · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning Open Problems in Mechanistic Interpretability

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:52152c90ef63bfac2605bb61641a09ed34519e1c4b1bebe7a3d16ac9fc63bc0a

Observation 8d03db4a-339b-41fd-9ec3-1604d377aff2 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Open Problems in Mechanistic Interpretability

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:ca8be91b1a3f27c987544110c29e43afb5021fb94fa0bb690687cab0d9f19149

Observation b176c6a8-7299-4588-afea-58e6699bcce2 · inbound

Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts cites this paper.

Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts Open Problems in Mechanistic Interpretability

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T18:48:04.046370Z digest=sha256:f3f5ac9dccc5e2db87df754a80a42eda70007859ce1f21ffb6503ec63c20ebfd

Observation fc33ff26-dc31-4f7c-9083-1c53c39ec565 · inbound

Linear-Readout Floors and Threshold Recovery in Computation in Superposition cites this paper.

Linear-Readout Floors and Threshold Recovery in Computation in Superposition Open Problems in Mechanistic Interpretability

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T14:20:04.839150Z digest=sha256:b38f75d03aaca372e589b480c655352b54d66ed1dfafc1f0e9a2bc639a1debe2

Observation 3f373b76-6617-469a-84b1-b2b6b46973fd · inbound

What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis cites this paper.

What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis Open Problems in Mechanistic Interpretability

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T18:33:43.998202Z digest=sha256:52e9bf766c3a91ee20408fd3c0cccc8e03613a942892919a63ec5b88be89d607

Observation 1ec700c8-2b4d-449a-a155-81d2976a55f3 · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Open Problems in Mechanistic Interpretability

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T15:35:08.202464Z digest=sha256:0a0bb7f9966f84baa44706a3dc1c7f9f3d4504e5c83125549673dbb5ff70ce05

Observation 6d6162f0-68f5-41c0-b5b5-238ac11a8d83 · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Open Problems in Mechanistic Interpretability

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:35:07.981613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T23:25:55.375541Z digest=sha256:ceee5499d6511d5ad59c794a969eda9a124a958ca2c3a92d9842ea2f7ed9c514

Observation 4d72dd99-4ea6-407e-a315-744b1ea3a925 · inbound

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior cites this paper.

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Open Problems in Mechanistic Interpretability

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T17:47:09.591001Z digest=sha256:6cb1a2da56886d8cd3226bf7ced01b615ad1664a246a2dc259da2d7ec2fab105

Observation 1554ce28-e0c3-4fd1-a1f7-e3ac49780471 · inbound

Bilinear autoencoders find interpretable manifolds cites this paper.

Bilinear autoencoders find interpretable manifolds Open Problems in Mechanistic Interpretability

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T01:25:35.920367Z digest=sha256:a2d5788631215f429372895d74c468c8acb4b93db41e9196c23633f63140349e

Observation 0d33df20-8fa4-4df4-a30d-a013c45da974 · inbound

From Mechanistic to Compositional Interpretability cites this paper.

From Mechanistic to Compositional Interpretability Open Problems in Mechanistic Interpretability

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:42:26.173782Z digest=sha256:356ae35737bcd747748abd6b9f27afa6ec63839a2fb533cb07cbdd9ade2e19ec

Observation ebcc1c64-06cd-4f63-8baa-fd9145dc60d2 · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? Open Problems in Mechanistic Interpretability

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:47:47.214726Z digest=sha256:3229f3cad2366f93863fd2741fec73762c37f57a9730d564678e938ce1b6cd5d

Observation 0d44401a-0fe5-4f4f-8d5b-8f57e2d138bb · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? Open Problems in Mechanistic Interpretability

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:07:40.471481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T17:06:22.481341Z digest=sha256:b2652b1754c88ca61f91e33e5a88eacdc977b3206eb75cb22021fb3734e80db0

Observation 35e42c23-a4a1-4037-98de-8da6505aa50a · inbound

fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery cites this paper.

fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery Open Problems in Mechanistic Interpretability

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:54:41.225199Z digest=sha256:fd1c17266ad20a17177b09edfc0d22f7e185b3cf77f020bab6553cb57577a027

Observation 20a61b1d-76e3-4180-bebd-16a5a2a48517 · inbound

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models cites this paper.

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models Open Problems in Mechanistic Interpretability

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:38:31.094834Z digest=sha256:22041128e429b39041746c6c0fd85e569d7bc83058b5a206888e4f5eb5a9c117

Observation 82a9a522-b342-4779-960f-0947f5b1bd04 · inbound

Metaphor Is Not All Attention Needs cites this paper.

Metaphor Is Not All Attention Needs Open Problems in Mechanistic Interpretability

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:59:39.485618Z digest=sha256:ee37a629e5a9c03f81db84259f5a57a653a4bf60169b8a9c2b6a0e218ddef59e

Observation a016c9a3-61bc-4973-8d54-32712ddabde8 · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Open Problems in Mechanistic Interpretability

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:30:06.348607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:46caffdd248ec231a109a1be43fe24d5cbf6adf8b813188a406e5aa57d963874

Observation 97721bca-9435-45d4-87fd-e19266247cc5 · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State Open Problems in Mechanistic Interpretability

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:59:28.715619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T20:53:40.666929Z digest=sha256:1c3cfb496ac065253bbb0f3586671ce0cbe440fef1db64c9815707f90e495580

Observation 3dcb8d84-c08f-4b65-a915-38604c7543ac · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State Open Problems in Mechanistic Interpretability

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T04:59:45.052674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T04:59:11.877068Z digest=sha256:f346c8970ba3b2834e662498399b71612385b710e86bd68d70461726abcd899a

Observation 57a03569-8ab6-4e6c-a037-ed4b5c2a5616 · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State Open Problems in Mechanistic Interpretability

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T21:53:47.195232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T21:49:47.934339Z digest=sha256:f6282cfa6e2d6ba0d16077a6b2262e9f9a667617ab3b0d3c01917a4c69212904

Observation 83788547-003f-413a-a697-dc1896affb6b · inbound

Tracing Persona Vectors Through LLM Pretraining cites this paper.

Tracing Persona Vectors Through LLM Pretraining Open Problems in Mechanistic Interpretability

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:29:27.899540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:28:17.086117Z digest=sha256:25e60422cf4836b9bb1728f5ebae051fc90074b365dd00355aa00379273e989d

Observation 896c9ec7-051e-4b98-af56-eef1f043f7c3 · inbound

Explainable AI Isn't Enough! Rethinking Algorithmic Contestability cites this paper.

Explainable AI Isn't Enough! Rethinking Algorithmic Contestability Open Problems in Mechanistic Interpretability

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T19:12:43.807569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T19:08:05.556651Z digest=sha256:dacb9ffdd4872f81e17a50447c67cd34ad05b031c6113e329201d4572a8ff301

Observation 28fdf147-299c-479d-9c97-ea8bffb1e399 · inbound

Mechanistic Interpretability for Learning Assurance of a Vision-Based Landing System cites this paper.

Mechanistic Interpretability for Learning Assurance of a Vision-Based Landing System Open Problems in Mechanistic Interpretability

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:59:45.603518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T06:56:23.842767Z digest=sha256:9156502acff4753ba863ab75b17a665842fa95fe37630655d0e8d991d885b890

Observation 5e5cc594-e20c-4b20-b906-73920776928a · inbound

Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift cites this paper.

Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift Open Problems in Mechanistic Interpretability

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:16:16.110708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T08:14:58.183289Z digest=sha256:05e47e13845fba2bfa13700f860cc32f886fc4b1ac874a63dbf03a6946e02f22

Observation 5f1e45e0-bde3-4f0b-816f-c2a5a2ebd8ee · inbound

Disentanglement Beyond Generative Models with Riemannian ICA cites this paper.

Disentanglement Beyond Generative Models with Riemannian ICA Open Problems in Mechanistic Interpretability

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:51:15.997118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T07:49:28.234614Z digest=sha256:666bc2683f6fe8cb8df79c4da8eb8363b2ca0d0ea167092fc5126436686588d2

Observation 5f0c9566-37d2-4ab2-b891-89a203de4cdf · inbound

Relational Linear Properties in Language Models: An Empirical Investigation cites this paper.

Relational Linear Properties in Language Models: An Empirical Investigation Open Problems in Mechanistic Interpretability

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:46:15.014015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T07:45:09.743615Z digest=sha256:439a527e7513654643e959c39079a3c8e8b01f2801dbec639e6d81e1bdfc6628

Observation 43b05f4a-3053-40c4-92e1-e2c0fc207530 · inbound

Relational Linear Properties in Language Models: An Empirical Investigation cites this paper.

Relational Linear Properties in Language Models: An Empirical Investigation Open Problems in Mechanistic Interpretability

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:14:56.781346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:13:03.620676Z digest=sha256:96af7b08f19dfb52f3a4d961346df2a0262166c0e78a178a20ffbdbcb827de62

Observation f7a3be11-03e7-4581-8494-14c97ca3fb21 · inbound

Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing cites this paper.

Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing Open Problems in Mechanistic Interpretability

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:03:29.439485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T14:00:31.962058Z digest=sha256:6a09b6e63ee02ea4f885a643cf34ab294248a51b526673630ae9074e14cf7d42

Observation 9155af1c-7b97-4a31-a89d-d98f687d60f9 · inbound

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability cites this paper.

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability Open Problems in Mechanistic Interpretability

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-28T02:11:29.091386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:07:18.198225Z digest=sha256:24f7edd6f0850dd9bb8ed8b92395c0e6af83bb017a6719589bd3fb5ab2525589

Observation 33e2b1c3-99b2-4f9c-a1f8-9cc668a95aaf · inbound

Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims cites this paper.

Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims Open Problems in Mechanistic Interpretability

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:47:27.708321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:33:19.875024Z digest=sha256:4a3aeac83712a0df81aaa445e84a6656d3c8bfe69f0d96eac976a01f32d9e07a

Observation c4f2eab3-fea5-4e2f-845e-f7ef0c831a57 · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Open Problems in Mechanistic Interpretability

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:57:32.398548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:99656798842b0ad7581f6a42a16b4a1fd48d26e658864183aab4f8b2e33feea1

Observation 55d6c3bb-6377-4b1c-91a0-2aadc661fb90 · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Open Problems in Mechanistic Interpretability

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:07:47.895988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:32343296bb79fe11ace498141b3500e7a5dfbb45e10c1e2cc11d502db831c00d

Observation a1c88b3d-f970-4ee2-a3c0-cf859e63efc9 · inbound

Learning Dynamics of Chain-of-Thought State Tracking in a Solvable Transformer Model cites this paper.

Learning Dynamics of Chain-of-Thought State Tracking in a Solvable Transformer Model Open Problems in Mechanistic Interpretability

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:39:05.249303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:49:03.896740Z digest=sha256:a8ff5a26f931ccf1c833b8dfb3d8de36a2db2ce082b3cbd2836aee6fe4cd7072

Observation 0aff0b28-aaa9-421b-8ed5-7d955597a588 · inbound

Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds cites this paper.

Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds Open Problems in Mechanistic Interpretability

Reference 116

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:58.962774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T23:54:29.531368Z digest=sha256:23b281a6dcb5074df55799025766bba0d65741264d1ef88343743ec105f2819b

Observation 49197e46-f0cc-4439-ba2a-693b072f57a3 · inbound

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders cites this paper.

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders Open Problems in Mechanistic Interpretability

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:55.788579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T20:16:48.256807Z digest=sha256:4ba3260389732203545c659067fc538110e037eb55f62b25b1dfa4b7866ca108

Observation e20f4c55-112c-4abd-9bca-ff9adfac4a08 · inbound

Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources cites this paper.

Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources Open Problems in Mechanistic Interpretability

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T03:00:02.282468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:00:02.282468Z digest=sha256:489ce7eb86f266c925e46e43f803a05f34ce5ebe59e63233ffb294741c19dbf8

Observation 48c5a2a1-ba7b-4c2b-a206-d130a5e166fe · inbound

Unsupervised Features Mining via Activation Geometry cites this paper.

Unsupervised Features Mining via Activation Geometry Open Problems in Mechanistic Interpretability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T20:51:42.052454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:51:42.052454Z digest=sha256:0913539ba2e93369763c251278c0750b1f744518242f67d1b1b0929c2cf3642c

Observation 925f8d3b-88dd-4495-aa4a-3fe6e71c89fb · inbound

Weight-Space Physics: Interpretable Hypernetworks for Lattice Quantum Field Theories cites this paper.

Weight-Space Physics: Interpretable Hypernetworks for Lattice Quantum Field Theories Open Problems in Mechanistic Interpretability

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-09T19:46:28.683089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T19:45:01.669063Z digest=sha256:74145b56f25e137b9af0417f819fa83159b3ee54e4da13891bf5a5ea41e51c5c

Observation 14bd7dad-d5fd-480c-9cea-ef3020aba4bc · inbound

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning cites this paper.

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning Open Problems in Mechanistic Interpretability

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:06:18.032089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T14:58:58.363330Z digest=sha256:2d9c41ca0277d3c0150859f4333c5020b6c9e60a90db0ca7d0fca109b785ce32

Observation 8255463b-7d81-4aa7-a958-0bdf24a853db · inbound

Laguerre Geometry for Interpreting Large Language Models cites this paper.

Laguerre Geometry for Interpreting Large Language Models Open Problems in Mechanistic Interpretability

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T10:42:30.055074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:42:30.055074Z digest=sha256:7dcdf2a99a7a7edb53db9d3010fa25fe33587548755dfe691e83cb57c54cd863

Observation 90bedcb2-6a35-442e-a894-69f72a24cca6 · inbound

Targeted Recovery of Weight-Space Mechanisms From Neural Networks cites this paper.

Targeted Recovery of Weight-Space Mechanisms From Neural Networks Open Problems in Mechanistic Interpretability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:39:48.676625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:39:48.676625Z digest=sha256:23e2a6e4dda32318a5f4ee1907978c4a24049d94dfad92aeb43b0de4fddba06a

Observation f612be05-850c-4c8f-bb3a-4aa956228a1b · inbound

Targeted Recovery of Weight-Space Mechanisms From Neural Networks cites this paper.

Targeted Recovery of Weight-Space Mechanisms From Neural Networks Open Problems in Mechanistic Interpretability

Reference 175

Resolution
unresolved
no resolver link, observed 2026-08-02T10:40:03.484884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:40:03.484884Z digest=sha256:4fb4555e23acf9d0b8b76bdcadd45899c3702cb8ae706fd0accc9abd3a467205

Observation 83e2ae79-c125-498c-9332-d3f102e15528 · inbound

Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods cites this paper.

Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods Open Problems in Mechanistic Interpretability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T10:46:51.069511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:46:51.069511Z digest=sha256:f6d0e181d51885065b970ec2291c855e6610e4e0fc998e29fbb9d28b1b0c17b3

Observation 5c622027-b414-473d-abdf-cdeb1c968870 · inbound

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control cites this paper.

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control Open Problems in Mechanistic Interpretability

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:42:34.779473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:42:34.779473Z digest=sha256:6148cc806d3d2bde58890e6b7b761136d64ed26c906e9e142688689a13d21f8a

Observation 6250a2e1-3dde-4b5d-ba0e-b62f793ffa87 · inbound

An Analysis of Residual-Stream Geometry Across Transformer Depth cites this paper.

An Analysis of Residual-Stream Geometry Across Transformer Depth Open Problems in Mechanistic Interpretability

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T17:18:44.974316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:18:44.974316Z digest=sha256:5b6b042c3fb38a361deefd8b562a169e074826552d544f89a53babdb1e1e893e

Observation 8675a0bb-2478-488f-99a4-f67be5a1fa3d · inbound

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models cites this paper.

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models Open Problems in Mechanistic Interpretability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:14:23.380605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:14:23.380605Z digest=sha256:b44fe2d9c2aafcd3178a3212618f9416d0734cbdd7473df35700011cc04d2216

Observation 865ee866-0a31-4a2e-a205-505620e1f5a6 · inbound

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model cites this paper.

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model Open Problems in Mechanistic Interpretability

Reference 125

Resolution
unresolved
no resolver link, observed 2026-08-01T12:16:42.768478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:16:42.768478Z digest=sha256:c1c365b4ef28ad27f4faee012dc2eabf9bf88a908d6a1bb43a8ac67ab470455f

Observation 084596d9-7682-4911-b82e-071205ac7e7b · inbound

Emergent Latent-State Computation under Stochastic Volatility cites this paper.

Emergent Latent-State Computation under Stochastic Volatility Open Problems in Mechanistic Interpretability

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:27:03.485740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:27:03.485740Z digest=sha256:1a8ff064e45e9f95efe3a3dc8bdb27cb549c397dd78720bde4ce5ba4508e8409

Observation bcfd39c0-cf36-4cde-8840-92ba590592f8 · inbound

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs cites this paper.

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs Open Problems in Mechanistic Interpretability

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T23:07:40.684953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:07:40.684953Z digest=sha256:01843c1fda8a19561984eab1f9eaa2b9146abae5da39d015d21864cfadb67635