Pith. sign in

Paper Citation Record · LEDGER

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

As of 16 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 1 inbound Pith citation observation for arXiv:2505.01372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01372 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:25:06.108473Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:42:26.173782Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact5
  • verified fuzzy33
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7be2a119-13da-4a7a-a1d7-649e3984d9d1 · outbound

This paper cites The Computational Complexity of Circuit Discovery for Inner Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Computational Complexity of Circuit Discovery for Inner Interpretability

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.518727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.518727Z digest=sha256:42f87ce3505413e64c4709047198af50a201e3cec391f9cb2520ecaca6d50121

Observation e96be6eb-3233-415f-bede-dba250dc41a7 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.526366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.526366Z digest=sha256:fbc9bfcb03c90e5efb8eb38316e06b1d3906569671677b6b482e7f5ddddea1a5

Observation 61d387c8-950c-4507-82b2-d3eb05f6a283 · outbound

This paper cites Physics of language models: Part 3.1, knowledge storage and extraction.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Physics of language models: Part 3.1, knowledge storage and extraction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.531896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.531896Z digest=sha256:182a64a38df274e5b1a606473b1eb7c8b00ac1a73d7ae882290ebac3811b1026

Observation aae00ee6-8ee0-419b-adff-b23b52bf598b · outbound

This paper cites The urgency of interpretability, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The urgency of interpretability, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.538644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.538644Z digest=sha256:a6e04fa91933019b04b8ee6894aea74562962a865d2d68e5db8b91d8e5f5d35d

Observation deff4b33-b745-4e3a-9e7a-86e1815bb557 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.544723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.544723Z digest=sha256:5440c817cc626dae46295597347d8f709265b191857133b1612f8605bb4a640b

Observation 964da59b-029d-4588-8631-c9b228d56532 · outbound

This paper cites Ai as systems, not just models, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Ai as systems, not just models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.550537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.550537Z digest=sha256:d5dfdeb6c0e5b4622e06c29b10ccf90d69d423cd37e58c53fa93714fecc13a59

Observation 7c179344-ca3a-48d8-82d1-3ba93e274d41 · outbound

This paper cites Standard saes might be incoherent: A choosing problem & a “concise” solution.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Standard saes might be incoherent: A choosing problem & a “concise” solution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.556998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.556998Z digest=sha256:8637053dda360523e3c52174bee52eb7836f2bef7c7fdb6896463c7496ee970e

Observation 1000c22c-0b9a-4e30-8e2f-592afc155ec8 · outbound

This paper cites Position: Interpretability is a bidirectional communication problem.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Position: Interpretability is a bidirectional communication problem

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.563027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.563027Z digest=sha256:b1ed5ccf6cc2468feab0d5b7c67b66d7af6dbac15b402ab2715c03ec5c239f82

Observation 237b7281-268c-41a9-b03e-66dee8e8639d · outbound

This paper cites A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.569755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.569755Z digest=sha256:c07f264d3e3754d0a30e74a260c8bfd5afb1ba98c81c58dfacb3eb99eb0e3f35

Observation 2f798255-8dd9-4727-8730-208b6e6b64e5 · outbound

This paper cites Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.575857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.575857Z digest=sha256:ff82881438bd5845a61c50b693ae06f39fe9c3fd7f87259feac7cc11094ed6a2

Observation e326204e-4a5b-4139-967e-9d642129ebdf · outbound

This paper cites Novum Organum.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Novum Organum

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.581703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.581703Z digest=sha256:c049590285e22a5f63dc8ee6bf6dab98413bf4685072894aa0ea61eb525183cf

Observation 221e6ee1-3828-4e38-ba66-1b3543b46e24 · outbound

This paper cites Simplicity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Simplicity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.586827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.586827Z digest=sha256:7fed91ce0b6f54c8605f8b223f86a04b62c2f1bc51a5800b0bb20428a020f5a8

Observation becd414b-4fff-4ff3-954f-598d5526b840 · outbound

This paper cites Design Rules: The Power of Modularity Volume 1.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Design Rules: The Power of Modularity Volume 1

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.592337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.592337Z digest=sha256:bf2d5519ec60fa63b43bbf616af7506650cbd5c359841754f2848c441f3ee164

Observation 98a5edf3-3bec-4513-b1e0-5c1ecb7f32cc · outbound

This paper cites Icml 2024 mechanistic interpretability workshop, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Icml 2024 mechanistic interpretability workshop, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.597081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.597081Z digest=sha256:a7f99b521e9b6f4e1b42ff5518b3e22ff145bf85aa98a0c138ffb5aae4db2800

Observation 64e4beca-6ae9-4858-b525-195cbebf0549 · outbound

This paper cites Local vs. Global Interpretability: A Computational Complexity Perspective.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Local vs. Global Interpretability: A Computational Complexity Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.601776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.601776Z digest=sha256:293880a34ebfe2a447bc305a47e61059ed6f0f3600ce8265c80bfd41b7da62df

Observation 3fa773a3-d922-4b64-81cd-e726fb950bf8 · outbound

This paper cites Explanation: A mechanist alternative.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Explanation: A mechanist alternative

Reference 16

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.262246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.606747Z digest=sha256:8039ccdfe69d463aa4ce9dbef6d7c23e008ffeacb62965c11c339e246dbe2efc

Observation 3c2e7bee-d2fd-4024-bc7c-1cb2370aa993 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.612261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.612261Z digest=sha256:3f9cfc81c10aa0dabd9e5b74c57c5a4dddfa5776e8266211112b12b12a8096d5

Observation c90eb841-85a5-44f5-b15d-ee4d768cc0c8 · outbound

This paper cites International AI Safety Report.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii International AI Safety Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.617272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.617272Z digest=sha256:9ef45ad4beda8aaa954c29ef79d7b5de934760fbc93f9813e8ef28738a90d6da

Observation ec496079-ff91-465c-b9db-71c5d491d4cc · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic Interpretability for AI Safety -- A Review

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.622447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.622447Z digest=sha256:8ab3a34cc3f9e8c281736b7efa8fe5ba0396fe9299c3d046b734a2b3f354832a

Observation 03f4349d-828d-465b-9065-1c382d45670d · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.239659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.627697Z digest=sha256:9fe240e35f0e0a7ad36420c62c2e3ece2e18c41149223fa625e7da7ad538dc1a

Observation c960ebea-2022-4f48-a643-0b57db62aa68 · outbound

This paper cites Auditing local explanations is hard.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Auditing local explanations is hard

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.632693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.632693Z digest=sha256:2ef35fd4f0129034d83176eb2f2ba6ed38b1a3b691013d846a21dd0dafe65160

Observation ed2767c7-3100-4b87-beb9-87dce637ba3a · outbound

This paper cites Language models can explain neurons in language models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Language models can explain neurons in language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.637441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.637441Z digest=sha256:9411902bd48ed749d02fd5abfd84bca19cbf3e7eeb8ed2638b356070cb7a3383

Observation 87cc22af-9945-4683-b20e-ef73fe38ec6a · outbound

This paper cites An Interpretability Illusion for BERT.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii An Interpretability Illusion for BERT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.642287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.642287Z digest=sha256:a7e490c9c9f1ba975740dcb7009908b2637373a16ab4d467d7b9b6bfb8a03b3a

Observation a2b41a39-4f83-4cfd-a9f0-3663c232c3da · outbound

This paper cites Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.647126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.647126Z digest=sha256:d6e1d2816a4246197e06c783f85cc4f17e8d26c7695011014ea8cf0648181f31

Observation 3a76e277-6588-436b-b83d-5be24a5b09ec · outbound

This paper cites Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.652281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.652281Z digest=sha256:37f6bfb1ae87d8adcce4a566d839f44b7711b6380bc5aab3319a992e2a8e3372

Observation 135608c9-8774-44fb-bf3e-18692e656f90 · outbound

This paper cites Towards Monosemanticity : Decomposing Language Models With Dictionary Learning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards Monosemanticity : Decomposing Language Models With Dictionary Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.658022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.658022Z digest=sha256:bb587be22a1544e3202f80a98ebf299c5d8dceb1903aed341059c38282da208a

Observation 34747481-f21c-4452-9e82-01aac17a063e · outbound

This paper cites Propositional Interpretability in Artificial Intelligence.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Propositional Interpretability in Artificial Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.663766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.663766Z digest=sha256:ba42b9c564781aed3a2bbbf12647bad511976194dbb61a05e75ff695cd8212ef

Observation 8c79eff4-85ec-4984-8618-b2188376c0c3 · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A toy model of universality: Reverse engineering how networks learn group operations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.669143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.669143Z digest=sha256:6162d3dc4de92100dd16fe34896e596b0c08977601ede60ad77aed776d20dbec

Observation ab6e7c15-147c-49e3-a67e-2ee2331be54d · outbound

This paper cites The evolutionary origins of modularity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The evolutionary origins of modularity

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.674171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.674171Z digest=sha256:f931e5f1c70732c924d443c9562e0e57356b69b28cd6100c7ad8ce52f23fbee4

Observation 0bbc4ffc-89aa-4a8a-9228-f21bfc00c894 · outbound

This paper cites Towards Automated Circuit Discovery for Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards Automated Circuit Discovery for Mechanistic Interpretability

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.678871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.678871Z digest=sha256:c67e2445e9bb4432b4857b582372e4779a2c2b71951854868c778385c730c8dc

Observation 288ced46-da29-40db-ae02-4244fd08479c · outbound

This paper cites Central dogma of molecular biology.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Central dogma of molecular biology

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.684525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.684525Z digest=sha256:4adb4cba55b30c1a87f67f2820605e808dfb29ce2ec7e83cb695f525d27be4e8

Observation 0076b1cc-574b-4515-bbd6-8b68134604ec · outbound

This paper cites The beginning of infinity: Explanations that transform the world.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The beginning of infinity: Explanations that transform the world

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.690760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.690760Z digest=sha256:736f4a078cb0cf4754f122382f4c93847df662b4e88bd7275eb9b634e4ede5a9

Observation 9e2e0718-96c4-4dee-82a6-24b4caa1f973 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.696683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.696683Z digest=sha256:101a8231befb7dfd154f083fb7633e5fd0aa6d8a19a8af608b62aef8fc8faf36

Observation 1643f022-9fd7-4021-af3f-81f37f4af433 · outbound

This paper cites Einstein.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Einstein

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.702151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.702151Z digest=sha256:675ce996c29fbc3dade76578d8e612d3d53fa29f8457df9e2e43fca310812ca2

Observation 6badf7f6-0bd7-48dd-b6d1-dbc160e7f37b · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.707415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.707415Z digest=sha256:7832558818b0d06a8f65bea4038298fbb1b7b690f74a38a3fe67c2c71bff0fe6

Observation 465f48a9-16d1-4854-9556-2b1edce37af5 · outbound

This paper cites Clusterability in Neural Networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Clusterability in Neural Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.712412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.712412Z digest=sha256:169b31f4eaa7b5556e7a0043f441712b0ab35238421372b81ba48da369265bd3

Observation d72172f6-6ee7-4cf2-a87e-c8aa458a6a2e · outbound

This paper cites Interpretability illusions in the generalization of simplified models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability illusions in the generalization of simplified models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.718000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.718000Z digest=sha256:ee7342f0e684a52a08454111979a7a6ca3ec04ad2eb806bf5316ef987c182410

Observation f97a349d-416e-4ec6-ba95-63ed4aa9bd7f · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Scaling and evaluating sparse autoencoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.723010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.723010Z digest=sha256:02d8ccb84a7debf085e739ea385466eadf335577c9aaec067a6fa5015464de2c

Observation 712695dc-797c-4577-a954-6c0f803ec29b · outbound

This paper cites Causal abstractions of neural networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal abstractions of neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.728174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.728174Z digest=sha256:075b993c21b7b83122d036c850770cf0d02e247f9b18540c4f34abb966044d77

Observation 22aea982-0987-44f7-ab5b-592ababe55dd · outbound

This paper cites Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.733991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.733991Z digest=sha256:27fb1703af7821b0ed9891a1b8f3d295e76974c07be448df7e9434d3c68def9a

Observation 25eafe53-c2b4-44c6-a494-fd448e1bf91c · outbound

This paper cites Clustering algorithms.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Clustering algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.740650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.740650Z digest=sha256:3f3f33c1620987c300a432d700b5f82f668297257c6cb8b688564589e812315f

Observation 5d25dd86-a4ae-4bdc-bad7-135601377bcb · outbound

This paper cites Compact proofs of model performance via mechanistic interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Compact proofs of model performance via mechanistic interpretability

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.746673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.746673Z digest=sha256:0cef8c401683c927f473c8bbf6c96ecc16ebf9f117c798dd93b9bc1e801125dd

Observation 33317061-ea4a-44d4-a1cc-f3220fce6bba · outbound

This paper cites Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.875795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.752590Z digest=sha256:a7d5b2bf28427d4ff01562ab614d002ae25440e9ce6eca249d5c712dbaa17f1b

Observation 6c1fbada-0a31-460d-8864-8ab9e751bc15 · outbound

This paper cites Hastie, R.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hastie, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.856843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.757293Z digest=sha256:2154378b27245892813727801da83cc41236b384a16a9857a887a6d98e106bca

Observation 6bb4f5b1-e9b0-4de3-b18b-8a9922be0e2e · outbound

This paper cites Hempel and Paul Oppenheim.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hempel and Paul Oppenheim

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.837001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.762261Z digest=sha256:f4f5f2e7158c033c7d96ef7b28199ed90068e4a9aaf91001ed16962555a00a43

Observation 0fc0361d-3d60-41e5-9e5c-c8186b9233c4 · outbound

This paper cites Philosophy of Natural Science.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Philosophy of Natural Science

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.816805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.767579Z digest=sha256:2ceb6d70f40fdb8ed557d964beb886ed6e08891676637c63e86381aacffbea34

Observation e355e61b-deec-4081-8248-fe70d5984b8b · outbound

This paper cites Bayesianism and inference to the best explanation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Bayesianism and inference to the best explanation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.797867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.773005Z digest=sha256:17a19b86ef248bd2d92e28957ff5688d60950726ac21e695df893ae6a55f2c3d

Observation b0510d13-33cf-4cb9-bfb5-b1bc7a98b276 · outbound

This paper cites Loss Landscape Degeneracy and Stagewise Development in Transformers.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Loss Landscape Degeneracy and Stagewise Development in Transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.779468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.779468Z digest=sha256:ee1803c37753925c54933dbd186be75fdfa4cc3f0b5b2e1cb69d5dcc5331aa7d

Observation c5ec30f4-92f1-40f8-9acf-a47b1637be70 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse autoencoders find highly interpretable features in language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.785160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.785160Z digest=sha256:d0ab500b0b679967ef335f10a5d44409338f1339be0916e8d4ba4bbd6fab0dce

Observation 40da69cb-7913-46ec-9031-94e2f0f81322 · outbound

This paper cites Hutter, E.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hutter, E

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.767475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.791367Z digest=sha256:b92f6324b12c386afa7e6ec9f8da37b10c1d05a9bdf8c6b9d9393292bcc66b83

Observation 1c3adc74-c6de-4fdf-b320-a17fec806d6d · outbound

This paper cites Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.749269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.796480Z digest=sha256:187f5bc592afe686aedd04bb451b8bb6d8540a92c83ae32768f3c76127e2d018

Observation 7ff9effe-ae4b-46a8-be2e-3ce3887cac71 · outbound

This paper cites Kandel, J.H.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Kandel, J.H

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.729487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.801635Z digest=sha256:bfc1a3d9c7fa602447aa74635bfc4d54ac2648145f75201b6113468966ba2660

Observation 6d585e48-eb4a-4bc0-9aac-d340d76ab5f0 · outbound

This paper cites Saebench: a comprehensive benchmark for sparse autoencoders, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Saebench: a comprehensive benchmark for sparse autoencoders, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.705111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.807554Z digest=sha256:3538595f0a98eb6e47bed9af6acf05341036b443b8952a1bcfc8e4389fd168e2

Observation fd84da5c-f9e0-43b8-96b3-ed972dd73d61 · outbound

This paper cites Kennefick.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Kennefick

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.687071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.814962Z digest=sha256:12379fc7dc683f5ad22c7e680c13480def8a788a3c1ed10cc1bf9f792e4cfc78

Observation 16b197c7-a52f-41ad-bccc-66f48285f750 · outbound

This paper cites Explanatory unification.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Explanatory unification

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.667802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.820191Z digest=sha256:33f4ffb05133e44b6cf0f4e0bab5baa443ea6b0b9c42807aa507e59b420741b4

Observation 10e71510-9e4a-4a0e-ad3c-0aaf7a4baba4 · outbound

This paper cites Three approaches to the quantitative definition ofinformation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Three approaches to the quantitative definition ofinformation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.646435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.825015Z digest=sha256:392c7714ef4003673c8c4f09498424f1e600c2c9c7bf1381c3f5c92fb7679e46

Observation 17746ef1-137a-4c4f-b7fd-8d3fa293643b · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:25:07.627680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.830341Z digest=sha256:9ed5956a653edcc57f0da1c7e2639cbf8d30540dc34be1f53a2198d43dfe5a00

Observation 1587381a-db0b-4c29-8c54-0a0c152ca444 · outbound

This paper cites The Structure of Scientific Revolutions.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Structure of Scientific Revolutions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.607038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.835464Z digest=sha256:21ea155307ac68e661834ac088cf28708ac2f7da9ca7a272b0dc04d30363942b

Observation aaa33f02-0ff5-4a53-aada-a5d9be8f6d4f · outbound

This paper cites Falsification and the methodology of scientific research programmes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Falsification and the methodology of scientific research programmes

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.586373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.840481Z digest=sha256:dc42abe148d3829fd84020e69352003da2d5c766945324b1240b3babbf5c0cfb

Observation 91c964fc-0773-420c-8279-fc1eb53403b3 · outbound

This paper cites The Methodology of Scientific Research Programmes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Methodology of Scientific Research Programmes

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.567371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.845034Z digest=sha256:64e0b1d26ffc8b5d4de67a9ff08294880eda26150bec25c39d5ada7a58b15635

Observation f9885b31-5dd2-4dae-97f1-f880fb428f8b · outbound

This paper cites Sparse Autoencoders Do Not Find Canonical Units of Analysis.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse Autoencoders Do Not Find Canonical Units of Analysis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.850864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.850864Z digest=sha256:7bd12584f9a97ca52b8a65480669dc7c3566d3e64be1066027c56b932a8a55a7

Observation f7a9f288-c896-4fba-9c44-d8753e20975f · outbound

This paper cites Lindsay and David Bau.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Lindsay and David Bau

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.856242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.856242Z digest=sha256:00f57c063ce5c127c41ccc72d3dad49d05e4dcc2bc76cf2496818a43bfe62ba9

Observation b0ae1407-4a60-447c-8d2e-dada1f82ccfc · outbound

This paper cites The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.862843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.862843Z digest=sha256:71236eb76985013e0b05b8687db4287bba59ce185ea12adb4034805ce1eb58b2

Observation 9eb3975f-8e1f-4e66-a474-ec9c3b1a62af · outbound

This paper cites Mechanistic mode connectivity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic mode connectivity

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.539180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.868951Z digest=sha256:1ea2cbd5ae9a1d5b1e92e14020865c3645c569ab4b843b41d714ade941ad3ec1

Observation ca10ed64-93ed-4ac4-a197-12840824058c · outbound

This paper cites Information theory, inference and learning algorithms.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Information theory, inference and learning algorithms

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.874084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.874084Z digest=sha256:27eea0d23d7f13f1ecbfa3425afa603e222c556cd3087776d191817fbd205d8d

Observation 86793f78-8c9c-49a5-b6b4-0c5e352fa2df · outbound

This paper cites Is this the subspace you are looking for? an interpretability illusion for subspace activation patching.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Is this the subspace you are looking for? an interpretability illusion for subspace activation patching

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.507881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.879925Z digest=sha256:3be40756e8159bcf89b300b44ced2f5185afd68fa74b380d419c22095776c49f

Observation 2085c2d2-ad9b-49d1-9c1b-412de83b2a98 · outbound

This paper cites Downstream applications as validation of interpretability progress, March 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Downstream applications as validation of interpretability progress, March 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.486729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.885763Z digest=sha256:81a02779574bb4f4f94e19c32b05ca5b08142fa67dc3e07e0c42b77bad9a70e6

Observation 484ea65f-aec7-463d-80ea-4d5c9dc4fd29 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.891673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.891673Z digest=sha256:09c2cdfca5afc459ddc2e35a965db574775f5102845cf88933ca359a2f8fcb0f

Observation eb401795-6468-4d96-8e49-cfb9a832798b · outbound

This paper cites Cognitive styles in two cognitive sciences.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Cognitive styles in two cognitive sciences

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.468750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.898471Z digest=sha256:4391470ffa9fb77ad9450215364df529569dadafce98a810e3bff76eb9541dc4

Observation fb3e9be1-3e90-4402-bf7a-be0ded544915 · outbound

This paper cites Zoom in: An introduction to circuits.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Zoom in: An introduction to circuits

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.905678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.905678Z digest=sha256:9019c9122e7d20631de398061d0c01d4f4b35fac97e75a172d52b91984164329

Observation 536f7abc-d582-4db8-9057-fae24a5dba45 · outbound

This paper cites In-context Learning and Induction Heads.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii In-context Learning and Induction Heads

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.911761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.911761Z digest=sha256:5c87775c2c05bb2e165178935a91789a38c6ff098c5adc5faac3a679f81b0731

Observation 293d1e50-1950-4e08-84f3-51ad6d5bb269 · outbound

This paper cites Automatically Interpreting Millions of Features in Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Automatically Interpreting Millions of Features in Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.918441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.918441Z digest=sha256:8ef3e4cc7b9c9e845816b53490df62f9ef54dd1727e458a51f8056e44ae94dd0

Observation 3b0b26c7-9cc4-4b96-8dd8-8f02df58426d · outbound

This paper cites Causality.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causality

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.925700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.925700Z digest=sha256:b091c4fab1c7034f981cb68b16fee79de1334bf873083efda05d9148c14ccb37

Observation c6aebdff-6c3d-404e-b0fd-1062e6f53578 · outbound

This paper cites Poincar \'e.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Poincar \'e

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.416845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.932748Z digest=sha256:2f79b20f6e1e7342d8ae80543cacf2d742df9d5f4f06769195f8dedf526bafbd

Observation e2af5c31-1409-4f9f-b0c2-e63b82fc5688 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.944080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.944080Z digest=sha256:3b9c2433171a266aa7b25c997d2e803dec458d2e5ae4656d6215d47589b10626

Observation 5b60c837-b55c-4963-8912-cee876b825cf · outbound

This paper cites Hume on theoretical simplicity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hume on theoretical simplicity

Reference 76

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.221515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.951779Z digest=sha256:ed5f84e50e2969639f8095a0a15f2b43b33ed88226e9a8058b1fe3fa55ad95c1

Observation e7de4f8c-c5f8-4c6e-bdb3-6d566e58bd11 · outbound

This paper cites Escalation risks from language models in military and diplomatic decision-making.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Escalation risks from language models in military and diplomatic decision-making

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.962513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.962513Z digest=sha256:e0e4d2aa4bd6fc746366c67b6b697a6e858458c9731333090c6cedfb5384f624

Observation 4747c415-0ae3-4b04-b1ee-2029e9bd4a1f · outbound

This paper cites Four decades of scientific explanation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Four decades of scientific explanation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.384390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.968237Z digest=sha256:af07226af7b54c674debcdb1bea3ca0dd4ae41586414590c66b5f895331da224

Observation 80943169-dfbc-466b-a923-efe22d9afad1 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:25:07.365849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.973050Z digest=sha256:6b0faf15f18fee8b9ccf9adfecc187e605599c90eac57dd74890eb74c3a0d151

Observation 29c5b0f6-51c4-4735-99a3-2c9f05b2f048 · outbound

This paper cites Mechanistic?.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.978424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.978424Z digest=sha256:a4ced567ebdcef665b1ad146402808f59c4a2d1c8f8569b0f6069a8da6df700a

Observation 08075a01-c91b-4cc8-8e58-12d0004ac73c · outbound

This paper cites Theoretical Virtues in Science: Uncovering Reality Through Theory.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Theoretical Virtues in Science: Uncovering Reality Through Theory

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.346087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.984022Z digest=sha256:e9d3a28a5be6840241a5c774ae92ef1779e9b132f663356462e5d8af4e77efe0

Observation 740ce7be-8b7a-4b84-b752-dd27ef2f759c · outbound

This paper cites Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.990103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.990103Z digest=sha256:b51fd9727e86cb1971cf8d26f9edd106f2d1145e64638934120ace17ede8f29b

Observation 1b7789d5-9c8b-4ace-9512-ac09ab787734 · outbound

This paper cites A mathematical theory of communication.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A mathematical theory of communication

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.996982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.996982Z digest=sha256:d44cb2e98cc10785155ad3356e6041c6919614f9511252e6d8b17e85a6af9a33

Observation b417d3b7-a3fc-4a92-a138-97465a1fe086 · outbound

This paper cites Open Problems in Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Open Problems in Mechanistic Interpretability

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.003123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.003123Z digest=sha256:f325d17bb20d0ef5d5625300280a2091e852867808efcaf29ad61b4f324f1aad

Observation e07c0ebc-f326-427a-a03f-bbc9cbaf8612 · outbound

This paper cites Hypothesis testing the circuit hypothesis in LLM s.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hypothesis testing the circuit hypothesis in LLM s

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.296737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.010034Z digest=sha256:f8628535798f670bf984492a2dd4eed725a0da0d34217c6c22d5071707517680

Observation 1d512897-3ac1-4f50-b804-40995ead7c76 · outbound

This paper cites The golden mean of scientific virtues, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The golden mean of scientific virtues, 2024

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.276920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.016375Z digest=sha256:60a567f94efffad05abac505960504976968ef58b13bb4074cd6efcaf2d11467

Observation 1d7ce50c-4f21-4f8c-a2d0-82bfdc5f1ff6 · outbound

This paper cites Knowledge in Perspective: Selected Essays in Epistemology.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Knowledge in Perspective: Selected Essays in Epistemology

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.259403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.024761Z digest=sha256:0241b9db492b3d359c2879ade3103a547e59709ebf354e484dbb87ef0a2678fd

Observation 352006f4-c9b9-4bf7-bc3c-71372f66f6b8 · outbound

This paper cites Grokking group multiplication with cosets.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Grokking group multiplication with cosets

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.240232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.031559Z digest=sha256:148287255ecd014e604d6555ed4d6a7764529406f0d94ad4d5df34735c9688d7

Observation ca0eab37-098d-4e05-a20d-024b68c63bb8 · outbound

This paper cites Simplicity as Evidence of Truth.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Simplicity as Evidence of Truth

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.218842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.037028Z digest=sha256:9d9bf9b927e3a181b16560bde71946b64bee63944a144c728e996e7cbf4160bb

Observation 2e2cb470-0087-4b83-a9c2-b26796256c1f · outbound

This paper cites TracrBench: Generating Interpretability Testbeds with Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii TracrBench: Generating Interpretability Testbeds with Large Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:25:06.330001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.043348Z digest=sha256:6cbed16e80ee25d6601b3c2428cd3046ecb99f2ea02e28204af1868651142a9d

Observation 0a2c7ecf-696a-4a1a-9927-dc96616081c7 · outbound

This paper cites Interpretability in the wild: a circuit for indirect object identification in gpt-2 small.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability in the wild: a circuit for indirect object identification in gpt-2 small

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.193035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.049199Z digest=sha256:734d7ee98aee13d6f7593717a3453f5c96d344ce6ed058d896d46b8ba6043282

Observation 57fadde9-0811-439f-84ad-fa1f01846649 · outbound

This paper cites Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005

Reference 92

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.202634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.057524Z digest=sha256:9f96e53086630c3a4088463b0dec50604a34db5d4b2249e63630b38caad4a6ae

Observation b37e3d23-7ad3-4fdd-b774-1a560c33b6d1 · outbound

This paper cites Understanding as compression.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Understanding as compression

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.173192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.064714Z digest=sha256:55e6cd7c1cc5723639fcf0b487d041acb54443eb0584f6059da58218e8a02801

Observation 095fdfb3-67ab-48ae-a712-b6d6f312a35d · outbound

This paper cites Geschichte und Naturwissenschaft.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Geschichte und Naturwissenschaft

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.154724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.071370Z digest=sha256:d2998cbfcd984fb9b4e1480da39ff193943f45453e0e086fcd734e48ef4e0124

Observation 721adab6-90b7-48eb-81f9-c03e9c9c1263 · outbound

This paper cites From probability to consilience: How explanatory values implement bayesian reasoning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii From probability to consilience: How explanatory values implement bayesian reasoning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.136110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.076238Z digest=sha256:3d55b2757cea67488e11715167cf9c207be10a349738febcb0c26d508ff74834

Observation 0c046019-3c7f-41e4-9991-f9a048c1a470 · outbound

This paper cites Woodward.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Woodward

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.111381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.081518Z digest=sha256:f170e613105f108b41dfb551be157555055e48d7a6f730be82b172695c0a80a9

Observation 3ff95413-4b01-4574-8af5-5fffe4e5c756 · outbound

This paper cites Towards a unified and verified understanding of group-operation networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards a unified and verified understanding of group-operation networks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.086534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.086534Z digest=sha256:21ef1a2d9dead00f6fdca171a78bbd9dc2cde1bceeec77554024d82e8ea9c1bd

Observation 74a24061-0e17-4599-a309-ec652df7de8d · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.093358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.093358Z digest=sha256:f779efc66daaa5a4de8b677bb985d6b8180f12a037716dbdf587e9a6ad1f30aa

Observation 5268333d-c642-4bbb-b2eb-f957abac149a · outbound

This paper cites A theory of usable information under computational constraints.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A theory of usable information under computational constraints

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.101369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.101369Z digest=sha256:466c27b2c22289d3b445aef3e8467eeafe537901e455c906428e71ec346e18a7

Observation 7474a404-cacd-4e6c-bd4e-726f49f34e19 · outbound

This paper cites Locally decodable codes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Locally decodable codes

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.076567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.108473Z digest=sha256:a37dd7ae00ebe7a859c3542a89354d632e9539d0e8277c634e9e5f35d521fe5d

Pith citing papers

Observation c7a68c25-a842-4499-bca8-0cea977ccc1b · inbound

From Mechanistic to Compositional Interpretability cites this paper.

From Mechanistic to Compositional Interpretability Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.444788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T02:42:26.173782Z digest=sha256:2a41e3171478c251c26d0a0a9e8288157235fa5d25323ae23d1c5c5197bac982