Pith. sign in

Paper Citation Record · LEDGER

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

As of 21 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 1 inbound Pith citation observation for arXiv:2505.01372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01372 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:25:06.108473Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:42:26.173782Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact5
  • verified fuzzy33
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7be2a119-13da-4a7a-a1d7-649e3984d9d1 · outbound

This paper cites The Computational Complexity of Circuit Discovery for Inner Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Computational Complexity of Circuit Discovery for Inner Interpretability

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.518727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.518727Z digest=sha256:1d7874d6fa1b750f2e85e0cfc8863d4b1902111e40a35fe3330dda79da8f4f3c

Observation e96be6eb-3233-415f-bede-dba250dc41a7 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.526366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.526366Z digest=sha256:acf8b783968fb7c8c24bef66b5eaeded365d3b36a54c0a81fbcdfc240c0a9d0d

Observation 61d387c8-950c-4507-82b2-d3eb05f6a283 · outbound

This paper cites Physics of language models: Part 3.1, knowledge storage and extraction.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Physics of language models: Part 3.1, knowledge storage and extraction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.531896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.531896Z digest=sha256:753ccc0a2238abaeda3e954f6b671eaeb7c29e3925108a64be561b71e3739c57

Observation aae00ee6-8ee0-419b-adff-b23b52bf598b · outbound

This paper cites The urgency of interpretability, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The urgency of interpretability, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.538644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.538644Z digest=sha256:6b2bab85f36daba205f3521ad1a320ce5c48c90245f1967e7434ef2d92352716

Observation deff4b33-b745-4e3a-9e7a-86e1815bb557 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.544723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.544723Z digest=sha256:4db280eb7aaa85a8c4076471eb0f25c6f8832d6e21ba2b13ac4acf35fb736c01

Observation 964da59b-029d-4588-8631-c9b228d56532 · outbound

This paper cites Ai as systems, not just models, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Ai as systems, not just models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.550537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.550537Z digest=sha256:3b602009ca76329bed27b1df563d98326b20d6c047649d8fbe250a06b920747b

Observation 7c179344-ca3a-48d8-82d1-3ba93e274d41 · outbound

This paper cites Standard saes might be incoherent: A choosing problem & a “concise” solution.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Standard saes might be incoherent: A choosing problem & a “concise” solution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.556998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.556998Z digest=sha256:bbd3b3a3f0dd2a341b34e57b16209b60df681c928c661748a2a0b7349d8bd54a

Observation 1000c22c-0b9a-4e30-8e2f-592afc155ec8 · outbound

This paper cites Position: Interpretability is a bidirectional communication problem.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Position: Interpretability is a bidirectional communication problem

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.563027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.563027Z digest=sha256:f074d276b4d87ca58068e245eaf764c0bd392b2c797c15923e3166d236dd2c63

Observation 237b7281-268c-41a9-b03e-66dee8e8639d · outbound

This paper cites A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A mathematical philosophy of explanations in mechanistic interpretability: The strange science part i.i, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.569755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.569755Z digest=sha256:cda0bdf02cdeaddef5dc7d10ad8c1f5aa191be150091be1a4da770ea75c87291

Observation 2f798255-8dd9-4727-8730-208b6e6b64e5 · outbound

This paper cites Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.575857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.575857Z digest=sha256:990d316a50bbdbfe76717730d11663ffe8bc4a61368ec51108e6929807e68b0d

Observation e326204e-4a5b-4139-967e-9d642129ebdf · outbound

This paper cites Novum Organum.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Novum Organum

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.581703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.581703Z digest=sha256:baddc0b175bbbb0cd8d0d850046c53d4a9cc21b88aad6d5c763966034cd872c0

Observation 221e6ee1-3828-4e38-ba66-1b3543b46e24 · outbound

This paper cites Simplicity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Simplicity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.586827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.586827Z digest=sha256:09721354ea05cfef69f1c003e185c8538b44b47d12d1d8e8a204a8072e9faeef

Observation becd414b-4fff-4ff3-954f-598d5526b840 · outbound

This paper cites Design Rules: The Power of Modularity Volume 1.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Design Rules: The Power of Modularity Volume 1

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.592337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.592337Z digest=sha256:6f6fa2dda452cace5ba5c6b47e1a96eb525fe7c79db9cf829e11a065e06037fa

Observation 98a5edf3-3bec-4513-b1e0-5c1ecb7f32cc · outbound

This paper cites Icml 2024 mechanistic interpretability workshop, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Icml 2024 mechanistic interpretability workshop, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.597081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.597081Z digest=sha256:c3213156c9d3260bd8092fb85670f2043c0dcaee420ab70b918c0a844b36201b

Observation 64e4beca-6ae9-4858-b525-195cbebf0549 · outbound

This paper cites Local vs. Global Interpretability: A Computational Complexity Perspective.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Local vs. Global Interpretability: A Computational Complexity Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.601776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.601776Z digest=sha256:f7ba657f596d23f58392ee2bb8d3a62e4168648d2c8e22310e67f546b45750f1

Observation 3fa773a3-d922-4b64-81cd-e726fb950bf8 · outbound

This paper cites Explanation: A mechanist alternative.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Explanation: A mechanist alternative

Reference 16

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.262246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.606747Z digest=sha256:a6df09e5c46bfbeb8b3eebd126226fa8c2cb09681bf986c7428776770bcba0a3

Observation 3c2e7bee-d2fd-4024-bc7c-1cb2370aa993 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.612261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.612261Z digest=sha256:2b65990aec50e1eee582122e2590d7302e0a82c43b95ee1cfdb9d304a9fcffd6

Observation c90eb841-85a5-44f5-b15d-ee4d768cc0c8 · outbound

This paper cites International AI Safety Report.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii International AI Safety Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.617272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.617272Z digest=sha256:f4dbd70d33f5b58a4949ae5d9602d32bc7153a4540d655d3bcab32ac17e71357

Observation ec496079-ff91-465c-b9db-71c5d491d4cc · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic Interpretability for AI Safety -- A Review

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.622447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.622447Z digest=sha256:c5982f9f25c4af5e815e41f75773b76eab3bf1c5bc9f3221725b9d9b1d70ca96

Observation 03f4349d-828d-465b-9065-1c382d45670d · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.239659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.627697Z digest=sha256:7f5405229ba8f99de6d54aad0449f1620135e794f08f3a2ec481cfe08fd78076

Observation c960ebea-2022-4f48-a643-0b57db62aa68 · outbound

This paper cites Auditing local explanations is hard.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Auditing local explanations is hard

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.632693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.632693Z digest=sha256:6cb00f5d883edba83954b7e1545fb0dd441d871dd76c9715f3936d3f6f3901af

Observation ed2767c7-3100-4b87-beb9-87dce637ba3a · outbound

This paper cites Language models can explain neurons in language models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Language models can explain neurons in language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.637441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.637441Z digest=sha256:e424217cef26fb369ead2da5157079147eca04c329283ab57e596be025297700

Observation 87cc22af-9945-4683-b20e-ef73fe38ec6a · outbound

This paper cites An Interpretability Illusion for BERT.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii An Interpretability Illusion for BERT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.642287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.642287Z digest=sha256:5e8510bee79f8d639b0a7f7404b82ed6afc5ceb23a844ca4e698eac9eb03b420

Observation a2b41a39-4f83-4cfd-a9f0-3663c232c3da · outbound

This paper cites Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.647126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.647126Z digest=sha256:d059bd5cb4081a86a8ac2d9cf1b10563aa315e125f4cebe347166739f60fcb42

Observation 3a76e277-6588-436b-b83d-5be24a5b09ec · outbound

This paper cites Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.652281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.652281Z digest=sha256:a9b7163da71a9011ddcd33983efa5603c18ade2cb199364410242a775677b716

Observation 135608c9-8774-44fb-bf3e-18692e656f90 · outbound

This paper cites Towards Monosemanticity : Decomposing Language Models With Dictionary Learning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards Monosemanticity : Decomposing Language Models With Dictionary Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.658022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.658022Z digest=sha256:3b2373563bae1c11c9acecc77976580605cf6ea11dc74a3861ade43c55b87865

Observation 34747481-f21c-4452-9e82-01aac17a063e · outbound

This paper cites Propositional Interpretability in Artificial Intelligence.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Propositional Interpretability in Artificial Intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.663766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.663766Z digest=sha256:52adb9d1d5c86a2057277134167c2885c8b0d2c28e6ea2d5c4f076dbf70843b5

Observation 8c79eff4-85ec-4984-8618-b2188376c0c3 · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A toy model of universality: Reverse engineering how networks learn group operations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.669143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.669143Z digest=sha256:947b0561bc0f3ad8b4ed8e9078ac7f5555f2aaa2c4d231f765e1ce2b60d6c3b0

Observation ab6e7c15-147c-49e3-a67e-2ee2331be54d · outbound

This paper cites The evolutionary origins of modularity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The evolutionary origins of modularity

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.674171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.674171Z digest=sha256:33c4e17e8d9c45688f49f9136dba3e70da0d344d1cef801d766ff0307e8dc62b

Observation 0bbc4ffc-89aa-4a8a-9228-f21bfc00c894 · outbound

This paper cites Towards Automated Circuit Discovery for Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards Automated Circuit Discovery for Mechanistic Interpretability

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.678871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.678871Z digest=sha256:9d6dba39ee280f40e280db0f5b21c91172b7963b74fafbba304a2dffde5a3d8f

Observation 288ced46-da29-40db-ae02-4244fd08479c · outbound

This paper cites Central dogma of molecular biology.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Central dogma of molecular biology

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.684525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.684525Z digest=sha256:ab0a2cca0d39e850ad99a093dba21c8d1d2639946fc201580d942a7207fce6bd

Observation 0076b1cc-574b-4515-bbd6-8b68134604ec · outbound

This paper cites The beginning of infinity: Explanations that transform the world.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The beginning of infinity: Explanations that transform the world

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.690760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.690760Z digest=sha256:f3fa33af5a06ad68a588eaf7525700c5d77f3a1c61aceec3d9340debdaed68f8

Observation 9e2e0718-96c4-4dee-82a6-24b4caa1f973 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.696683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.696683Z digest=sha256:e79cd389ab66d1caf68db9cd26421c03929f7efa15ba62ab28564a1630f07dd8

Observation 1643f022-9fd7-4021-af3f-81f37f4af433 · outbound

This paper cites Einstein.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Einstein

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.702151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.702151Z digest=sha256:a9b959f1c8c0210395729c8659cf347237a66eec2c2564337e7ec55fc700e3d3

Observation 6badf7f6-0bd7-48dd-b6d1-dbc160e7f37b · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.707415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.707415Z digest=sha256:c6832e55d1a459c6d0f85a72985abdef7158517ad3e1199487596a8e11b88b7f

Observation 465f48a9-16d1-4854-9556-2b1edce37af5 · outbound

This paper cites Clusterability in Neural Networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Clusterability in Neural Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.712412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.712412Z digest=sha256:9c3fa1ba5eb007ce10e0f1581d910767bfce80774f9264d76a55330aaf05fe46

Observation d72172f6-6ee7-4cf2-a87e-c8aa458a6a2e · outbound

This paper cites Interpretability illusions in the generalization of simplified models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability illusions in the generalization of simplified models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.718000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.718000Z digest=sha256:cbd519ec74109a7f05bd3a4de880b930536d24242995452976f1bb802c4e4eea

Observation f97a349d-416e-4ec6-ba95-63ed4aa9bd7f · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Scaling and evaluating sparse autoencoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.723010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.723010Z digest=sha256:68af794370efb77856fb82a38a12bb99bdc0671a1403d1a3d34f1343d449f317

Observation 712695dc-797c-4577-a954-6c0f803ec29b · outbound

This paper cites Causal abstractions of neural networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal abstractions of neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.728174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.728174Z digest=sha256:b4d6dae588b7e0670e21a241433ad7327afa60220a6806e5accd59cb85ade0da

Observation 22aea982-0987-44f7-ab5b-592ababe55dd · outbound

This paper cites Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.733991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.733991Z digest=sha256:bb257df63e06b21a497365b2dabbaa912ee1e125cd9c9b3049d02d080c70108c

Observation 25eafe53-c2b4-44c6-a494-fd448e1bf91c · outbound

This paper cites Clustering algorithms.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Clustering algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.740650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.740650Z digest=sha256:b61ed8a59540709c8a6b5aca67dde70fef42e929b046a49fe070612dd38fa9da

Observation 5d25dd86-a4ae-4bdc-bad7-135601377bcb · outbound

This paper cites Compact proofs of model performance via mechanistic interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Compact proofs of model performance via mechanistic interpretability

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.746673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.746673Z digest=sha256:2dad1891d94d46a4ace2730829ccdf098b0fcefd155a4b63be560d11da0ebb8c

Observation 33317061-ea4a-44d4-a1cc-f3220fce6bba · outbound

This paper cites Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.875795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.752590Z digest=sha256:929923ebcc35fb130a153686e32abded1aa6187b2db7865fbb30182fdb084bbd

Observation 6c1fbada-0a31-460d-8864-8ab9e751bc15 · outbound

This paper cites Hastie, R.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hastie, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.856843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.757293Z digest=sha256:61d8bf09adb3d9952f5eefae8d03e2eee9cc2a40105aff5b4aa52eb0cdd6ba50

Observation 6bb4f5b1-e9b0-4de3-b18b-8a9922be0e2e · outbound

This paper cites Hempel and Paul Oppenheim.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hempel and Paul Oppenheim

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.837001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.762261Z digest=sha256:9938303dc684d4112c935bf4eb2f7d1d5e2c1c49a76362ba5980028c00c34815

Observation 0fc0361d-3d60-41e5-9e5c-c8186b9233c4 · outbound

This paper cites Philosophy of Natural Science.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Philosophy of Natural Science

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.816805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.767579Z digest=sha256:6700e2b0837c718a122d7a883cb87922e0b6d86ca12f75cdbdd0108ce232a0a2

Observation e355e61b-deec-4081-8248-fe70d5984b8b · outbound

This paper cites Bayesianism and inference to the best explanation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Bayesianism and inference to the best explanation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.797867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.773005Z digest=sha256:c85ff476c2e9015a09921c47878f3c6d788301b5b96fc1d835f84ce62b6ff858

Observation b0510d13-33cf-4cb9-bfb5-b1bc7a98b276 · outbound

This paper cites Loss Landscape Degeneracy and Stagewise Development in Transformers.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Loss Landscape Degeneracy and Stagewise Development in Transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.779468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.779468Z digest=sha256:09e52ecdbfc9c93334338944caf1258e3e1819aeebae0efb0eebf8b809e56230

Observation c5ec30f4-92f1-40f8-9acf-a47b1637be70 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse autoencoders find highly interpretable features in language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.785160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.785160Z digest=sha256:a6d025bd4d8edcb984a2d1419187dae84a111642b05423d6e3e26d569b4f5188

Observation 40da69cb-7913-46ec-9031-94e2f0f81322 · outbound

This paper cites Hutter, E.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hutter, E

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.767475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.791367Z digest=sha256:6ad05efdbd04bc5793a11c6534c023ce39de260474b4b912f7839c06c45a1105

Observation 1c3adc74-c6de-4fdf-b320-a17fec806d6d · outbound

This paper cites Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Fine-tuning neural networks to match their interpretation: Towards scaling compact proofs, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.749269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.796480Z digest=sha256:d789243907dbcfef676845c445ec6c48e50ebe1fb075942bb7e5f379d7852b0f

Observation 7ff9effe-ae4b-46a8-be2e-3ce3887cac71 · outbound

This paper cites Kandel, J.H.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Kandel, J.H

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.729487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.801635Z digest=sha256:93fa8f1a01a7f8d32561fd9e99615bb004faf7ad60707990d71f29d5ebc89846

Observation 6d585e48-eb4a-4bc0-9aac-d340d76ab5f0 · outbound

This paper cites Saebench: a comprehensive benchmark for sparse autoencoders, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Saebench: a comprehensive benchmark for sparse autoencoders, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.705111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.807554Z digest=sha256:45cc89c37316f49025972b5e289abc7a896cd38cbbf5fdcd06807310ce929cc8

Observation fd84da5c-f9e0-43b8-96b3-ed972dd73d61 · outbound

This paper cites Kennefick.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Kennefick

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.687071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.814962Z digest=sha256:809998a3161587029171db666b8333cae972221ae1fe0d53753d27ce0be943c2

Observation 16b197c7-a52f-41ad-bccc-66f48285f750 · outbound

This paper cites Explanatory unification.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Explanatory unification

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.667802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.820191Z digest=sha256:6ee2d3152be51190b398c65e7b2636a69c84359c4b0a6f6149f1fd3b9f553f21

Observation 10e71510-9e4a-4a0e-ad3c-0aaf7a4baba4 · outbound

This paper cites Three approaches to the quantitative definition ofinformation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Three approaches to the quantitative definition ofinformation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.646435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.825015Z digest=sha256:2fe39f84efd31cc17336017cdbfba2d5ec99c87d92a49ed9494ddcce3853e13e

Observation 17746ef1-137a-4c4f-b7fd-8d3fa293643b · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:25:07.627680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.830341Z digest=sha256:773915ae677de4203a7f514ffd5d097d783488c9596ba127281569f5766a963b

Observation 1587381a-db0b-4c29-8c54-0a0c152ca444 · outbound

This paper cites The Structure of Scientific Revolutions.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Structure of Scientific Revolutions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.607038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.835464Z digest=sha256:939d7f4f19155adb17eb029a49efb6b21e53c939e5170b82c5cf38b3eda9b52e

Observation aaa33f02-0ff5-4a53-aada-a5d9be8f6d4f · outbound

This paper cites Falsification and the methodology of scientific research programmes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Falsification and the methodology of scientific research programmes

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.586373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.840481Z digest=sha256:b201c32e96979ba9ef5c55991b5c968296f2e88c2830de5228cc44c343b98298

Observation 91c964fc-0773-420c-8279-fc1eb53403b3 · outbound

This paper cites The Methodology of Scientific Research Programmes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The Methodology of Scientific Research Programmes

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.567371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.845034Z digest=sha256:11da4254bfe2a4b0af1172b76aa0d7433ddd4bc053a32d34bfd945f2d17e4ee8

Observation f9885b31-5dd2-4dae-97f1-f880fb428f8b · outbound

This paper cites Sparse Autoencoders Do Not Find Canonical Units of Analysis.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse Autoencoders Do Not Find Canonical Units of Analysis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.850864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.850864Z digest=sha256:d8726bbb48302c717453bf63ac386b6bb322ef2b5be90dd5465cd06ab17775ee

Observation f7a9f288-c896-4fba-9c44-d8753e20975f · outbound

This paper cites Lindsay and David Bau.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Lindsay and David Bau

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.856242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.856242Z digest=sha256:2a3ef6a1b790b85199f2903125cbe3d28935da5834a385da1c11d90e1ba08702

Observation b0ae1407-4a60-447c-8d2e-dada1f82ccfc · outbound

This paper cites The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.862843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.862843Z digest=sha256:5e7f0cde2e651398dc737ade7c64ef3af8f1c6f760508d3f444307f8e43ae0f9

Observation 9eb3975f-8e1f-4e66-a474-ec9c3b1a62af · outbound

This paper cites Mechanistic mode connectivity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic mode connectivity

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.539180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.868951Z digest=sha256:a701fed53bf1bae8ac366c18022202872b62e5680ec877b14788a2065d329b34

Observation ca10ed64-93ed-4ac4-a197-12840824058c · outbound

This paper cites Information theory, inference and learning algorithms.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Information theory, inference and learning algorithms

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.874084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.874084Z digest=sha256:ddb267e257325658caecbfba49d9fe5fe3628f6e22e6893aa4b396678f30d0f9

Observation 86793f78-8c9c-49a5-b6b4-0c5e352fa2df · outbound

This paper cites Is this the subspace you are looking for? an interpretability illusion for subspace activation patching.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Is this the subspace you are looking for? an interpretability illusion for subspace activation patching

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.507881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.879925Z digest=sha256:e0f398729c151a4cac64f71e6925d82b6d63ebb7886ba2c7db8646127c96d189

Observation 2085c2d2-ad9b-49d1-9c1b-412de83b2a98 · outbound

This paper cites Downstream applications as validation of interpretability progress, March 2025.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Downstream applications as validation of interpretability progress, March 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.486729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.885763Z digest=sha256:72bb885346afdef207abbfeb3fbb316879b648900259013030c5c451633cbb09

Observation 484ea65f-aec7-463d-80ea-4d5c9dc4fd29 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.891673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.891673Z digest=sha256:21212b4830c74ac35d425c836d18bb8c53a003067e90ede06484959453081b37

Observation eb401795-6468-4d96-8e49-cfb9a832798b · outbound

This paper cites Cognitive styles in two cognitive sciences.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Cognitive styles in two cognitive sciences

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.468750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.898471Z digest=sha256:80c8e4cec8fc871c34d3c06fefe7112b7884939b381f431c41e7969154a79494

Observation fb3e9be1-3e90-4402-bf7a-be0ded544915 · outbound

This paper cites Zoom in: An introduction to circuits.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Zoom in: An introduction to circuits

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.905678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.905678Z digest=sha256:7ae9557d01c4171c6f1b668f3186a1f945f03d166939790efdb3e12256d96ec2

Observation 536f7abc-d582-4db8-9057-fae24a5dba45 · outbound

This paper cites In-context Learning and Induction Heads.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii In-context Learning and Induction Heads

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.911761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.911761Z digest=sha256:358dcdfe8510dbda23082264d433386ba24d10c2dfc1e11fed81ee0190c7b07b

Observation 293d1e50-1950-4e08-84f3-51ad6d5bb269 · outbound

This paper cites Automatically Interpreting Millions of Features in Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Automatically Interpreting Millions of Features in Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.918441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.918441Z digest=sha256:1125b22a284560cb9e8a01feaa575c1515fe391f452d3a24031a4393bc284d58

Observation 3b0b26c7-9cc4-4b96-8dd8-8f02df58426d · outbound

This paper cites Causality.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Causality

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.925700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.925700Z digest=sha256:1396643511acd64475080149398799125759b622f686e1dd419d4d925142c6b4

Observation c6aebdff-6c3d-404e-b0fd-1062e6f53578 · outbound

This paper cites Poincar \'e.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Poincar \'e

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.416845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.932748Z digest=sha256:3ed2ed806ac9c4e2af62a57bd9aa9ea06c04caee3263ae61358b49af0b62aff6

Observation e2af5c31-1409-4f9f-b0c2-e63b82fc5688 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.944080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.944080Z digest=sha256:c9f9b1d372fb09a35ad86fdf92d7921c3a276a7ecb4f002cf3cd870004065d80

Observation 5b60c837-b55c-4963-8912-cee876b825cf · outbound

This paper cites Hume on theoretical simplicity.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hume on theoretical simplicity

Reference 76

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.221515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.951779Z digest=sha256:619622ca10a4ef9c947e5704419928fd0de2346f2515ccb26e328ed68e2070be

Observation e7de4f8c-c5f8-4c6e-bdb3-6d566e58bd11 · outbound

This paper cites Escalation risks from language models in military and diplomatic decision-making.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Escalation risks from language models in military and diplomatic decision-making

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.962513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.962513Z digest=sha256:88faffe6c6a7cf2c0276e3a27eb100dda4148742854e0b9ae81d6425e4f9a944

Observation 4747c415-0ae3-4b04-b1ee-2029e9bd4a1f · outbound

This paper cites Four decades of scientific explanation.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Four decades of scientific explanation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.384390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.968237Z digest=sha256:7922ac267586c095e18bcb99d1112f3de9a4ddb401239d39146226559b545c36

Observation 80943169-dfbc-466b-a923-efe22d9afad1 · outbound

This paper cites an unresolved cited work.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:25:07.365849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.973050Z digest=sha256:908ab010122f2ea520edf13b1f211be39feba1381bdec5f8166703bda66629c4

Observation 29c5b0f6-51c4-4735-99a3-2c9f05b2f048 · outbound

This paper cites Mechanistic?.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Mechanistic?

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.978424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.978424Z digest=sha256:526fc30cf61fa9cc0d50f5fd5ba04b4c110f5453de4a59610b9c9aad543ceb0f

Observation 08075a01-c91b-4cc8-8e58-12d0004ac73c · outbound

This paper cites Theoretical Virtues in Science: Uncovering Reality Through Theory.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Theoretical Virtues in Science: Uncovering Reality Through Theory

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.346087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:05.984022Z digest=sha256:c8cd3111e8d34fa047d1398c540269bf5ce8279da21c6db73bac1735d2fbca2a

Observation 740ce7be-8b7a-4b84-b752-dd27ef2f759c · outbound

This paper cites Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Riechers, Lucas Teixeira, Alexander Gietelink Oldenziel, and Sarah Marzen

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.990103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.990103Z digest=sha256:c351cd6aefd650bc6c5869b9b23c41e50dde8081f8ac4b104ebef1001519c38a

Observation 1b7789d5-9c8b-4ace-9512-ac09ab787734 · outbound

This paper cites A mathematical theory of communication.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A mathematical theory of communication

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:05.996982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:05.996982Z digest=sha256:04db55f0ef1c11af25f075a75429f9f64cca81c428a49e6d6e6ce2c8797f8061

Observation b417d3b7-a3fc-4a92-a138-97465a1fe086 · outbound

This paper cites Open Problems in Mechanistic Interpretability.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Open Problems in Mechanistic Interpretability

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.003123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.003123Z digest=sha256:009acdc9b8691008eb4777ebc16fa9f2e6855f2fd141187079878d3f895f29ed

Observation e07c0ebc-f326-427a-a03f-bbc9cbaf8612 · outbound

This paper cites Hypothesis testing the circuit hypothesis in LLM s.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Hypothesis testing the circuit hypothesis in LLM s

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.296737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.010034Z digest=sha256:95fd0b9bc0025a104e644eb9b55709979a1fe67c256106d1074b9f5154cf60ce

Observation 1d512897-3ac1-4f50-b804-40995ead7c76 · outbound

This paper cites The golden mean of scientific virtues, 2024.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii The golden mean of scientific virtues, 2024

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.276920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.016375Z digest=sha256:7417051a601eda808779e542fbf9e6098897d479e351f0e2d5d30d2905a8008b

Observation 1d7ce50c-4f21-4f8c-a2d0-82bfdc5f1ff6 · outbound

This paper cites Knowledge in Perspective: Selected Essays in Epistemology.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Knowledge in Perspective: Selected Essays in Epistemology

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.259403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.024761Z digest=sha256:dbbf7f8d3dd42ee50d25388d1042f71fdb8c35b073d9a5ee9accda8dda4a179c

Observation 352006f4-c9b9-4bf7-bc3c-71372f66f6b8 · outbound

This paper cites Grokking group multiplication with cosets.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Grokking group multiplication with cosets

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.240232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.031559Z digest=sha256:8faeb1c6ec4aa56baafa55da05001a894c5dc9ea285b0cf8e9f6abc510656bb2

Observation ca0eab37-098d-4e05-a20d-024b68c63bb8 · outbound

This paper cites Simplicity as Evidence of Truth.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Simplicity as Evidence of Truth

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.218842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.037028Z digest=sha256:e4e369325bc8f85587892518a9d33e1ec4a1212aeef2559b4444e7798307bafc

Observation 2e2cb470-0087-4b83-a9c2-b26796256c1f · outbound

This paper cites TracrBench: Generating Interpretability Testbeds with Large Language Models.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii TracrBench: Generating Interpretability Testbeds with Large Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:25:06.330001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.043348Z digest=sha256:c6dd644bcbca28dda8b06acd235352040ea0e98ac583073a51ded047d91bb189

Observation 0a2c7ecf-696a-4a1a-9927-dc96616081c7 · outbound

This paper cites Interpretability in the wild: a circuit for indirect object identification in gpt-2 small.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Interpretability in the wild: a circuit for indirect object identification in gpt-2 small

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.193035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.049199Z digest=sha256:a5dbc2c3c310955725b570b3ea57cc55d48f01709f944b85be3c461cb3f8e136

Observation 57fadde9-0811-439f-84ad-fa1f01846649 · outbound

This paper cites Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Why favour simplicity? Analysis, 65 0 (3): 0 205--210, 2005

Reference 92

Resolution
verified exact
doi, observed 2026-08-16T04:25:06.202634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.057524Z digest=sha256:f072400a3ddbcb3eea9f370147cf8fbd38b36d2fb7038c239e2f79fb1a1d6f7d

Observation b37e3d23-7ad3-4fdd-b774-1a560c33b6d1 · outbound

This paper cites Understanding as compression.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Understanding as compression

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.173192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.064714Z digest=sha256:6cc1f89700f75227f28a9f0261b7dd2e47d2b5afd6e12711f9e8bad0ecfed59c

Observation 095fdfb3-67ab-48ae-a712-b6d6f312a35d · outbound

This paper cites Geschichte und Naturwissenschaft.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Geschichte und Naturwissenschaft

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.154724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.071370Z digest=sha256:734982de94a55592bc2b88ed730ff5cbc081aaeab06ce9b143989da917fff381

Observation 721adab6-90b7-48eb-81f9-c03e9c9c1263 · outbound

This paper cites From probability to consilience: How explanatory values implement bayesian reasoning.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii From probability to consilience: How explanatory values implement bayesian reasoning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.136110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.076238Z digest=sha256:32bd8cc676f29047b5cd2db810a0f1296f364bf082406e5655ea2018b4112141

Observation 0c046019-3c7f-41e4-9991-f9a048c1a470 · outbound

This paper cites Woodward.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Woodward

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.111381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.081518Z digest=sha256:be4e6245ece0b274ee4f99c62cc4c3700222d9e0b179505f246b77ced7af840b

Observation 3ff95413-4b01-4574-8af5-5fffe4e5c756 · outbound

This paper cites Towards a unified and verified understanding of group-operation networks.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Towards a unified and verified understanding of group-operation networks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.086534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.086534Z digest=sha256:2cc8b298e2b371f6528412132333cbc9c0df4b80ab99649e45c9688dc6280554

Observation 74a24061-0e17-4599-a309-ec652df7de8d · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.093358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.093358Z digest=sha256:8429ef0b828373cc217e0bff8e37625c6a1625ee6e768fa7134bbf14ed32f42b

Observation 5268333d-c642-4bbb-b2eb-f957abac149a · outbound

This paper cites A theory of usable information under computational constraints.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii A theory of usable information under computational constraints

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.101369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.101369Z digest=sha256:f78c7a0a4cd99b9522e97cefc1e3175ae3bd526d1260af797078c3bec8319e38

Observation 7474a404-cacd-4e6c-bd4e-726f49f34e19 · outbound

This paper cites Locally decodable codes.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii Locally decodable codes

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:25:07.076567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-16T04:25:06.108473Z digest=sha256:26206b57dfae905cb54683cdb9e4f12ba1cdc4496bca02ff0ce4b54c6c87aca1

Pith citing papers

Observation c7a68c25-a842-4499-bca8-0cea977ccc1b · inbound

From Mechanistic to Compositional Interpretability cites this paper.

From Mechanistic to Compositional Interpretability Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.444788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T02:42:26.173782Z digest=sha256:5257f6315ad5885b60c218f0d41fa228897ca473337e3c5263d53519395315e3