Pith. sign in

Paper Citation Record · LEDGER

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 100 inbound Pith citation observations for arXiv:2211.00593.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.00593 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T17:13:51.408311Z

measured 166 of 166 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 137 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:18:23.641247Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact13
  • verified fuzzy38
  • unresolved4
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch10

External citation measurements

50
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 98497d3c-b10e-453e-bf8d-596debd99853 · outbound

This paper cites Language models are few-shot learners.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.545755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:2f317d26db440a190fb06be275a1d22ff59325eaff42dcd0c0152dcbd30ac4f4

Observation f08b166d-a709-4013-bfc6-bfdd73254e67 · outbound

This paper cites A literature survey of recent advances in chatbots.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A literature survey of recent advances in chatbots

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.549189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:ec49daa14888be61bd736d6fe320d4c2b3f785d0c660e7f36e029590ba5999bd

Observation 37bfacbe-9204-4eb2-bea1-7d8fd4956da3 · outbound

This paper cites A mathematical framework for transformer circuits.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A mathematical framework for transformer circuits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.550951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:1c9e8098221e3f89129f0444aec7ba2a91d1ddd3a01dd9b7b002f505bef87563

Observation 279f0de2-acd0-49f9-a6b9-22cbb44214f2 · outbound

This paper cites Causal abstractions of neural networks.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Causal abstractions of neural networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.552747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:467bc2d36c8d9c41cb5dfdb9e0a1e25077e60280041874592275473a56ad3ef7

Observation 9f534c03-5115-4764-b4f2-8f1ad9c98b99 · outbound

This paper cites X-Risk Analysis for AI Research.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small X-Risk Analysis for AI Research

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.482305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:7b8d9fe632aa7888e36ed6e39c6943d03dd4db5dd3556b0b8a0ed5294231cda6

Observation 2c858597-b301-4dc6-906d-460d218ead17 · outbound

This paper cites Natural language descriptions of deep visual features.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Natural language descriptions of deep visual features

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.554390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:29dac04a77c872b3824597d0570859d0e61f66f88c0e16da20f9050a5fa40114

Observation a30b5c23-be6f-47a9-a6e5-036ea59af1d5 · outbound

This paper cites Are sixteen heads really better than one? Advances in neural information processing systems, 32.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Are sixteen heads really better than one? Advances in neural information processing systems, 32

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.556608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:92f70d256665c7fec08dba843d381e59c81e80935866a6b8ff53fce90f69f06f

Observation e7b7134a-9388-4c94-94a5-4c1dc9906e5f · outbound

This paper cites Compositional explanations of neurons.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Compositional explanations of neurons

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.558173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:9e6c371e884f7ea27d7ffd9d9f64b3661037932cd653110d292176c10fec02e6

Observation dd0c2e2d-b8be-4964-881f-6ac73b41eef2 · outbound

This paper cites A mechanistic interpretability analysis of grokking.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A mechanistic interpretability analysis of grokking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.559969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:f79e970d4b416d517afd3471dbd433d5c432ea40d356f1e6996c35e0522cdffb

Observation 403d4df1-c521-44a9-b835-c04f501e2754 · outbound

This paper cites Mechanistic interpretability, variables, and the importance of interpretable bases.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Mechanistic interpretability, variables, and the importance of interpretable bases

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.561992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:cf71838fe6c5988aa01b266a8939d5819403db690250d28ac9ba7651b022f040

Observation f8d5f64c-cf8f-4a71-8f2a-54eba66d1f20 · outbound

This paper cites Distill , year =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Distill , year =

Reference 16

Resolution
metadata mismatch
doi, observed 2026-05-13T17:13:51.448603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:f39535795ed56c781f2465d8d7aa19be17999214f0cdf0cf7f5e70619c8f75ca

Observation f436f922-3967-4911-80db-da477de89f2e · outbound

This paper cites Language models are unsupervised multitask learners.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Language models are unsupervised multitask learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.564180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:f3fe553b8f7a457310259687a8b044f35f0263b9e5b52fd42e442831ea38aae5

Observation aca4400f-afb2-45ac-a7bb-c249b2619fb9 · outbound

This paper cites Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.504124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:481151a9e750b1b5203aca5a3f8852ec6790b1076ca9c4157979d2f29f755df3

Observation a96bd0e0-3982-45cb-93ba-5e824f313142 · outbound

This paper cites Attention is all you need.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Attention is all you need

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.566263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:094a6e9ed49493840f8274d7b5be83512c986fd9f6a379d85a3e87be13b2b2cd

Observation 3169fad0-860b-48ac-b990-01b8e1e848e4 · outbound

This paper cites Investigating gender bias in language models using causal mediation analysis.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Investigating gender bias in language models using causal mediation analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.568214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:d05ed0b14d43f7aa670c4ad2b800a4e7a62b1408d16a800a1e3240aa0cd15695

Observation d0cd3606-387a-4966-b84e-6cd3416991a9 · outbound

This paper cites Emergent Abilities of Large Language Models.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Emergent Abilities of Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:13:51.497753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:1fad5a56ec23171b84a04397a730b936fa6673b1d308584b8a1c5d66dde9839e

Observation 4e3bd3c9-e00a-434a-a5a9-f123127dc550 · outbound

This paper cites Shifting machine learning for healthcare from development to deployment and from models to data.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Shifting machine learning for healthcare from development to deployment and from models to data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.569974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:e4b10679ac136c87c85c594e0133a905c35974c51aefbedc89025d960d1ce73d

Observation 17f4ee92-3301-4a0c-8146-1327c7c89752 · outbound

This paper cites Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.442503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:b0cb840f0409e894d72d6d6b5bc1861962a9fa7bcb61c091b8ba52a9378eb214

Observation 2d3a1646-5ffe-4ddb-a530-9b526f550ba3 · outbound

This paper cites BERT Rediscovers the Classical NLP Pipeline.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small BERT Rediscovers the Classical NLP Pipeline

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.446046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:6b668ff2553cbd0669f7b1e1aeff4eac824fa479e1d89956e578f24c6662e2d7

Observation 3d26bd71-74cf-4a5d-8c6c-369a06076975 · outbound

This paper cites 2019 , subtitle =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2019 , subtitle =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.571557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:f218da05d18a223851e13a32d08ce6354a0256f055d54ddd8706545578f83f28

Observation bf0796bc-9d44-4fb9-8347-a1754cf14598 · outbound

This paper cites Learning to Generate Reviews and Discovering Sentiment.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Learning to Generate Reviews and Discovering Sentiment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.452248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:f3818d6bcda1c3275ab299a07f3d254e10a5e1da6856f5aa162878e922c894ba

Observation cd5c7a01-5b11-47d7-8b76-2285a6854c9a · outbound

This paper cites Implicit Representations of Meaning in Neural Language Models.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Implicit Representations of Meaning in Neural Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.455778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:c3f6358fa2ac2f76b72744d5874296323366ae49372c6aa3ce6639fe07fd7a2a

Observation a175f726-d97d-4094-8d9e-db3489192e8d · outbound

This paper cites An Interpretability Illusion for BERT.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small An Interpretability Illusion for BERT

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.507126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:e50ac096d69d3ef54a0ce7ace8d68d4f73df526e8b769ad646b095c5f11c1d05

Observation 7a387258-f02b-4a01-ae25-77cd1a6846ed · outbound

This paper cites Proceedings of the 57th Conference of the Association for Computational Linguistics.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Proceedings of the 57th Conference of the Association for Computational Linguistics

Reference 30

Resolution
verified exact
doi, observed 2026-05-13T17:13:51.458733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:ae3c35323f57350955e91f59e8000ee58d145c53e1e8efe7f0d2a90a10160c06

Observation 68a0559c-f58b-49cf-9553-2b7f506d0b5a · outbound

This paper cites International Conference on Machine Learning , pages=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small International Conference on Machine Learning , pages=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.573549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:ec57a53e77e28d69d985d9094d7e4641680cf605e126382f2dc30ab8a1c3bfd6

Observation 02091e11-3423-48e0-ae27-57225f28dfd2 · outbound

This paper cites Unveiling Transformers with LEGO: a synthetic reasoning task.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unveiling Transformers with LEGO: a synthetic reasoning task

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.491461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:6a340b8c2700e678b5bce47f9040591bef1da13621e23d3d73cf00b618c5eeb1

Observation fbc2d07d-597a-42d0-9b66-0e56c0f3a6b1 · outbound

This paper cites Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.494465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:8c0e6cfca0dc25f758b09c45b2a329f92dfe4de18f74901683aea3aff6f0105c

Observation 7f3fa369-c0d2-41a8-bb75-9482dfaf0adc · outbound

This paper cites arXiv , year=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small arXiv , year=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.575342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:36c20cbbbd182af1d655c2e7f4ce245b33a29c164aa70787aa51683eaf409b3f

Observation a4280cf2-37d2-4e9a-88ec-6d2c4dd77aa7 · outbound

This paper cites In-context Learning and Induction Heads.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small In-context Learning and Induction Heads

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T17:13:51.501018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:0c272c121e19068ed31074f8389722369c32cd9a5a382cae3f4ec1af82a6510f

Observation d0820cd9-5070-49fe-8f27-45396e23d4e7 · outbound

This paper cites Information , volume=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Information , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.577082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:2bc38f0c0dd26dba3548561d47bb238fe8ba7dfb34b19934286e7011556bfbfa

Observation 5d100709-e2bc-449b-963f-4322db7bd8d1 · outbound

This paper cites Nature Biomedical Engineering , pages=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Nature Biomedical Engineering , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.578696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:40636747ef17e65c43fd29579d3b81ef025c54882346c34f93421ece39655e67

Observation 63ae8442-2c60-4ab8-a8bd-741f909d8abd · outbound

This paper cites ArXiv , year=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small ArXiv , year=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.580503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:4480bf23a20ebde23413ff116978fdcebac0c96bdc5fa4a8b3307abac8202ee1

Observation 0f9dba28-98ec-4cc6-96b0-f2beed018285 · outbound

This paper cites Unsolved Problems in ML Safety.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unsolved Problems in ML Safety

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:45:28.030299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:2b0daf92df97c461cc8dec170bb26be1abdbd289e864e96a02eb8a8ec39084c7

Observation 83cfc6fa-cb3e-40e9-8621-b986875ff7cd · outbound

This paper cites International Conference on Learning Representations , year=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small International Conference on Learning Representations , year=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.582070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:adbe81803dc62b5dc4585505acbba48959a9eabf6ff3b2319c75c9aea9d66943

Observation 44a715b7-5696-4efb-bead-36a3fdf5c9d2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , volume=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.583817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:f21acc076411caf084e6f647c59c8afe06c08cee4e639d1eb105aae45a027f9e

Observation 0bd8dc7d-0387-4449-b96d-a2ae2a04cfad · outbound

This paper cites Implicit Representations of Meaning in Neural Language Models.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Implicit Representations of Meaning in Neural Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.479376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:7b34f219b2351dba2affde68d7b52388b9de839e36e65c98abbb346eff62b7e2

Observation a1cb9589-f8aa-417a-9e46-03544b49338a · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.585785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:892e5e698bec435e41b0c143529c7220033268211634b34e6220f4568f88c41d

Observation 91650fe7-4249-4906-8bf1-82e44e07705b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , volume=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.587697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:a1a575a1b5754dd91746d9f77d19871711906782d381d8af98bdd04854bb2ad8

Observation 602f23f7-0f08-449c-adc1-1ddda406da47 · outbound

This paper cites On the Pitfalls of Analyzing Individual Neurons in Language Models.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small On the Pitfalls of Analyzing Individual Neurons in Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.467058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:c9f2cc88df113fe4302ce640d18053c461edaf05512f3ec48f1a8d6ca0fa1851

Observation 2ee3c801-d141-4173-bcaa-57024fc6ff31 · outbound

This paper cites Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.473392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:dd2b52f1756e45170caed1b58e75fe0d67a5ff86d37d6efd58026f6a89075de0

Observation 88621bf4-43d0-4af1-a283-d2ae7a49a5bf · outbound

This paper cites ArXiv , year=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small ArXiv , year=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.509351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:074f051ca97b64645ba9f17f6b7cc7cd9071af8ab4bbcfc0badd22522c896980

Observation 047ba2cb-c305-4d86-96f3-3e6a46377144 · outbound

This paper cites Advances in neural information processing systems , volume=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in neural information processing systems , volume=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.511277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:d85f08ce9ecb9442dca8d7ebb18cb19cd90cea5fe9c1b752ff49ab2fc5a50445

Observation c7072732-3db9-4ae0-b905-d61d79902e7f · outbound

This paper cites Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.437696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:f45c3b45ca7b6c00fa02faa008369862eba704477ac37a7d27db8c229eef92b5

Observation ffe00a29-866d-49e0-b687-5fcbbec3926f · outbound

This paper cites Analyzing Transformers in Embedding Space.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Analyzing Transformers in Embedding Space

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.476307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:0b45817bfe92ac24ff2e6bee196e7d475fe975dc7229a9b337d078a60023f3f4

Observation ea483a46-26a0-4d1f-82c7-33a2bd8eaa54 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Transformer Feed-Forward Layers Are Key-Value Memories

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:35:50.837220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:52a411328d27e7d3ebf99d34c0bbebeff866a58e14e8f4eaa0454e07d7122036

Observation 3ec52625-53b6-4a57-98cb-bfecf08c22ed · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Locating and Editing Factual Associations in GPT

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:58:57.694057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:20860deb9a78dfc6e3d6d6aeb043def918533bc8cbfe744c0d0d8aae1aeebbf9

Observation 2eeb5839-af2f-4a12-b879-b8301d1766ea · outbound

This paper cites Distill , year =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Distill , year =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.513258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:d2fbcfa3617e8812d09f89637d4ade50278f8fea7105672554720a85f11c077e

Observation aa4caf6d-a2c5-4beb-9462-bd4c3efeed6d · outbound

This paper cites Scaling Learning Algorithms Towards.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Scaling Learning Algorithms Towards

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.515215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:ede0812e652fa7954ddd7eb72336636b7c9f4b1aed3165f6369eaeb8a274e501

Observation 0bb6b7e6-a66e-430a-84ca-5f5a148e293f · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small and Osindero, Simon and Teh, Yee Whye , journal =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.517180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:16ef33db238887f72a5c615079767bc144b8e98d174926d1e739d128caaf156f

Observation c4d53185-b7e1-4f5e-a6d2-c3c023af842a · outbound

This paper cites 2016 , publisher=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2016 , publisher=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.519294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:845b709153b94ed3b849c48c2a025652a86c72b2e3ebef37982131d818adc0ec

Observation 25fd0f4f-fec7-4df4-9f3a-f64fef9fc1cc · outbound

This paper cites an unresolved cited work.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-13T17:13:51.521096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:ee36238a2e15efd8261d6abfe03c4a44a942f67ca98540c1b6e82b5c2cb60174

Observation 62c4ebf3-4e48-4c0a-9285-de5df982d558 · outbound

This paper cites 2021 , journal=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2021 , journal=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.523060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:7bbbe8b02eaff08879a8ab066817ff2f002f1e5557c1e1d90bd97428603c58c8

Observation c0987cb3-bed0-4b35-99b5-fc5ddbc65541 · outbound

This paper cites an unresolved cited work.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-13T17:13:51.524818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:4b544cf1275926104c5fe1a50b6a1ba0cf5a8588dff12b1d4ab68befe0243421

Observation f4af84f0-fff2-4cdd-8fff-2ac8df776418 · outbound

This paper cites Advances in Neural Information Processing Systems , editor=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , editor=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.526724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:09aec4de28963c1af5926a9b28425a51ccbbe133f0a9f494bf0049ddf9f5a00a

Observation b18e8659-7b93-4b86-9abe-7341cee8e46c · outbound

This paper cites Advances in Neural Information Processing Systems , description =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , description =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.528723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:3f5e17e4b54bc040f6df1642b0e04d9ff122ddb83c30dc7b0dc5f5f158b9f00c

Observation 0cc76915-2cdf-4107-bcab-6983442b219a · outbound

This paper cites Language Models are Few-Shot Learners , url =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Language Models are Few-Shot Learners , url =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.530521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:0401f63252c452dafcf7b53b22867b3aaebfc9c2d2b421ca874e8c50ad5bcf33

Observation 5991bd5f-0b1f-47cf-b598-1d91be5347a0 · outbound

This paper cites Distill , year =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Distill , year =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.532356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:d262bee518997d47e6f2a0c1be3f2cdcb5071504c3a880d876b6617126affd10

Observation b0c57649-e011-4981-a64d-b48d65a89eb4 · outbound

This paper cites an unresolved cited work.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-13T17:13:51.534412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:d221714282396212f0b2331375cdcb019839804054b950b37f5d9fd585f4c9a5

Observation c482a09a-28d9-46ee-ab98-a981aff2145a · outbound

This paper cites Sanity Checks for Saliency Maps , url =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Sanity Checks for Saliency Maps , url =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.537065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:134ad8d601068d546bd3208d865fe6e3b8dc279fa00a77e15fa055147bfba9da

Observation 1d6857b2-b1d1-47bb-a8fd-528a0afa2dc3 · outbound

This paper cites GitHub repository , howpublished =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small GitHub repository , howpublished =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.538794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:8ec5b0e1f68f55c05307964e486ba5657cdd46af68e5ba6f3ac5d67cf3192525

Observation 81bd157d-3446-4f4c-bed0-dc027aa9cdbc · outbound

This paper cites an unresolved cited work.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-13T17:13:51.540381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:ff1a0187f9e25c73611f781c9e7510eb35fce77c07a4b27780a107ce2b7572d9

Observation 15f9fe67-6c11-4ac6-ab47-fc57bdde959b · outbound

This paper cites an unresolved cited work.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work

Reference 68

Resolution
parse uncertain
raw_fallback, observed 2026-05-13T17:13:51.542011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:753b4d6e75bac7b1b94d74185640960d8b38a6cb157ac70699bf83392a0d4f78

Observation c49db7c6-8185-4259-95d3-9fff7cd80529 · outbound

This paper cites An overview of 11 proposals for building safe advanced AI.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small An overview of 11 proposals for building safe advanced AI

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.461849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:56a3b647547983e5f42ac1e62adfd5030f157534ed779f08196a7883e93ceb52

Observation 06f26e36-fbd1-47b3-aeed-72ee147eb75d · outbound

This paper cites A ttention is not E xplanation.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A ttention is not E xplanation

Reference 70

Resolution
verified exact
doi, observed 2026-05-13T17:13:51.464099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:fad79c4aad818f0da775239172b9bc841671ac4f456e1e8f3b8e163130d4ec76

Observation 878d1a18-f906-423c-ad06-2020161a1e8d · outbound

This paper cites Optimal Brain Damage , url =.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Optimal Brain Damage , url =

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.543568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:a3f8ac6405cf839f988c9b174eb4cc9c74b8b773a7785836a14b7e85278476bc

Observation 26bf98f5-93e2-45b6-8e3c-209314a33acc · outbound

This paper cites 2022 , journal=.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2022 , journal=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T17:13:51.547451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:4bf7e6e60e0b9e459b5658af1d7a4ac3ccb6fb26c651325df1f4b8a7565a6b92

Pith citing papers

Observation bd07232b-7a32-41ac-ac35-9ab414b28ddb · inbound

Progress measures for grokking via mechanistic interpretability cites this paper.

Progress measures for grokking via mechanistic interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:52:56.157111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T21:52:56.040569Z digest=sha256:9ddef16c314357eb5d3f5f54a8017726733c630664833f19a6ba41bf5d5d51ae

Observation 16867590-ebb7-473c-bb39-a4221f87c667 · inbound

Eliciting Latent Predictions from Transformers with the Tuned Lens cites this paper.

Eliciting Latent Predictions from Transformers with the Tuned Lens Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T16:54:37.382049Z digest=sha256:fb8adcc263a86afaf8e572323c92123e57900729c126b33fca2141de32dcde11

Observation 9b262013-a5bf-47ad-a29c-b8745b7cc891 · inbound

Sparse Autoencoders Find Highly Interpretable Features in Language Models cites this paper.

Sparse Autoencoders Find Highly Interpretable Features in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:44:02.219327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T06:42:23.274826Z digest=sha256:37aa197f9e910093135a6060bdefb28455aae6ef2877594d957a285900a529c0

Observation 06c1682f-2943-4fdf-9c88-a1c241b5c5a9 · inbound

Massive Activations in Large Language Models cites this paper.

Massive Activations in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.926612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:6f2e0aead184a4c562a57eb4a932deca58b8611dae693aa5781d618b35b2ccbc

Observation 5eb24e17-4314-437f-9002-6296ee8a9797 · inbound

How to use and interpret activation patching cites this paper.

How to use and interpret activation patching Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:34:06.947508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:34:06.888683Z digest=sha256:07dbadd5f87bccbb969410bc08bb491410108863c3b1d26362399f42181fba3a

Observation 73623492-eea4-4e80-a147-caa65323a6c5 · inbound

Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 cites this paper.

Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:47:20.043114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T05:47:19.953111Z digest=sha256:0a705607eb5da27ccdb9abc362e46f7852c46bc09ee6ae334afa1841fe3eaae0

Observation 74c04229-41bc-4e49-a463-c0e69e89325a · inbound

METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring cites this paper.

METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:23.641247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:18:23.641247Z digest=sha256:d12d5395a8f2b18d74cfc554cc9ee691fa4a1f800905dd20cbc898388b4c0b5d

Observation 506f5d05-097b-47e5-8de2-c7c498d0d4b7 · inbound

Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words cites this paper.

Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:28:16.037905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:28:16.037905Z digest=sha256:739b5fc1f0cdf839a127109dc83f727074ab9d7679bca1f08a7fa1883b9f694c

Observation fb34adb7-acd4-4b6a-83e0-145a026b5f36 · inbound

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit cites this paper.

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.427015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.427015Z digest=sha256:fcc6a251acc621ad7c81de38aa392141359050fb6058b7026bc259a94ae1d668

Observation afab162a-b487-4298-9b9a-441122dd8fd5 · inbound

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition cites this paper.

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T14:55:02.270268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:55:02.270268Z digest=sha256:69f57f199df16b13a858d07a5f0f0f726bed2368d05b4a187c3b580fd55d013c

Observation 2ac73ebc-e5f0-4645-b3e4-e0a97897b17c · inbound

Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models cites this paper.

Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:43:14.745081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:43:14.745081Z digest=sha256:3b132da4133cde0ba9c6e985fc1751263ac4346004c1f4e118de282ab0a85484

Observation 0556a0d9-92db-4033-b964-1d228dfcbde1 · inbound

Structure Development in List-Sorting Transformers cites this paper.

Structure Development in List-Sorting Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T23:32:04.152449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:32:04.152449Z digest=sha256:7194f851b82dfa447590f9cd73e683c64ef38cbf43928379e802561be70e1249

Observation 6699a44f-329f-4556-90f4-bbebab7ace34 · inbound

Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability cites this paper.

Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-09T22:08:49.514280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:08:49.514280Z digest=sha256:ff3fa829fd6627fe5514da8e0173ecb2ece2aaa0af982e19739e338b2a579f7c

Observation dbe5dc59-dc06-4da0-aeb6-7b7df836aba7 · inbound

Discovering Chunks in Neural Embeddings for Interpretability cites this paper.

Discovering Chunks in Neural Embeddings for Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T14:29:20.143673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:29:20.143673Z digest=sha256:8ba475947ee963dd9438fe1d09b8d56637649b9b332b069ede06e30e4ffaa63e

Observation b5677b01-3877-4b22-961c-ac1a439d4c43 · inbound

Studying Cross-cluster Modularity in Neural Networks cites this paper.

Studying Cross-cluster Modularity in Neural Networks Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T12:08:22.889317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:08:22.889317Z digest=sha256:bff18e0ac17e4e6a568bdc7cfb6e38320f6d70868e9bddbfbccbf2f9d428d4ba

Observation 3faaf77c-0529-4876-97a7-607826861f51 · inbound

Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging cites this paper.

Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T23:51:29.437526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:51:29.437526Z digest=sha256:5a806965d9529339a78b95b03c438b1b4dffa5589bc037a44b6d2812e1dd6c62

Observation 2ab72c83-bf32-4747-89cd-887135d947a0 · inbound

Sparse Autoencoders Do Not Find Canonical Units of Analysis cites this paper.

Sparse Autoencoders Do Not Find Canonical Units of Analysis Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:13.878104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:13.878104Z digest=sha256:b9d4c6392dbaa3403ef5640008732b85a95d5ea8247b38fb969af8a7e03fc2d9

Observation 9d1b64ae-978c-4a90-b109-83f02abc0433 · inbound

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation cites this paper.

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:55.076336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:15:55.076336Z digest=sha256:a797321038885f04ddc9f8c1140f6b1ef508c77e494755023b90eed4a7cf42c0

Observation c7766895-4d45-4b23-ab7f-ee6e2c54a899 · inbound

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management cites this paper.

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:44.813113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:48:44.813113Z digest=sha256:f7ccfa32aca18dff4ed7dbd23f8400b490f8e717c85d6842f02eb6c33f0db4a6

Observation b43ceb3a-bc7d-49f0-8ed7-1d01645c071e · inbound

On Mechanistic Circuits for Extractive Question-Answering cites this paper.

On Mechanistic Circuits for Extractive Question-Answering Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T11:01:28.934946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:01:28.934946Z digest=sha256:f9aa044b89091746a7e2c627248f8035b58a7505f46b1267e94e9f3f1a05bc29

Observation edb7054f-6d58-4845-83de-4659040a41f8 · inbound

Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning cites this paper.

Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T22:56:40.300816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:56:40.300816Z digest=sha256:fbcfe7ec55d7d4a7f992735df528f8ab7b174a2047ac70131e4716e2f50549ce

Observation 55c3bffa-c6e7-440d-88ef-8f86b156cbb9 · inbound

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis cites this paper.

Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:19.823946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:19.823946Z digest=sha256:e9eea28c01c2ce3a8384b41ffbcf2a4e1a0a9b9124ad8c710506f0f5b50e101d

Observation b2fd7b59-8fd2-4e9c-9f18-1af68ddb3d58 · inbound

Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models cites this paper.

Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:01.163498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:01.163498Z digest=sha256:69649159f981e91b1678b367fc4198f5657c128fbd418f96e10a80623406eaa5

Observation 498a66e8-0e85-4739-925d-87adffcf6165 · inbound

Void in Language Models cites this paper.

Void in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:03.554660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:03.554660Z digest=sha256:4531ab990a0f0bcb138681c363b0fc42a40304bfd0b78c6c7420a91d2db82601

Observation 3e2c7f5a-3e4c-447f-84e2-551669f66c45 · inbound

Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence cites this paper.

Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:00:58.165439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:00:58.165439Z digest=sha256:ddd2559507d81a4fc63bce4cd15f4fe79e92753748176c03b96b990fcfd94f5c

Observation 726d0a80-e229-4ab5-8293-e2c5e2121d26 · inbound

How Syntax Specialization Emerges in Language Models cites this paper.

How Syntax Specialization Emerges in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:02.941819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:18:02.941819Z digest=sha256:93a58343414a6c221c7ccb3592960cb204b315d96e8c9779dbc33048fd960825

Observation 10673e84-3e3d-497b-91a0-9cf3f210ee49 · inbound

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities cites this paper.

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:23.280878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:43:23.280878Z digest=sha256:f313f6d93527134733823c5f1db91b72ad699c08e375f1f7566844ffb24c45e0

Observation a44dc603-4a85-406b-9742-f1b0b0604356 · inbound

Expert Survey: AI Reliability & Security Research Priorities cites this paper.

Expert Survey: AI Reliability & Security Research Priorities Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:31:01.134755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:31:01.134755Z digest=sha256:23f3d0862afe4bc52a6452fdf89f53851caf9af8f1595b30698ccaca78aa29a7

Observation e34ae475-2f13-495c-9b01-b3495503fece · inbound

Mamba Knockout for Unraveling Factual Information Flow cites this paper.

Mamba Knockout for Unraveling Factual Information Flow Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:19.671001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:19.671001Z digest=sha256:513ff82f3f622d536983502c4f4b20a407d533dba739d9545a9f4b1165f5515a

Observation 25085deb-1076-4066-a0b5-1f3681d0afe1 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.352932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.352932Z digest=sha256:4a5782ececfd2ae0344df5e1471abdc51543282e516bbf15ee1e6eabc62b2188

Observation 795ad8e5-0f9c-470e-b00f-ec2eff7fd9fe · inbound

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective cites this paper.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.987123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.987123Z digest=sha256:26d4433d05eb09ce1c3b6c875651f40e5162784b6a3c4e59173b133ebad8cb44

Observation f2b18c6f-035b-4e61-bfb7-531a1aa1829f · inbound

InverseScope: Scalable Activation Inversion for Interpreting Large Language Models cites this paper.

InverseScope: Scalable Activation Inversion for Interpreting Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:34.868186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:34.868186Z digest=sha256:af2711ded4ebe04d4cc4d14d4e0f57e4d9fef05e287dac63a3268a9980ca5e17

Observation f79164a6-a64d-4590-95be-a8493a301f24 · inbound

Extrapolation by Association: Length Generalization Transfer in Transformers cites this paper.

Extrapolation by Association: Length Generalization Transfer in Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:23.693310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:23.693310Z digest=sha256:72398b598104c80e57dc9bf0ffbf1ac35e02d9ba489804dd7d9ef466d65cdb70

Observation 39793d58-ccfa-490b-b150-96a62e62d4c3 · inbound

Stochastic Parameter Decomposition cites this paper.

Stochastic Parameter Decomposition Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:39.358150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:48:39.358150Z digest=sha256:a49d6f37430e4f800abd9e6c72bfc059fa80bcfabd8f6a8f65d65661e710bd44

Observation 85216b34-3275-4f8e-8904-190d27a6f832 · inbound

How Do Vision-Language Models Process Conflicting Information Across Modalities? cites this paper.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:07.394735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:07.394735Z digest=sha256:a6f9facd785bf65f7259bc353eba94bc863fdd1da555e84c5988e0b0fde22394

Observation a1e61622-ebda-4274-882d-000351d125d4 · inbound

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability cites this paper.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.798498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.798498Z digest=sha256:566bd2fdff6564f20120f496f27e9a5df6da1613b409c5007888af3db5f5c9a4

Observation 92c1d813-3288-4c50-858a-a1b32c741045 · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:59.242409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:59.242409Z digest=sha256:939e028d0f100d2ad7ff992af0f138af48656cf459ef908bc2d4645bdd5daf89

Observation 7b4444bd-35fc-4680-8718-f0de87b4c9c1 · inbound

A Survey on Latent Reasoning cites this paper.

A Survey on Latent Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:31.787447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:31.787447Z digest=sha256:c39ca8716e2f57ee952098c2d10ffcae6ae56d8538069e35d3a344cc955f7293

Observation 82a0b61e-b35f-4d95-a5da-5a758d3c941d · inbound

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers cites this paper.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.725380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.725380Z digest=sha256:37bbb7a15955bf29c42eeaa835d0ddb057dcb4e23fe069fbce1f6203ba7ace63

Observation 3a97b7dc-8cdf-41a1-9c7d-070759ea9c63 · inbound

Algorithm Development in Neural Networks: Insights from the Streaming Parity Task cites this paper.

Algorithm Development in Neural Networks: Insights from the Streaming Parity Task Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-06T17:51:44.752240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:51:44.752240Z digest=sha256:98d141b2ecf6443076ecf9b24656e5bd5654c9b1f1a1421782e2309d97f4ea6b

Observation 0d1d4c40-723c-48b0-92a5-ca7436e0cd66 · inbound

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them cites this paper.

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.676363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:01.676363Z digest=sha256:fc8688918922bac089007d613c2a9cd5d10efd1ff2771b4e3e97ba8aff7c988f

Observation 2201a38b-4fcd-4c41-8296-90544649bed1 · inbound

Scaling laws for activation steering with Llama 2 models and refusal mechanisms cites this paper.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.509549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.509549Z digest=sha256:44c3a9ce5a2e22d57fbaffb65a604a971b5de1e231a4af9296bc65f075f4cf31

Observation 51295c71-28af-4e5a-b7f4-85e9fa1a7d44 · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:16.793689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:16.793689Z digest=sha256:40940cc38c6312637a0366b9861f6bed31b54badf7568970b6ca37333a529b4c

Observation f76758b2-cf9a-4c8c-8444-5c3b49d110d9 · inbound

Insights into a radiology-specialised multimodal large language model with sparse autoencoders cites this paper.

Insights into a radiology-specialised multimodal large language model with sparse autoencoders Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:44.385051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:40:44.385051Z digest=sha256:cbe2c70ab0cb58282574b66ca318db84da898fa0b03ddd905419c0de620a94a4

Observation 64e739de-8800-4a24-926d-73ab73659a3f · inbound

On the transferability of Sparse Autoencoders for interpreting compressed models cites this paper.

On the transferability of Sparse Autoencoders for interpreting compressed models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:45.806155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:24:45.806155Z digest=sha256:baec05df8ca90160f16dc6af9a8caa87968e342fad5c4c9e4acc854307c0a3a7

Observation 500d88f3-9155-479b-b9c1-1b0dec2ce635 · inbound

Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations cites this paper.

Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:30:19.195128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:30:19.195128Z digest=sha256:95bbf218c79e9ec15eed2b961bc63f05417f6c42028619ace3c8ed5e47173204

Observation 8c86245e-2728-4fdc-affe-b2b977a3905b · inbound

NEAT: Concept driven Neuron Attribution in LLMs cites this paper.

NEAT: Concept driven Neuron Attribution in LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T18:01:19.714746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:01:19.714746Z digest=sha256:41a13f500e5d188522c0c57f65afe53fa5b28c32db2c85e04f386ef9d967f093

Observation cd91d2a0-b672-47c7-bdc3-51f66fabe751 · inbound

From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits cites this paper.

From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:37:53.695806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:37:53.695806Z digest=sha256:8c69f9e7399df467dfd247f096c66e96351ff355eee779828de18ec0fcb7786b

Observation 019fb5de-08b8-4526-ab5f-2d97bed4325b · inbound

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation cites this paper.

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T17:16:33.076002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:16:33.076002Z digest=sha256:e44816b4bc46f9bec9c01f9d5a93aab24fe85f73fa8cda73c62b034c5aca746c

Observation dac7148a-b2ec-45d9-a02f-1e1b7f3be1e3 · inbound

All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens cites this paper.

All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T18:50:56.711463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:50:56.711463Z digest=sha256:660ca0d1af86fa6a9852ed0f5d024eab0e226bc52757dcc6d03d63dd2637ec52

Observation 78f72428-eba0-4732-a9e6-1ca4723e0d93 · inbound

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal cites this paper.

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:56:45.794536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T18:56:13.680353Z digest=sha256:fa4190e40485b399d5d2207e60790bfc616e4ea3754fdb8127dbfef4622cd726

Observation 4b27728d-2133-4603-9263-342399895f24 · inbound

Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content cites this paper.

Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T16:40:31.230569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:40:31.230569Z digest=sha256:1bc8b2f7d33a10631d44a5267f9f22e408ec32286f1f127371875fc319ca6d09

Observation b19b3df4-f598-416f-ac83-e9462e2f911c · inbound

From Features to Actions: Explainability in Traditional and Agentic AI Systems cites this paper.

From Features to Actions: Explainability in Traditional and Agentic AI Systems Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T03:50:17.552737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:50:17.552737Z digest=sha256:0c8b4c4d95b704258010eb9a4c40cf527874c99421ba98a21db793d6110b20ee

Observation 26f2a1fc-a6e3-4916-ab4e-87dbfafc3d56 · inbound

Prototype Transformer: Towards Language Model Architectures Interpretable by Design cites this paper.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.942638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.942638Z digest=sha256:ecfadec19f0c00c9b25290516d48697c7add3f867008019a8cb10dfc7e153e6d

Observation 2e26656e-1c79-49ac-b8b8-7fb793272cb3 · inbound

Transformers converge to invariant algorithmic cores cites this paper.

Transformers converge to invariant algorithmic cores Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T20:46:32.835730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:46:32.835730Z digest=sha256:3fc7252aa8878aa62dc5f44b1df367c0da52da7fb7b6c5c3d3dcfe026cc70129

Observation bb22fcdc-9130-478d-995c-dfdbbe5d6e62 · inbound

Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs cites this paper.

Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T16:59:55.496930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:59:55.496930Z digest=sha256:fa0a9e1c0526e0d3552276d99a3ac8cb721ea46954af71684d16ffa99438888a

Observation 44187884-afe5-412c-a825-b27170415412 · inbound

PhiNet: Speaker Verification with Phonetic Interpretability cites this paper.

PhiNet: Speaker Verification with Phonetic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:13:16.407720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T21:12:27.804372Z digest=sha256:668844c019528fb0a873de3cbcad4ac92177f6ce5c2c3292bac1852142868482

Observation 5b7272f9-483c-48c7-88b3-15cefa33a3f9 · inbound

Speaking of Language: Reflections on Metalanguage Research in NLP cites this paper.

Speaking of Language: Reflections on Metalanguage Research in NLP Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:28:14.115088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:23:55.117913Z digest=sha256:3bb047a57562ba246c43ad03ed9c96f6d54c67fd9afc0c16d71b3b11cc894c86

Observation 1ca5196b-4a7d-425b-9f27-2468f91b1fea · inbound

CURE:Circuit-Aware Unlearning for LLM-based Recommendation cites this paper.

CURE:Circuit-Aware Unlearning for LLM-based Recommendation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T16:47:25.256063Z digest=sha256:29d2654628004aa50e9d5abc945318205abf03e220fff9495759a8cf1edcf114

Observation 10d8aaa4-b06a-482a-bec8-22e97b4d6c68 · inbound

A Numerical PDEs Approach to Evolution Equations in Shape Analysis Based on Regularized Morphoelasticity cites this paper.

A Numerical PDEs Approach to Evolution Equations in Shape Analysis Based on Regularized Morphoelasticity Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T12:04:13.336341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:04:13.336341Z digest=sha256:0f3800f1ccd7888f4c9a16e9401f8ec8ca9651170c1051be85c52b8e975195e4

Observation 8a2fadcb-f376-4d46-8e0b-9b5e98569108 · inbound

Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings cites this paper.

Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:28:49.855280Z digest=sha256:9c6ac759a4c108bb7fe25dba9175ef03a882b9f35c646d8ca6e162004e77afcc

Observation 24aed8c1-dc66-43b6-acd5-c4ed346fcbca · inbound

Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs cites this paper.

Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T15:23:53.906870Z digest=sha256:ec837eceb76c85042747747f36b41dcde4746c94b0d2cf8d3c3730437571c2a3

Observation 0c9a623c-fb56-483c-b557-82967e02c085 · inbound

The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference cites this paper.

The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:23:37.593409Z digest=sha256:aa9cf8b77b2b7b3ce998c7c3eed09110fb9c10399716a818ba777976a981daad

Observation 9da0b9f1-91f3-469b-8bc6-b7a1b140cd55 · inbound

Grokking of Diffusion Models: Case Study on Modular Addition cites this paper.

Grokking of Diffusion Models: Case Study on Modular Addition Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:28:08.221886Z digest=sha256:fbb2dce27ca7eaf718f2717f8ab1d06b862a557daf373818cf098a0299e73dbf

Observation 95491ea6-98c9-4d5e-b15a-dcaccddc61d3 · inbound

Cell-Based Representation of Relational Binding in Language Models cites this paper.

Cell-Based Representation of Relational Binding in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T02:21:04.591556Z digest=sha256:acc37b3732423384ba63a0ce3ba10803833cee75fceb83019bfd019da15cabcd

Observation f4f6c0fb-a348-48aa-aebc-3816b7489f5a · inbound

Graph Memory Transformer (GMT) cites this paper.

Graph Memory Transformer (GMT) Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:27:32.790728Z digest=sha256:7382c53984f1f40d361c8b96cb899267c270c77be1cea8c1d31141917141b47a

Observation 0ade4e87-c831-40a6-b1c2-565b4b6acb68 · inbound

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs cites this paper.

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T08:34:14.310656Z digest=sha256:59af110fc0547f79fbcd50f11098016b3a50aaec1b5426bc4a202cf3eac1088e

Observation 295211e9-5eef-4df4-9d7e-4673d45abf88 · inbound

Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B cites this paper.

Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T00:29:17.274052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T00:26:22.118037Z digest=sha256:03929fa250600f1a4608baa8ee916f3d729f6b6e6cc4e2c9077606142ac1405f

Observation f1e482b1-5b50-4bb2-a4d2-320534f6f623 · inbound

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models cites this paper.

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T18:51:59.402906Z digest=sha256:4fdb9a1b1635dbf5d8302107f904390768f8e0078fab67659c6b8daeb3c1b35f

Observation e32a04dc-ddcd-4fd8-a4c8-97ea88e54705 · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T15:35:08.202464Z digest=sha256:929eba106a4b0bc96eade9cabf49ebcd4ec98a353686276e0cdd76699701b579

Observation c6febe8e-9fbb-410c-96b4-25db81a0247d · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:15:45.970687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T23:25:55.375541Z digest=sha256:fc69012be51948a98bb41fd1c20a496d0a2ee25b4a1a0bbcb31297bf470f9c4d

Observation 8d9d24ba-9a3f-4b3e-99f6-c11c8e7502dc · inbound

Negative Before Positive: Asymmetric Valence Processing in Large Language Models cites this paper.

Negative Before Positive: Asymmetric Valence Processing in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T11:13:15.576248Z digest=sha256:3634edafca2e832b74e01f90a583ad597a568bb34b3bc1cde26f653ca9159c26

Observation 60d326e6-7a38-4a6a-87a5-ab761f42f2bf · inbound

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List cites this paper.

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:46:46.241768Z digest=sha256:f9e77093129de98d0f2dc71ae65b6573d08e7987c0c3e02355eabb1d64684075

Observation 9bac8d96-335f-4482-96c3-bb725c4417a6 · inbound

Hallucination Detection via Activations of Open-Weight Proxy Analyzers cites this paper.

Hallucination Detection via Activations of Open-Weight Proxy Analyzers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:43:07.105363Z digest=sha256:72e2872ee90b267cada820a0875c2bbb0187637125b56fa95dc6bf6d60228d27

Observation 624c863e-9194-4106-b8a0-391d003b3199 · inbound

Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions cites this paper.

Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T03:14:17.072349Z digest=sha256:fcf31002633aa8f64b1f0ac8b93ae6404e329a255d85f4a3ca79f4f9ea6683ed

Observation 06313e4a-423a-41a2-b486-a89aa6c50ec0 · inbound

Tool Calling is Linearly Readable and Steerable in Language Models cites this paper.

Tool Calling is Linearly Readable and Steerable in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T03:09:11.013914Z digest=sha256:3d23731326b37c8a97c7ab0887ea0fd281ad119b048a27ec753ffda0f100d684

Observation b8a59202-11fd-4fe6-8ee4-3e64790d516c · inbound

Architecture, Not Scale: Circuit Localization in Large Language Models cites this paper.

Architecture, Not Scale: Circuit Localization in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:19:23.901130Z digest=sha256:d8dad73fd488c6b5d1131de2202310c8b2d41c2b225db02a2ee52ad764532b75

Observation 70994fb7-94ea-4d5d-96c8-85c819f99643 · inbound

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations cites this paper.

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 38

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:55:32.734498Z digest=sha256:9f86eec549c1506793dd1ee26199a3ea77aa674c2e2ddcb49debd21ed880f0b4

Observation 2e2b1232-1b81-4600-84a8-3288a850731b · inbound

Dissecting Jet-Tagger Through Mechanistic Interpretability cites this paper.

Dissecting Jet-Tagger Through Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:09.296991Z digest=sha256:83ec480437a44663dc3692827632b4534c1bbf854cd10e32f43482b13f004433

Observation e30a75e0-8f51-4920-8f12-ef35c4eeb491 · inbound

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining cites this paper.

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:12:20.784347Z digest=sha256:1aee870373de0e33a9c461dc80252c55bb2a8210c8fa55d55beb02fa370e23c5

Observation 6701c411-84e5-4ab4-b5f3-a6b8f4729de6 · inbound

Not How Many, But Which: Parameter Placement in Low-Rank Adaptation cites this paper.

Not How Many, But Which: Parameter Placement in Low-Rank Adaptation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:13:55.497799Z digest=sha256:4751861dd8dbb5259356bfa193753b2657bde80f8fd69ec2ca066a46c5ef0df2

Observation ba5cb654-c967-4782-a1b5-f4f2a2ad089c · inbound

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender cites this paper.

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:13:51.588363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T04:50:43.547037Z digest=sha256:a01a469fa169154a3ddc8031290ffc356d60dbdc9c90da1b6f057ecfe35ecd69

Observation 144bb3e7-cbed-4722-99e3-b7f307725fff · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:17:55.505913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:b1296d80d06c21ded7bbca3a4aeacdb00dc9d8e4ead702071e3b1a8166df9038

Observation df8da93c-ed81-44b9-bbf3-ef552741e46c · inbound

How to Interpret Agent Behavior cites this paper.

How to Interpret Agent Behavior Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:27:35.915812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:23:25.269217Z digest=sha256:68efa2e95cb575eb0f81f91109ee60827c43c54d88bc4f2975c83605660d1368

Observation a254d540-c15f-4cb6-bff0-884c2fcae2d8 · inbound

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands cites this paper.

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T20:55:03.906712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T20:53:04.274840Z digest=sha256:0636823eeeea85ad5420207861d470e0a73683d47ac95cb60578a79653501b31

Observation 7a215416-a70c-4cab-a9be-a0aead2162ef · inbound

Beyond Linear Superposition: Discovering Climate Features in AI Weather Models with KAN-SAE cites this paper.

Beyond Linear Superposition: Discovering Climate Features in AI Weather Models with KAN-SAE Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:48:23.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T14:48:03.631853Z digest=sha256:1ba5a3c6ec0e1db7d1fa8d86ab266871e55786770133a20bdce174fa87d46509

Observation fad22ae8-4744-45ba-a445-697e11bd43c6 · inbound

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination cites this paper.

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:28:05.509681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T06:25:43.369455Z digest=sha256:f388abddda5602b8ffb457cc519082e606edb5caf30ea3d500a33f2dbfb13eb6

Observation d0b69103-8ec1-4b6f-a902-5c481152ec00 · inbound

Manifold-Guided Attention Steering cites this paper.

Manifold-Guided Attention Steering Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:14:45.333190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:14:12.029680Z digest=sha256:7ea74ce2b2a3f1bd8e136f92bd0052940c92aff654356b22362e2d83df96ebd3

Observation 4ecd92ed-d6ef-4ffd-8726-637478f30ca8 · inbound

From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models cites this paper.

From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:01:11.963153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T06:56:18.332754Z digest=sha256:6ba58dbd1a9c76e3cc1efa92c587b2a5dbb146b4ef336d25b174f53593e72ea3

Observation c25b2f2a-f7fb-4b4f-a9eb-49ae70c531fa · inbound

Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability cites this paper.

Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:04:38.219967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T11:58:53.868902Z digest=sha256:e6c652e6c65cc0d6f57756c18d5dbd17b3c9ffe3f97a7f202a1c8da3b47466b3

Observation 318a0fc9-99c2-44c0-a366-d59510d6a87f · inbound

Binding Visual Features Point by Point cites this paper.

Binding Visual Features Point by Point Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:44:03.323191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:56:44.793896Z digest=sha256:0765be5631b866817fff844a4b280a9e38d6378d9ccb504dab2cc98658782258

Observation fabec0ee-ae7b-4e8e-a546-cf076f665328 · inbound

Tracing Computation Density in LLMs cites this paper.

Tracing Computation Density in LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:23:50.638460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T18:20:33.014687Z digest=sha256:d315a40a7d5fd811f78d482bb847adb85656f5ffde767ca2946903ed424bf80b

Observation 6f718af9-9b38-4c18-a23d-36789c7c6ed3 · inbound

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection cites this paper.

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:13:16.594466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:09:12.429033Z digest=sha256:02fc67c71a833a710fc720f655934d76eb43ad48e5f82c422c578f83b271ffd5

Observation 9d4c1b38-3e18-4c41-9c1e-be69f6bb65e6 · inbound

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning cites this paper.

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:16:24.594953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T14:29:34.220222Z digest=sha256:cc9b3868a561c8f4cb4454e06dec598b34a8a6fb0eb1580ba2353114781f1756

Observation 63e274be-e5d8-4e06-b022-a065b5ec8a76 · inbound

A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting cites this paper.

A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:36:29.800981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T09:52:58.807254Z digest=sha256:2384cdc6b6b229dbc1310a9bf293ff8c54c05348c930a3e9535f0e97aa3656fc

Observation 2c07424f-6a92-4b99-a595-ceb5bbfca97d · inbound

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs cites this paper.

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.475224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:08:36.451147Z digest=sha256:16e6f002aa841cb4617c6c1fa0c8a39ae1dbbf0689138c4fa58d2d7573852729

Observation 0bb5532a-d2b0-472f-aaa0-dfc50bd77957 · inbound

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations cites this paper.

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:41.392147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T07:39:16.719604Z digest=sha256:3c509c33b98a94b37cbf70492217ba647e8ba8acab842716304456d83bcffe42

Observation 367da854-a7ea-46fe-a2d0-b520df8d094b · inbound

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models cites this paper.

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:06:44.456142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T07:07:58.170765Z digest=sha256:948fe33bf4587b1e6c32cc92e5db6bcbb6a326de42f15e40bd2bbd42bbfa4ccf

Observation 0390f0d2-67fb-423e-ac7e-3417df484cc7 · inbound

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability cites this paper.

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-28T02:11:29.080954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:07:18.198225Z digest=sha256:5afd4661b5410a522ce38f4d87b4f38668dbfa74564552066ac495e21202de43

Observation fe30d291-3b8f-4a69-8570-afdda8478029 · inbound

Sparsely gated tiny linear experts cites this paper.

Sparsely gated tiny linear experts Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:17:09.306780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:49:49.299925Z digest=sha256:749a73800593419ea16e5a0d5f91a5418d65014bd1302301f3c9ac3aba3f5f03