Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T17:13:51.408311Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 100 inbound Pith citation observations for arXiv:2211.00593.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T17:13:51.408311Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:18:23.641247Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
66 of 66 outbound references displayed
External citation measurements
50
pith, observed 2026-08-05T02:28:24.338817Z
Observation 98497d3c-b10e-453e-bf8d-596debd99853 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Language models are few-shot learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f08b166d-a709-4013-bfc6-bfdd73254e67 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A literature survey of recent advances in chatbots
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37bfacbe-9204-4eb2-bea1-7d8fd4956da3 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A mathematical framework for transformer circuits
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 279f0de2-acd0-49f9-a6b9-22cbb44214f2 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Causal abstractions of neural networks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9f534c03-5115-4764-b4f2-8f1ad9c98b99 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small X-Risk Analysis for AI Research
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c858597-b301-4dc6-906d-460d218ead17 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Natural language descriptions of deep visual features
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a30b5c23-be6f-47a9-a6e5-036ea59af1d5 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Are sixteen heads really better than one? Advances in neural information processing systems, 32
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e7b7134a-9388-4c94-94a5-4c1dc9906e5f · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Compositional explanations of neurons
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd0c2e2d-b8be-4964-881f-6ac73b41eef2 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A mechanistic interpretability analysis of grokking
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 403d4df1-c521-44a9-b835-c04f501e2754 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Mechanistic interpretability, variables, and the importance of interpretable bases
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f8d5f64c-cf8f-4a71-8f2a-54eba66d1f20 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Distill , year =
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f436f922-3967-4911-80db-da477de89f2e · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Language models are unsupervised multitask learners
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aca4400f-afb2-45ac-a7bb-c249b2619fb9 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a96bd0e0-3982-45cb-93ba-5e824f313142 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Attention is all you need
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3169fad0-860b-48ac-b990-01b8e1e848e4 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Investigating gender bias in language models using causal mediation analysis
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0cd3606-387a-4966-b84e-6cd3416991a9 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Emergent Abilities of Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e3bd3c9-e00a-434a-a5a9-f123127dc550 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Shifting machine learning for healthcare from development to deployment and from models to data
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17f4ee92-3301-4a0c-8146-1327c7c89752 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d3a1646-5ffe-4ddb-a530-9b526f550ba3 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small BERT Rediscovers the Classical NLP Pipeline
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3d26bd71-74cf-4a5d-8c6c-369a06076975 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2019 , subtitle =
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bf0796bc-9d44-4fb9-8347-a1754cf14598 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Learning to Generate Reviews and Discovering Sentiment
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd5c7a01-5b11-47d7-8b76-2285a6854c9a · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Implicit Representations of Meaning in Neural Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a175f726-d97d-4094-8d9e-db3489192e8d · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small An Interpretability Illusion for BERT
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a387258-f02b-4a01-ae25-77cd1a6846ed · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Proceedings of the 57th Conference of the Association for Computational Linguistics
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 68a0559c-f58b-49cf-9553-2b7f506d0b5a · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small International Conference on Machine Learning , pages=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02091e11-3423-48e0-ae27-57225f28dfd2 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unveiling Transformers with LEGO: a synthetic reasoning task
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbc2d07d-597a-42d0-9b66-0e56c0f3a6b1 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f3fa369-c0d2-41a8-bb75-9482dfaf0adc · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small arXiv , year=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4280cf2-37d2-4e9a-88ec-6d2c4dd77aa7 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small In-context Learning and Induction Heads
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0820cd9-5070-49fe-8f27-45396e23d4e7 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Information , volume=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d100709-e2bc-449b-963f-4322db7bd8d1 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Nature Biomedical Engineering , pages=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63ae8442-2c60-4ab8-a8bd-741f909d8abd · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small ArXiv , year=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0f9dba28-98ec-4cc6-96b0-f2beed018285 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unsolved Problems in ML Safety
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83cfc6fa-cb3e-40e9-8621-b986875ff7cd · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small International Conference on Learning Representations , year=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44a715b7-5696-4efb-bead-36a3fdf5c9d2 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , volume=
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bd8dc7d-0387-4449-b96d-a2ae2a04cfad · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Implicit Representations of Meaning in Neural Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a1cb9589-f8aa-417a-9e46-03544b49338a · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 91650fe7-4249-4906-8bf1-82e44e07705b · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , volume=
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 602f23f7-0f08-449c-adc1-1ddda406da47 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small On the Pitfalls of Analyzing Individual Neurons in Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ee3c801-d141-4173-bcaa-57024fc6ff31 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88621bf4-43d0-4af1-a283-d2ae7a49a5bf · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small ArXiv , year=
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 047ba2cb-c305-4d86-96f3-3e6a46377144 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in neural information processing systems , volume=
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7072732-3db9-4ae0-b905-d61d79902e7f · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ffe00a29-866d-49e0-b687-5fcbbec3926f · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Analyzing Transformers in Embedding Space
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea483a46-26a0-4d1f-82c7-33a2bd8eaa54 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Transformer Feed-Forward Layers Are Key-Value Memories
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3ec52625-53b6-4a57-98cb-bfecf08c22ed · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Locating and Editing Factual Associations in GPT
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2eeb5839-af2f-4a12-b879-b8301d1766ea · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Distill , year =
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa4caf6d-a2c5-4beb-9462-bd4c3efeed6d · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Scaling Learning Algorithms Towards
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bb6b7e6-a66e-430a-84ca-5f5a148e293f · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small and Osindero, Simon and Teh, Yee Whye , journal =
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c4d53185-b7e1-4f5e-a6d2-c3c023af842a · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2016 , publisher=
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25fd0f4f-fec7-4df4-9f3a-f64fef9fc1cc · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 62c4ebf3-4e48-4c0a-9285-de5df982d558 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2021 , journal=
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0987cb3-bed0-4b35-99b5-fc5ddbc65541 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f4af84f0-fff2-4cdd-8fff-2ac8df776418 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , editor=
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b18e8659-7b93-4b86-9abe-7341cee8e46c · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Advances in Neural Information Processing Systems , description =
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0cc76915-2cdf-4107-bcab-6983442b219a · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Language Models are Few-Shot Learners , url =
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5991bd5f-0b1f-47cf-b598-1d91be5347a0 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Distill , year =
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0c57649-e011-4981-a64d-b48d65a89eb4 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c482a09a-28d9-46ee-ab98-a981aff2145a · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Sanity Checks for Saliency Maps , url =
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1d6857b2-b1d1-47bb-a8fd-528a0afa2dc3 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small GitHub repository , howpublished =
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 81bd157d-3446-4f4c-bed0-dc027aa9cdbc · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 15f9fe67-6c11-4ac6-ab47-fc57bdde959b · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c49db7c6-8185-4259-95d3-9fff7cd80529 · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small An overview of 11 proposals for building safe advanced AI
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06f26e36-fbd1-47b3-aeed-72ee147eb75d · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small A ttention is not E xplanation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 878d1a18-f906-423c-ad06-2020161a1e8d · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small Optimal Brain Damage , url =
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26bf98f5-93e2-45b6-8e3c-209314a33acc · outbound
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 2022 , journal=
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd07232b-7a32-41ac-ac35-9ab414b28ddb · inbound
Progress measures for grokking via mechanistic interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16867590-ebb7-473c-bb39-a4221f87c667 · inbound
Eliciting Latent Predictions from Transformers with the Tuned Lens Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b262013-a5bf-47ad-a29c-b8745b7cc891 · inbound
Sparse Autoencoders Find Highly Interpretable Features in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06c1682f-2943-4fdf-9c88-a1c241b5c5a9 · inbound
Massive Activations in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5eb24e17-4314-437f-9002-6296ee8a9797 · inbound
How to use and interpret activation patching Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73623492-eea4-4e80-a147-caa65323a6c5 · inbound
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 74c04229-41bc-4e49-a463-c0e69e89325a · inbound
METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506f5d05-097b-47e5-8de2-c7c498d0d4b7 · inbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb34adb7-acd4-4b6a-83e0-145a026b5f36 · inbound
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afab162a-b487-4298-9b9a-441122dd8fd5 · inbound
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac73ebc-e5f0-4645-b3e4-e0a97897b17c · inbound
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0556a0d9-92db-4033-b964-1d228dfcbde1 · inbound
Structure Development in List-Sorting Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6699a44f-329f-4556-90f4-bbebab7ace34 · inbound
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe5dc59-dc06-4da0-aeb6-7b7df836aba7 · inbound
Discovering Chunks in Neural Embeddings for Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5677b01-3877-4b22-961c-ac1a439d4c43 · inbound
Studying Cross-cluster Modularity in Neural Networks Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3faaf77c-0529-4876-97a7-607826861f51 · inbound
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab72c83-bf32-4747-89cd-887135d947a0 · inbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1b64ae-978c-4a90-b109-83f02abc0433 · inbound
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7766895-4d45-4b23-ab7f-ee6e2c54a899 · inbound
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43ceb3a-bc7d-49f0-8ed7-1d01645c071e · inbound
On Mechanistic Circuits for Extractive Question-Answering Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb7054f-6d58-4845-83de-4659040a41f8 · inbound
Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c3bffa-c6e7-440d-88ef-8f86b156cbb9 · inbound
Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2fd7b59-8fd2-4e9c-9f18-1af68ddb3d58 · inbound
Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498a66e8-0e85-4739-925d-87adffcf6165 · inbound
Void in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e2c7f5a-3e4c-447f-84e2-551669f66c45 · inbound
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726d0a80-e229-4ab5-8293-e2c5e2121d26 · inbound
How Syntax Specialization Emerges in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10673e84-3e3d-497b-91a0-9cf3f210ee49 · inbound
Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44dc603-4a85-406b-9742-f1b0b0604356 · inbound
Expert Survey: AI Reliability & Security Research Priorities Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e34ae475-2f13-495c-9b01-b3495503fece · inbound
Mamba Knockout for Unraveling Factual Information Flow Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25085deb-1076-4066-a0b5-1f3681d0afe1 · inbound
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 795ad8e5-0f9c-470e-b00f-ec2eff7fd9fe · inbound
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b18c6f-035b-4e61-bfb7-531a1aa1829f · inbound
InverseScope: Scalable Activation Inversion for Interpreting Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f79164a6-a64d-4590-95be-a8493a301f24 · inbound
Extrapolation by Association: Length Generalization Transfer in Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39793d58-ccfa-490b-b150-96a62e62d4c3 · inbound
Stochastic Parameter Decomposition Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85216b34-3275-4f8e-8904-190d27a6f832 · inbound
How Do Vision-Language Models Process Conflicting Information Across Modalities? Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e61622-ebda-4274-882d-000351d125d4 · inbound
Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c1d813-3288-4c50-858a-a1b32c741045 · inbound
Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4444bd-35fc-4680-8718-f0de87b4c9c1 · inbound
A Survey on Latent Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a0b61e-b35f-4d95-a5da-5a758d3c941d · inbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a97b7dc-8cdf-41a1-9c7d-070759ea9c63 · inbound
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1d4c40-723c-48b0-92a5-ca7436e0cd66 · inbound
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2201a38b-4fcd-4c41-8296-90544649bed1 · inbound
Scaling laws for activation steering with Llama 2 models and refusal mechanisms Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51295c71-28af-4e5a-b7f4-85e9fa1a7d44 · inbound
LLMs Encode Harmfulness and Refusal Separately Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76758b2-cf9a-4c8c-8444-5c3b49d110d9 · inbound
Insights into a radiology-specialised multimodal large language model with sparse autoencoders Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64e739de-8800-4a24-926d-73ab73659a3f · inbound
On the transferability of Sparse Autoencoders for interpreting compressed models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 500d88f3-9155-479b-b9c1-1b0dec2ce635 · inbound
Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c86245e-2728-4fdc-affe-b2b977a3905b · inbound
NEAT: Concept driven Neuron Attribution in LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd91d2a0-b672-47c7-bdc3-51f66fabe751 · inbound
From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019fb5de-08b8-4526-ab5f-2d97bed4325b · inbound
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac7148a-b2ec-45d9-a02f-1e1b7f3be1e3 · inbound
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f72428-eba0-4732-a9e6-1ca4723e0d93 · inbound
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b27728d-2133-4603-9263-342399895f24 · inbound
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19b3df4-f598-416f-ac83-e9462e2f911c · inbound
From Features to Actions: Explainability in Traditional and Agentic AI Systems Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f2a1fc-a6e3-4916-ab4e-87dbfafc3d56 · inbound
Prototype Transformer: Towards Language Model Architectures Interpretable by Design Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e26656e-1c79-49ac-b8b8-7fb793272cb3 · inbound
Transformers converge to invariant algorithmic cores Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb22fcdc-9130-478d-995c-dfdbbe5d6e62 · inbound
Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44187884-afe5-412c-a825-b27170415412 · inbound
PhiNet: Speaker Verification with Phonetic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5b7272f9-483c-48c7-88b3-15cefa33a3f9 · inbound
Speaking of Language: Reflections on Metalanguage Research in NLP Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ca5196b-4a7d-425b-9f27-2468f91b1fea · inbound
CURE:Circuit-Aware Unlearning for LLM-based Recommendation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10d8aaa4-b06a-482a-bec8-22e97b4d6c68 · inbound
A Numerical PDEs Approach to Evolution Equations in Shape Analysis Based on Regularized Morphoelasticity Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a2fadcb-f376-4d46-8e0b-9b5e98569108 · inbound
Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 24aed8c1-dc66-43b6-acd5-c4ed346fcbca · inbound
Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c9a623c-fb56-483c-b557-82967e02c085 · inbound
The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9da0b9f1-91f3-469b-8bc6-b7a1b140cd55 · inbound
Grokking of Diffusion Models: Case Study on Modular Addition Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95491ea6-98c9-4d5e-b15a-dcaccddc61d3 · inbound
Cell-Based Representation of Relational Binding in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f4f6c0fb-a348-48aa-aebc-3816b7489f5a · inbound
Graph Memory Transformer (GMT) Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ade4e87-c831-40a6-b1c2-565b4b6acb68 · inbound
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 295211e9-5eef-4df4-9d7e-4673d45abf88 · inbound
Borrowed Geometry: Cross-Distribution Head-Importance Fingerprints of Frozen Pretrained Gemma 4 31B Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f1e482b1-5b50-4bb2-a4d2-320534f6f623 · inbound
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e32a04dc-ddcd-4fd8-a4c8-97ea88e54705 · inbound
High-Dimensional Statistics: Reflections on Progress and Open Problems Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c6febe8e-9fbb-410c-96b4-25db81a0247d · inbound
High-Dimensional Statistics: Reflections on Progress and Open Problems Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d9d24ba-9a3f-4b3e-99f6-c11c8e7502dc · inbound
Negative Before Positive: Asymmetric Valence Processing in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60d326e6-7a38-4a6a-87a5-ab761f42f2bf · inbound
The Position Curse: LLMs Struggle to Locate the Last Few Items in a List Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9bac8d96-335f-4482-96c3-bb725c4417a6 · inbound
Hallucination Detection via Activations of Open-Weight Proxy Analyzers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 624c863e-9194-4106-b8a0-391d003b3199 · inbound
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06313e4a-423a-41a2-b486-a89aa6c50ec0 · inbound
Tool Calling is Linearly Readable and Steerable in Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8a59202-11fd-4fe6-8ee4-3e64790d516c · inbound
Architecture, Not Scale: Circuit Localization in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 70994fb7-94ea-4d5d-96c8-85c819f99643 · inbound
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e2b1232-1b81-4600-84a8-3288a850731b · inbound
Dissecting Jet-Tagger Through Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e30a75e0-8f51-4920-8f12-ef35c4eeb491 · inbound
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6701c411-84e5-4ab4-b5f3-a6b8f4729de6 · inbound
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ba5cb654-c967-4782-a1b5-f4f2a2ad089c · inbound
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 144bb3e7-cbed-4722-99e3-b7f307725fff · inbound
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df8da93c-ed81-44b9-bbf3-ef552741e46c · inbound
How to Interpret Agent Behavior Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a254d540-c15f-4cb6-bff0-884c2fcae2d8 · inbound
Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a215416-a70c-4cab-a9be-a0aead2162ef · inbound
Beyond Linear Superposition: Discovering Climate Features in AI Weather Models with KAN-SAE Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fad22ae8-4744-45ba-a445-697e11bd43c6 · inbound
Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0b69103-8ec1-4b6f-a902-5c481152ec00 · inbound
Manifold-Guided Attention Steering Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4ecd92ed-d6ef-4ffd-8726-637478f30ca8 · inbound
From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c25b2f2a-f7fb-4b4f-a9eb-49ae70c531fa · inbound
Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 318a0fc9-99c2-44c0-a366-d59510d6a87f · inbound
Binding Visual Features Point by Point Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fabec0ee-ae7b-4e8e-a546-cf076f665328 · inbound
Tracing Computation Density in LLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6f718af9-9b38-4c18-a23d-36789c7c6ed3 · inbound
Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d4c1b38-3e18-4c41-9c1e-be69f6bb65e6 · inbound
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63e274be-e5d8-4e06-b022-a065b5ec8a76 · inbound
A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c07424f-6a92-4b99-a595-ceb5bbfca97d · inbound
Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bb5532a-d2b0-472f-aaa0-dfc50bd77957 · inbound
STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 367da854-a7ea-46fe-a2d0-b520df8d094b · inbound
Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0390f0d2-67fb-423e-ac7e-3417df484cc7 · inbound
Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe30d291-3b8f-4a69-8570-afdda8478029 · inbound
Sparsely gated tiny linear experts Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.