Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:17:52.535540Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2506.17052.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:17:52.535540Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T14:41:21.582578Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:48:32.559321Z
81 of 81 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 64ab2eca-44f6-48a9-99fd-20e5b87d9c7c · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Interpretable machine learning–a brief history, state-of-the-art and challenges
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b643256d-d1e9-4627-926b-87ffc578a1fd · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Explainable ai: A review of machine learning interpretability methods
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ff3d312b-6801-4a5e-bd20-8d9fd9fc6294 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e2b43f0-461b-4252-a661-6d930a520d88 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Understanding Neural Networks Through Deep Visualization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 642212fa-24cd-4ad3-b00d-3bff67f46833 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Visualizing deep neural network decisions: Prediction difference analysis
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b9ac762-14d7-4e77-8bac-2a686f14ad36 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Grad-cam: Visual explanations from deep networks via gradient-based localization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b3044b-6645-4246-9b13-bdd5ea2172f9 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Axiomatic attribution for deep networks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 237e3feb-ef2f-46de-8d7d-2f62a0885404 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Attention is all you need
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9761cfd0-fe32-434d-a5ea-4fd9d92ba08f · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Rethinking Interpretability in the Era of Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b4bd07-1b98-45f0-ab26-e81f09eee511 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Transformer feed-forward layers are key-value memories
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c796c4d-e5e5-4d7e-a76c-5e6daa386b1a · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5a469460-6d1a-4e4e-b51d-f9904a857e38 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers The emergence of number and syntax units in lstm language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 221d03ec-4c1d-4460-8928-c38e345c0c20 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Locating and editing factual associations in gpt
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d3f0020-5787-4ebc-a371-fd24cf007610 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Mass-Editing Memory in a Transformer
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe043fa-b81c-43bd-bba2-69f7f47568d6 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Lima: Less is more for alignment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2691784c-8244-46e8-9204-be5d940657d0 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers HarmBench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 819be478-7c94-44f6-b8e0-d3bf80a4262a · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Training Verifiers to Solve Math Word Problems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c62cfb-af2f-4037-8945-8fe333d55559 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers A mathematical framework for transformer circuits
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6cce585-25e4-44ca-ae2b-ff260f21d0e8 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Efficient Estimation of Word Representations in Vector Space
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd49423f-e7be-4217-8a54-32e790c4e333 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Daniel Freeman, Theodore R
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 83b3e45c-ae52-4120-a566-991f529ade7a · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9626cbd-ac7f-4a70-b9e8-ea500237e2fb · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Linear Representations of Sentiment in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ac3ff6e-50cb-46cf-a1a5-3aba4bea503c · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fbd40d7-8645-4ec1-95d8-3521f64a2617 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Steering Llama 2 via Contrastive Activation Addition
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2124474-298c-46db-94ea-6adb117f5cf8 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Improving Activation Steering in Language Models with Mean-Centring
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d569c0c0-f6bf-4ada-aa56-97d22086cf4e · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Refusal in Language Models Is Mediated by a Single Direction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58f1d1e-39a8-4eb5-bd90-16e91e99310b · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers interpreting gpt: the logit lens
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6919faa0-7e13-41b7-8ac2-92568ab038c5 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Investigating gender bias in language models using causal mediation analysis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bcdc1773-2f24-431e-b9e6-994e4b714c29 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Localizing Model Behavior with Path Patching
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7f17ba-d12a-4f79-8e33-91b390a63f50 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Inducing causal structure for interpretable neural networks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 909a3951-0689-420c-8119-5d57319de618 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Transformerlens
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation db1964c7-2d63-417f-85b7-abbfe89e212b · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Vit prisma: A mechanistic interpretability library for vision transformers
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d3fe72d7-9fc3-49cc-a5bd-fa86c1c47c03 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers How do Large Language Models Handle Multilingualism?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99bd0b90-0298-4efb-97cc-60a315ee15be · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Do Llamas Work in English? On the Latent Language of Multilingual Transformers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb8239f-9043-4c1c-b6d0-7581405ae700 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d257a39-dcee-4505-9392-c535afc26dcf · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Scaling and evaluating sparse autoencoders
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5e9480-8f7b-4658-933d-fb99fa067da3 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Gemma: Open Models Based on Gemini Research and Technology
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2f4150-01fe-4982-b398-40de9ffdb133 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Towards monosemanticity: Decomposing language models with dictionary learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b299e770-df4c-4212-aed7-614042f39d1b · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Chain-of-thought prompting elicits reasoning in large language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73ebeb61-65e3-4b43-a40b-2d146a86ddb3 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers A comprehensive survey on test-time adaptation under distribution shifts
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f473369-48c2-4778-8403-99a9eac5c26f · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers CommonsenseQA: A question answering challenge targeting commonsense knowledge
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 687ac887-d8c1-4734-9783-fabfd09df58a · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Evaluating Large Language Models Trained on Code
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e0b8fc-f7a5-488e-9a38-a31038f39cdc · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Program Synthesis with Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a1ab93-b49f-47ca-ba9d-fa00e0a9023e · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8075964b-0708-4cb9-b11e-d80c6d4c5c93 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b7d8c2e-f06d-4c25-bc2d-ddc1ad5348cf · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers A framework for few-shot language model evaluation, 07 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 32333543-dfa8-417c-8d82-0f8846f9c097 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Interpretability dreams
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9064121e-3617-45a1-adf7-3cd50690821e · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce5b0fc8-b6da-4bbe-b3ed-d1147dba9c51 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Qwen Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c96340-ff62-45bd-8b02-11fe58fe6bf3 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a798688-a531-43f5-8999-672f930d95ff · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f00f00d-7a0a-4be8-81eb-a06372aa70ed · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Vision transformers need registers
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cefe7a3d-2aab-41ed-bbbe-6e8af1ef90a2 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Visualizing and understanding convolutional networks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2137c90c-b7f3-4931-998d-6da6021f36ea · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Network dissec- tion: Quantifying interpretability of deep visual representations
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d88e2d76-bb37-4b71-a6f0-b9b7248221be · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Interpreting deep visual representa- tions via network dissection
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aa756b2a-7760-441c-a14a-5d5c9283c4ee · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Interpretable basis decomposition for visual explanation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4fa9617f-09e8-45bb-aa96-e094ef2ced7e · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Understanding the role of individual units in a deep neural network.Proceedings of the National Academy of Sciences, 117(48):30071–30078, 2020
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e2db74ee-c989-4540-b24d-8c16022b83bf · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Dissecting recall of factual associations in auto-regressive language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a0d43275-b1e8-40e1-b4b5-f1e6e0a668fe · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Does localization inform editing? surprising differences in causality-based localization vs
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2e6a373-50b8-476c-8736-978ca4eba91b · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Editing common sense in transformers
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3e7da4eb-9637-4224-b636-9d2aac054a2d · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Massive editing for large language models via meta learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b549e9ad-8aed-47db-874e-adf4fadb6513 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Pmet: Precise model editing in a transformer
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f29ad2ba-d8dd-4f5e-ad0b-33c035e61128 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d1b6891a-83c8-4bc3-b3b1-41f22a11dcb6 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16d825b-4161-4458-b8d0-0e7b859f6f3b · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Sparse Interventions in Language Models with Differentiable Masking
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b192fa-5f25-4c66-a537-0bd45eb0313a · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Knowledge Neurons in Pretrained Transformers
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6a5fd1-ef38-48c3-a8c3-a80f97a42aef · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Finding and editing multi- modal neurons in pre-trained transformers
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f6a365e1-e286-470a-9e25-1a23d00a9f89 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Towards neuron attributions in multi-modal large language models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94bf6f2b-6176-4f0c-922b-f484219d0d79 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Retrieval Head Mechanistically Explains Long-Context Factuality
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aed7111-a7da-4ae8-ba57-ed04aa60d13c · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Successor Heads: Recurring, Interpretable Attention Heads In The Wild
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b586e7ea-7c1e-4318-93cc-4662c38e4816 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7300ee97-b5f2-4e4a-8dd3-616d91825639 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Zoom in: An introduction to circuits
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d80426b-0b0d-41e9-8963-a3827e1d2793 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Mechanistic Interpretability for AI Safety -- A Review
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b99e52d2-27e3-4523-b6f1-919dd739f733 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Open Problems in Mechanistic Interpretability
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be5256e-4ce6-475a-88c2-b11c428be253 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers In-context learning and induction heads
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8c13d93a-1e39-41d0-a3ea-ae34912b38ab · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e5519bd1-d3fd-4315-a5cb-0501924c2b6d · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Progress measures for grokking via mechanistic interpretability
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ca93db5b-6680-463e-9a74-f5884919918a · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Iteration Head: A Mechanistic Study of Chain-of-Thought
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5636ce6-db96-401c-8c66-602474f4cdd5 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers the Golden Gate Bridge
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 702bfc56-2627-4152-8ce0-b9ec394778f0 · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Tabby cat
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8c24b08e-4f0d-4a80-89b8-f589c85e364d · outbound
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a16380cb-f3e9-4345-9f56-68ffd849ae43 · inbound
Towards Robustness against Typographic Attack with Training-free Concept Localization From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.