Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:00:51.737979Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2507.09406.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:00:51.737979Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
15 of 15 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 12d5f8d8-3aa0-4fc5-80ad-1915e43d841c · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac2f7d0-f8ea-46b0-9f18-77c51008f7e5 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Advances in Tabulating Carmichael Numbers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e4a4539-e5c8-4b12-b1a2-387df5881be0 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd598a0e-193b-40c0-8e9c-e236ded13ec8 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Generative Adversarial Transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fccdbb04-1984-4470-9093-8fde84001115 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Wav-KAN: Wavelet Kolmogorov-Arnold Networks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03f4675-9e8a-470a-af2a-b3a9d1a67ef9 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d94b862f-9d0a-46a8-8e48-25a4526a7381 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76678cc1-2453-42e8-9002-4947eb8d2b69 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44b6f56d-e98c-43a7-8ad2-cc33bdc55e65 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Guidelines to Develop Trustworthy Conversational Agents for Children
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82a0b61e-b35f-4d95-a5da-5a758d3c941d · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77779af3-9269-45de-8d9f-813008f96dae · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db60480d-bbfd-433e-994d-e15860c42690 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524799c9-cfad-4af9-8e34-35d4c6851b49 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Mechanistic Interpretability for AI Safety -- A Review
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff18f5d-3676-41d1-9e20-d22d8d0d5311 · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23b4ea6-c812-4b64-8cbd-3fec2f95ea8b · outbound
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Attribution Patching Outperforms Automated Circuit Discovery
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98504e4f-8bf1-491c-9c42-e344c02384b3 · inbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
Reference 257
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd2e7047-a80d-4aca-bb02-c463207269ec · inbound
Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.