Pith. sign in

Paper Citation Record · LEDGER

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2507.09406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09406 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:00:51.737979Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 12d5f8d8-3aa0-4fc5-80ad-1915e43d841c · outbound

This paper cites Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.639435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.639435Z digest=sha256:ae17fdc2f109e6d5d236040294de809766293c51e0f718f9801ee7c4818ae080

Observation 7ac2f7d0-f8ea-46b0-9f18-77c51008f7e5 · outbound

This paper cites Advances in Tabulating Carmichael Numbers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Advances in Tabulating Carmichael Numbers

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:52.053060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:00:51.660510Z digest=sha256:c5a15607c24a204510c2a8c5091521e91fd764c2b3e62f743d3daa45739c54a1

Observation 4e4a4539-e5c8-4b12-b1a2-387df5881be0 · outbound

This paper cites Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.667312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.667312Z digest=sha256:4b0e23cac753d153fc8b691bd85d25c937f5ebadff56e4cb26331b42adb1b346

Observation dd598a0e-193b-40c0-8e9c-e236ded13ec8 · outbound

This paper cites Generative Adversarial Transformers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Generative Adversarial Transformers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:00:52.006017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:00:51.674334Z digest=sha256:408290d97d49d8e1d3c8349289b5021f4fa2a1976a94fbeba4245142dbec3445

Observation fccdbb04-1984-4470-9093-8fde84001115 · outbound

This paper cites Wav-KAN: Wavelet Kolmogorov-Arnold Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Wav-KAN: Wavelet Kolmogorov-Arnold Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.691306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.691306Z digest=sha256:2d3b27d1d5e6f3d851bc119fbbe29b7bd1b84b02916291f40295c9e030256bfa

Observation b03f4675-9e8a-470a-af2a-b3a9d1a67ef9 · outbound

This paper cites Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.937745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:00:51.699652Z digest=sha256:5a04f09828129cdd06e33098d6807bb9413e255c1ab1c9135c8b9476007f986e

Observation d94b862f-9d0a-46a8-8e48-25a4526a7381 · outbound

This paper cites When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.707342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.707342Z digest=sha256:8956c8ca886715fddf51b184773b36f4f8af3847fd35d09719ae10dd490f631d

Observation 76678cc1-2453-42e8-9002-4947eb8d2b69 · outbound

This paper cites Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.868458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:00:51.720004Z digest=sha256:3a06b5e2ef9d63dd0d7a6408f34bc2a7c70d933ec25c628e9893965ca639ac2c

Observation 44b6f56d-e98c-43a7-8ad2-cc33bdc55e65 · outbound

This paper cites Guidelines to Develop Trustworthy Conversational Agents for Children.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Guidelines to Develop Trustworthy Conversational Agents for Children

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.798306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:00:51.737979Z digest=sha256:c0ea57fd5c2a5f2721b829b5fde7b88c41047b674ca74586d5d1495f8d738ebe

Observation 82a0b61e-b35f-4d95-a5da-5a758d3c941d · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.725380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.725380Z digest=sha256:ff9d1c3ef1c52150144705a9c098dbd1a7428a79ab3f9f4e9dc07b6d554864a2

Observation 77779af3-9269-45de-8d9f-813008f96dae · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.683998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.683998Z digest=sha256:669e86d12e6c514cc2c707e764bbfe4014f88bb83d0d40dee35805bd71f08972

Observation db60480d-bbfd-433e-994d-e15860c42690 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.731106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.731106Z digest=sha256:c2c35d0ff83f8a4ab0c44b9ec9f228d88775677ed9a2ef05948f963c3f76ba6e

Observation 524799c9-cfad-4af9-8e34-35d4c6851b49 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Mechanistic Interpretability for AI Safety -- A Review

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.646078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.646078Z digest=sha256:c5d848c8755630fe28b151f56957e70661c6f58c56903dfd020e04e7764f6458

Observation cff18f5d-3676-41d1-9e20-d22d8d0d5311 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.654377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.654377Z digest=sha256:4f136fd599ab52f2e22ddd72a903e9bfce9d914e06b8e9a3acc4db865b61ab41

Observation b23b4ea6-c812-4b64-8cbd-3fec2f95ea8b · outbound

This paper cites Attribution Patching Outperforms Automated Circuit Discovery.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Attribution Patching Outperforms Automated Circuit Discovery

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.713937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.713937Z digest=sha256:857a99cf8802c8f08b5e74bac1160575b10b2376177b0cfe807f06515ca02532

Pith citing papers

Observation 98504e4f-8bf1-491c-9c42-e344c02384b3 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.785630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:eb560924c335f64cf3067c626a8fdd81089ac0b6e319cffd571fea4019322607

Observation dd2e7047-a80d-4aca-bb02-c463207269ec · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.976229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:330d9ce0cd6bc677a151a2c4b24e295869bae0fd126802242e5ee7e786ebb4aa