Pith. sign in

Paper Citation Record · LEDGER

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2507.09406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09406 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:00:51.737979Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 12d5f8d8-3aa0-4fc5-80ad-1915e43d841c · outbound

This paper cites Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.639435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.639435Z digest=sha256:a40f27e3e2edf97ed08b1aeb99c884d05553c1bfd003e45bbeb9a9ca7bf995e2

Observation 7ac2f7d0-f8ea-46b0-9f18-77c51008f7e5 · outbound

This paper cites Advances in Tabulating Carmichael Numbers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Advances in Tabulating Carmichael Numbers

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:52.053060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:00:51.660510Z digest=sha256:ccb1ca10ce03de0989c56f5741c3fe3ac5bb362bedde1a642f4191878cac536c

Observation 4e4a4539-e5c8-4b12-b1a2-387df5881be0 · outbound

This paper cites Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.667312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.667312Z digest=sha256:3d591192b96b8ad9adc981793222cfc25ed37bf767d258045004e4d92c227652

Observation dd598a0e-193b-40c0-8e9c-e236ded13ec8 · outbound

This paper cites Generative Adversarial Transformers.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Generative Adversarial Transformers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:00:52.006017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:00:51.674334Z digest=sha256:9d44bc9b3b88f78a7f4d9dd0b4686258f5a4b7a1a91795c66db26cb2e82f6aca

Observation fccdbb04-1984-4470-9093-8fde84001115 · outbound

This paper cites Wav-KAN: Wavelet Kolmogorov-Arnold Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Wav-KAN: Wavelet Kolmogorov-Arnold Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.691306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.691306Z digest=sha256:e74ef9b41344b1141b105cc911ac1094113f3dd73033be8f477755e048db59da

Observation b03f4675-9e8a-470a-af2a-b3a9d1a67ef9 · outbound

This paper cites Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Workload Assessment of Human-Machine Interface: A Simulator Study with Psychophysiological Measures

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.937745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:00:51.699652Z digest=sha256:8e84fb26d494b68c3645ffc02295282e9c6d0bd85fe3e92a7acef40770d8bd11

Observation d94b862f-9d0a-46a8-8e48-25a4526a7381 · outbound

This paper cites When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.707342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.707342Z digest=sha256:41bd72f78faf04d3d8ca13f1b68d254948b171cf61f67cd4f9b51a16ba2f0bc2

Observation 76678cc1-2453-42e8-9002-4947eb8d2b69 · outbound

This paper cites Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Episodic Memory Theory for the Mechanistic Interpretation of Recurrent Neural Networks

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.868458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:00:51.720004Z digest=sha256:21b942ab5433f2c211f3f449d1fcbf4c5241b8945e8b2914336333b9d3751e1d

Observation 44b6f56d-e98c-43a7-8ad2-cc33bdc55e65 · outbound

This paper cites Guidelines to Develop Trustworthy Conversational Agents for Children.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Guidelines to Develop Trustworthy Conversational Agents for Children

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:00:51.798306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:00:51.737979Z digest=sha256:90a73740b0ae3169f4aea814cea9984252bbba89b96de50e8742ea239bbe0ba1

Observation 82a0b61e-b35f-4d95-a5da-5a758d3c941d · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.725380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.725380Z digest=sha256:999ffb8e80f286ba9ccae05110931e8094d327fbe236e410586865bfc9393883

Observation 77779af3-9269-45de-8d9f-813008f96dae · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.683998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.683998Z digest=sha256:d32d2063eaf622b46862b139081cbcfd5bb1244d46b463486988f08ec15c3a22

Observation db60480d-bbfd-433e-994d-e15860c42690 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.731106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.731106Z digest=sha256:b8993f9f9336bd494b9fc9e47f1ad8be7a18cae9c0b324ee867538ecfd6f7740

Observation 524799c9-cfad-4af9-8e34-35d4c6851b49 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Mechanistic Interpretability for AI Safety -- A Review

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.646078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.646078Z digest=sha256:f92be5393dc22c9c133c172647ece2f8c3d46c62ee28842c4a3355375c853e9e

Observation cff18f5d-3676-41d1-9e20-d22d8d0d5311 · outbound

This paper cites The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A".

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.654377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.654377Z digest=sha256:a20465f0092ca99219a664b864db9e163821969bc861211882b309230f9898bd

Observation b23b4ea6-c812-4b64-8cbd-3fec2f95ea8b · outbound

This paper cites Attribution Patching Outperforms Automated Circuit Discovery.

Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers Attribution Patching Outperforms Automated Circuit Discovery

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:51.713937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:51.713937Z digest=sha256:10d77af0eb5115788aef42be6a86c7a6fa4aeeb1e72aa979d4e237820502f6e5

Pith citing papers

Observation 98504e4f-8bf1-491c-9c42-e344c02384b3 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.785630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:612d4a2ff5f59218c9112f0b817a3b40c4dbeafa54428bfefb341179b2304d62

Observation dd2e7047-a80d-4aca-bb02-c463207269ec · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.976229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:db9cffe19892be56214313310565f2de30e8f462538d105341a58c9faf9d278a