Pith. sign in

Paper Citation Record · LEDGER

Detecting Safety Training Modification in Language Models via Activation Analysis

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.05578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05578 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:13:51.695862Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1145ef7f-047a-4ef1-8ea8-a333b10b3283 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp.

Detecting Safety Training Modification in Language Models via Activation Analysis Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.220640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.637323Z digest=sha256:732006b31a0554da3d3361d1fd4579dc3b26ac51e7a013124801666415969bf3

Observation 85033a13-9694-4ca9-b036-b4c508dc651b · outbound

This paper cites Dolphin: An uncensored, unbiased language model.

Detecting Safety Training Modification in Language Models via Activation Analysis Dolphin: An uncensored, unbiased language model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.209167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.641526Z digest=sha256:0fb69ed70ed620f27439618ae3008b99dc0190a56d21946b67fa11e0710b1a9f

Observation a8ccaf23-60e3-4135-81da-5273afaaff09 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Detecting Safety Training Modification in Language Models via Activation Analysis Representation Engineering: A Top-Down Approach to AI Transparency

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.645097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.645097Z digest=sha256:74779ac47910c680f8914587a294feea5b4c262503550138502ca26799066490

Observation ed92bf25-8632-4f6f-acfa-18a076b6f4ce · outbound

This paper cites Steering Language Models With Activation Engineering.

Detecting Safety Training Modification in Language Models via Activation Analysis Steering Language Models With Activation Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.649072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.649072Z digest=sha256:98d43bfc68cd65f45efffe15c01c561a0799427f735fbeae5b71b51a035e5ba0

Observation d430b2db-35fb-49e0-89b2-2b967986b074 · outbound

This paper cites Instructional Fingerprinting of Large Language Models.

Detecting Safety Training Modification in Language Models via Activation Analysis Instructional Fingerprinting of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.652802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.652802Z digest=sha256:dc29fd6fa2018d51d42e140048a458cb8746b0b224d85535ceb231028d2f7ede

Observation d7dba1c4-a689-4636-a7e1-22455d71b97e · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Detecting Safety Training Modification in Language Models via Activation Analysis HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.656374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.656374Z digest=sha256:4dba4d10069cb8c0d40b6ccd9941d6e36588982804600aaf442bf34c9421e403

Observation 912bbfdd-f564-45b6-b6f1-af2451286584 · outbound

This paper cites ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Detecting Safety Training Modification in Language Models via Activation Analysis ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.198859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.660030Z digest=sha256:2f7d2b68fcd0557fa463349d7d4a3f0f2610b15cad6837009f82ba57dcbbdbb5

Observation 3f998b65-7b57-40f4-8e23-58d45a7e8509 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Detecting Safety Training Modification in Language Models via Activation Analysis WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.663609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.663609Z digest=sha256:33556f0a8d05a327c7d393c86c692e0d3cad1f27c302dbf3ede1143867b040df

Observation 5d4c3a31-a753-4cdb-8bf0-9a98a4617f0f · outbound

This paper cites Sokhansanj.

Detecting Safety Training Modification in Language Models via Activation Analysis Sokhansanj

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.188435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.667184Z digest=sha256:fc5d99265a65d9a4cf2303abf996dea879173ba4977c51fef26dbdd5cff77871

Observation 36a9fd31-55df-475f-a3f3-8fa8785b160b · outbound

This paper cites XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025.

Detecting Safety Training Modification in Language Models via Activation Analysis XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.670840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.670840Z digest=sha256:aa39047d20258a797b3203cef631cb98dafe5cee6160fed530dbe249f5499f4a

Observation 9c0e3c76-2e48-4f09-8a88-634e2dbaacd8 · outbound

This paper cites Defending Large Language Models Against Attacks With Residual Stream Activation Analysis.

Detecting Safety Training Modification in Language Models via Activation Analysis Defending Large Language Models Against Attacks With Residual Stream Activation Analysis

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:13:51.949039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.673991Z digest=sha256:5e9c64b14017119ca0eb6947063c9532b1158f33e5261a3d46330f43fc08c765

Observation dd234149-1813-4a52-9f04-abdb222f3955 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Detecting Safety Training Modification in Language Models via Activation Analysis Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.677511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.677511Z digest=sha256:c707b1cb694f9624332aa14dded7b3e0bd594a278e3d7222e75ae05bb4b94f5f

Observation 8b6de393-72fc-4fe3-8aea-7a496d7ddcde · outbound

This paper cites Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022.

Detecting Safety Training Modification in Language Models via Activation Analysis Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.178320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.680835Z digest=sha256:22c83e678acf17e2d96befe5b062a1c0ee02cbe8c95603ef9316057ece02f972

Observation a315e742-ee52-403d-b15e-bc24df6b0682 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Detecting Safety Training Modification in Language Models via Activation Analysis Gonzalez, Hao Zhang, and Ion Stoica

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.683805Z digest=sha256:367c50770c5c1a3393bb6951e125b0943a7b497a42ef1d198c69e13fde50fa67

Observation 9fbef769-486b-4882-acad-c3121e68a00b · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Detecting Safety Training Modification in Language Models via Activation Analysis The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.686798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.686798Z digest=sha256:3b349d9adac7f2ef7493d5d14a17b909640b9ce24c105dbd7201289bd360b3a4

Observation 706e3b47-05cf-4bee-b850-6f0c68768f80 · outbound

This paper cites Lo- cating and Editing Factual Associations in GPT.

Detecting Safety Training Modification in Language Models via Activation Analysis Lo- cating and Editing Factual Associations in GPT

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.159556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.690142Z digest=sha256:f1e0fb1c9c78d750d3dcc529ec6d3f56de4c65a46c9ecb86880480fac85c5bc0

Observation 87cae2ed-8127-4b6c-a6c9-d1effc757772 · outbound

This paper cites ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation.

Detecting Safety Training Modification in Language Models via Activation Analysis ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:13:51.917259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.693027Z digest=sha256:2a240a1f63cac7c27707f849075929d2905c183cc8cf919e11ff581bfcb4c0c1

Observation 11f04dd4-9a9c-42ba-86b3-792f5256a686 · outbound

This paper cites VLLM_ALLOW_INSECURE_SERIALIZATION.

Detecting Safety Training Modification in Language Models via Activation Analysis VLLM_ALLOW_INSECURE_SERIALIZATION

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.148643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.695862Z digest=sha256:dbaa82e6dcb7893ab999a20bd9c0d0001e3d3e01d6f51b74e63e4466576c7dbd

Pith citing papers

No inbound Pith citation observations are available.