Pith. sign in

Paper Citation Record · LEDGER

Detecting Safety Training Modification in Language Models via Activation Analysis

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.05578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05578 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:13:51.695862Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1145ef7f-047a-4ef1-8ea8-a333b10b3283 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp.

Detecting Safety Training Modification in Language Models via Activation Analysis Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.220640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.637323Z digest=sha256:89242b881f6d4da3e26f7f5c2d759ac1a932257e865658c28bdcbdce6735450e

Observation 85033a13-9694-4ca9-b036-b4c508dc651b · outbound

This paper cites Dolphin: An uncensored, unbiased language model.

Detecting Safety Training Modification in Language Models via Activation Analysis Dolphin: An uncensored, unbiased language model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.209167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.641526Z digest=sha256:b6134483235fe0319c1f0197f1ed12e431f23f62253727943954d73952cf4a01

Observation a8ccaf23-60e3-4135-81da-5273afaaff09 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Detecting Safety Training Modification in Language Models via Activation Analysis Representation Engineering: A Top-Down Approach to AI Transparency

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.645097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.645097Z digest=sha256:a819cb019c80517b22f3e239426d6f3eb4cf9d7ed8b12024a28b0e1c0ecec304

Observation ed92bf25-8632-4f6f-acfa-18a076b6f4ce · outbound

This paper cites Steering Language Models With Activation Engineering.

Detecting Safety Training Modification in Language Models via Activation Analysis Steering Language Models With Activation Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.649072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.649072Z digest=sha256:b87954f1b966f5674909ad11338b220585644a348f198b4e6296e04835b276c1

Observation d430b2db-35fb-49e0-89b2-2b967986b074 · outbound

This paper cites Instructional Fingerprinting of Large Language Models.

Detecting Safety Training Modification in Language Models via Activation Analysis Instructional Fingerprinting of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.652802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.652802Z digest=sha256:f860386a2fe756d30093f79371855203ef17b4c62ff03849d5555461551e54c2

Observation d7dba1c4-a689-4636-a7e1-22455d71b97e · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Detecting Safety Training Modification in Language Models via Activation Analysis HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.656374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.656374Z digest=sha256:c918250f04b28bb4152cd11c633d368fae5db63f9bb0f110b48475c818f0bd1d

Observation 912bbfdd-f564-45b6-b6f1-af2451286584 · outbound

This paper cites ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Detecting Safety Training Modification in Language Models via Activation Analysis ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.198859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.660030Z digest=sha256:133f76c72b17241506dc604a71f4726dc70a96a3d36be17d919b43c237acfac7

Observation 3f998b65-7b57-40f4-8e23-58d45a7e8509 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Detecting Safety Training Modification in Language Models via Activation Analysis WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.663609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.663609Z digest=sha256:51eb7dbdff465855102187b21db9747166062e6e147fd7d6e67d78e18a11cb1d

Observation 5d4c3a31-a753-4cdb-8bf0-9a98a4617f0f · outbound

This paper cites Sokhansanj.

Detecting Safety Training Modification in Language Models via Activation Analysis Sokhansanj

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.188435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.667184Z digest=sha256:6013d88665c49e9f4e2230b19fb1a6dd6dd8a65a154954f5287f6735cf32e0bc

Observation 36a9fd31-55df-475f-a3f3-8fa8785b160b · outbound

This paper cites XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025.

Detecting Safety Training Modification in Language Models via Activation Analysis XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.670840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.670840Z digest=sha256:29a63dfa189ea7779bdfa03d43f957999d7bad923ca2a8730e6cde4b8a1ee212

Observation 9c0e3c76-2e48-4f09-8a88-634e2dbaacd8 · outbound

This paper cites Defending Large Language Models Against Attacks With Residual Stream Activation Analysis.

Detecting Safety Training Modification in Language Models via Activation Analysis Defending Large Language Models Against Attacks With Residual Stream Activation Analysis

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:13:51.949039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.673991Z digest=sha256:f740e32d7880d417f0a0f9f87cb253dd5ce2ca8cafbdcc90ae9a63bb3ed4a2c7

Observation dd234149-1813-4a52-9f04-abdb222f3955 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Detecting Safety Training Modification in Language Models via Activation Analysis Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.677511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.677511Z digest=sha256:673b9464187e934a21e7f9d809a2e989ce9512a86e4998526980508af7e4a558

Observation 8b6de393-72fc-4fe3-8aea-7a496d7ddcde · outbound

This paper cites Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022.

Detecting Safety Training Modification in Language Models via Activation Analysis Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.178320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.680835Z digest=sha256:7ef6fc3a1f1c2e668470593b4f62073717b212c706469f98a3dce8a94b3b1963

Observation a315e742-ee52-403d-b15e-bc24df6b0682 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Detecting Safety Training Modification in Language Models via Activation Analysis Gonzalez, Hao Zhang, and Ion Stoica

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.683805Z digest=sha256:24efb2b9956b9a315b3f89bab43d68c433ead82bea37fc9accca01214c1c7cf2

Observation 9fbef769-486b-4882-acad-c3121e68a00b · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Detecting Safety Training Modification in Language Models via Activation Analysis The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.686798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.686798Z digest=sha256:6b343933103f1d19e2da6e6a865b34c85e5bad009a973def3eedb906e5e0270b

Observation 706e3b47-05cf-4bee-b850-6f0c68768f80 · outbound

This paper cites Lo- cating and Editing Factual Associations in GPT.

Detecting Safety Training Modification in Language Models via Activation Analysis Lo- cating and Editing Factual Associations in GPT

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.159556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.690142Z digest=sha256:58472f125e877cd15f950a8a6ad7de4cdd01384cb35939e00ce72117a5699926

Observation 87cae2ed-8127-4b6c-a6c9-d1effc757772 · outbound

This paper cites ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation.

Detecting Safety Training Modification in Language Models via Activation Analysis ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:13:51.917259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.693027Z digest=sha256:492ebff3418cc69c731f739e3c61e339f8aa044b24e83128b00f65603e546aca

Observation 11f04dd4-9a9c-42ba-86b3-792f5256a686 · outbound

This paper cites VLLM_ALLOW_INSECURE_SERIALIZATION.

Detecting Safety Training Modification in Language Models via Activation Analysis VLLM_ALLOW_INSECURE_SERIALIZATION

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.148643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.695862Z digest=sha256:98629d2ad58585ef7ec7aa5aab22119edc94950371d670ecd2935796bff1fc1c

Pith citing papers

No inbound Pith citation observations are available.