Pith. sign in

Paper Citation Record · LEDGER

Detecting Safety Training Modification in Language Models via Activation Analysis

As of 10 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.05578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05578 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:13:51.695862Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1145ef7f-047a-4ef1-8ea8-a333b10b3283 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp.

Detecting Safety Training Modification in Language Models via Activation Analysis Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.220640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.637323Z digest=sha256:adb69b9ac201e645886c6bfde78342b7a765e32c9c18f4ea5398cc53a4da6f33

Observation 85033a13-9694-4ca9-b036-b4c508dc651b · outbound

This paper cites Dolphin: An uncensored, unbiased language model.

Detecting Safety Training Modification in Language Models via Activation Analysis Dolphin: An uncensored, unbiased language model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.209167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.641526Z digest=sha256:053e224ca17fe46442ac9ca2e898fe8c3652819762d06477e4c6544e21754a78

Observation a8ccaf23-60e3-4135-81da-5273afaaff09 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Detecting Safety Training Modification in Language Models via Activation Analysis Representation Engineering: A Top-Down Approach to AI Transparency

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.645097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.645097Z digest=sha256:60b2cef5e6cb8ec6d1ecd090fc76e5ac9844bcf758457094969628ec71936df1

Observation ed92bf25-8632-4f6f-acfa-18a076b6f4ce · outbound

This paper cites Steering Language Models With Activation Engineering.

Detecting Safety Training Modification in Language Models via Activation Analysis Steering Language Models With Activation Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.649072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.649072Z digest=sha256:233db8aa280db3665e0074204504d07b0b252ecad9f2dc23a6340a072d3fa315

Observation d430b2db-35fb-49e0-89b2-2b967986b074 · outbound

This paper cites Instructional Fingerprinting of Large Language Models.

Detecting Safety Training Modification in Language Models via Activation Analysis Instructional Fingerprinting of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.652802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.652802Z digest=sha256:85a3c52f17722fbda533f0ccb05146800a5e0bfd40184d87d04c5df2514cbeee

Observation d7dba1c4-a689-4636-a7e1-22455d71b97e · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Detecting Safety Training Modification in Language Models via Activation Analysis HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.656374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.656374Z digest=sha256:8227805691791b6f9007309da35f5ec8da0bc32df42beb5d421905bae7c64316

Observation 912bbfdd-f564-45b6-b6f1-af2451286584 · outbound

This paper cites ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Detecting Safety Training Modification in Language Models via Activation Analysis ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.198859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.660030Z digest=sha256:061faa6c84524e091e679ea49f19d2011f799427d742e44bdbe626ee535436f2

Observation 3f998b65-7b57-40f4-8e23-58d45a7e8509 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Detecting Safety Training Modification in Language Models via Activation Analysis WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.663609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.663609Z digest=sha256:7fe52e1d1627021694394f3dbbec91c86ebdf35731fed2681d1ba8a8c2013b2a

Observation 5d4c3a31-a753-4cdb-8bf0-9a98a4617f0f · outbound

This paper cites Sokhansanj.

Detecting Safety Training Modification in Language Models via Activation Analysis Sokhansanj

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.188435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.667184Z digest=sha256:ef48ece485308ea825e4f68c78464e99ead25856c8968efe7457cec45e911c37

Observation 36a9fd31-55df-475f-a3f3-8fa8785b160b · outbound

This paper cites XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025.

Detecting Safety Training Modification in Language Models via Activation Analysis XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.670840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.670840Z digest=sha256:5f7a1ecc257920245ad05855762c94dba0e23f3b3e8a5084f5fbd71054389fec

Observation 9c0e3c76-2e48-4f09-8a88-634e2dbaacd8 · outbound

This paper cites Defending Large Language Models Against Attacks With Residual Stream Activation Analysis.

Detecting Safety Training Modification in Language Models via Activation Analysis Defending Large Language Models Against Attacks With Residual Stream Activation Analysis

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:13:51.949039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.673991Z digest=sha256:e1d008366f1f835017d4db13714c99cf8563175129807a2315b8108ad07842db

Observation dd234149-1813-4a52-9f04-abdb222f3955 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Detecting Safety Training Modification in Language Models via Activation Analysis Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.677511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.677511Z digest=sha256:5f2877af4a6f4707d84f7363bd19ab05ca08258580b82c4ada4bf2af57faaa36

Observation 8b6de393-72fc-4fe3-8aea-7a496d7ddcde · outbound

This paper cites Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022.

Detecting Safety Training Modification in Language Models via Activation Analysis Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.178320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.680835Z digest=sha256:b7d917c9ea7f908755e3ffe87430d2c5b9ad54b4ace768f1c6931fab37afc460

Observation a315e742-ee52-403d-b15e-bc24df6b0682 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Detecting Safety Training Modification in Language Models via Activation Analysis Gonzalez, Hao Zhang, and Ion Stoica

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.683805Z digest=sha256:5f61a663f206a6c056e4d704665db06fb069a24a762b6f7474fe334de1b54632

Observation 9fbef769-486b-4882-acad-c3121e68a00b · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Detecting Safety Training Modification in Language Models via Activation Analysis The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.686798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.686798Z digest=sha256:6922150db2bbcd35ea3acf5330f5fe025bee4477ab405b3b32ae2725647bb047

Observation 706e3b47-05cf-4bee-b850-6f0c68768f80 · outbound

This paper cites Lo- cating and Editing Factual Associations in GPT.

Detecting Safety Training Modification in Language Models via Activation Analysis Lo- cating and Editing Factual Associations in GPT

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.159556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.690142Z digest=sha256:adcccd26bdb0b83f1fb422bdf32c588dea63179faff4bab48f51015dc0dfb161

Observation 87cae2ed-8127-4b6c-a6c9-d1effc757772 · outbound

This paper cites ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation.

Detecting Safety Training Modification in Language Models via Activation Analysis ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:13:51.917259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.693027Z digest=sha256:a07dadffa8c8781e2df6bdd9ce8dc5d09273839676876dbe6f51a08a0e2aa197

Observation 11f04dd4-9a9c-42ba-86b3-792f5256a686 · outbound

This paper cites VLLM_ALLOW_INSECURE_SERIALIZATION.

Detecting Safety Training Modification in Language Models via Activation Analysis VLLM_ALLOW_INSECURE_SERIALIZATION

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.148643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.695862Z digest=sha256:639720a36067fffa5bc73f39e334a35a3260a1f0712114b44c31d9cba5112acc

Pith citing papers

No inbound Pith citation observations are available.