Pith. sign in

Paper Citation Record · LEDGER

A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2401.01967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.01967 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:46:29.654391Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39a753dc-e3e7-4ba2-b3fb-147aff02c503 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 145

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:47:56.007880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:c586a490f813dac669bb890fc03be3bdb84948576cdbfcc26ae5591ccfe31645

Observation a1f7d08c-4030-4d6a-8592-d22be6238419 · inbound

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering cites this paper.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.654391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.654391Z digest=sha256:69b0af616cd030150a6c6434b75ca62fb659feaaf8b8b73f2e3a7af1c481e3ce

Observation 7646f38f-246d-4cbe-8366-4d53756e1f44 · inbound

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions cites this paper.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.292674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.292674Z digest=sha256:05e090e71d83797a4ae360fa79388ab8e1310cbbd025176a0878fdb5a4202fb2

Observation f36f4cbe-ae25-4a6e-bb70-b6b9ea3c6815 · inbound

Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs cites this paper.

Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:16.715517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:16.715517Z digest=sha256:1fd156e5ac68e46c8ed6fb9eaa65ac7f0a7e6db8deefecceaa154a045613c6c2

Observation c7d9fbcc-b2c3-4b1e-9051-d532b70b21de · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.589127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:ca5112eabfac97e396f10ee4da3971384b68cc5868a12cf2eeb82d22b1a580a7

Observation e8c1a24c-31d4-49f9-8c43-df6bbe522c10 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.735447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.735447Z digest=sha256:95762c08b86355b7ec8dde119bd269403060974218fef9d0b5db56c025d68f8b

Observation 49cceb30-af28-4b3a-b8f0-90f368234dec · inbound

NEAT: Concept driven Neuron Attribution in LLMs cites this paper.

NEAT: Concept driven Neuron Attribution in LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T18:01:19.652668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:01:19.652668Z digest=sha256:26b1f30d0b56907df48a1aa8a96a6dcabc9be466adaccd4b3befd6522e7344bd

Observation d9769c2f-bf46-42c4-b8b9-c147b6fdeca4 · inbound

Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight cites this paper.

Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:21:23.755068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T13:21:18.699465Z digest=sha256:574519abb29de76b69d5ca38e786a2247ebaee109688cbe918f3ede9cf3c7094

Observation 4dd04c29-8bed-4f24-b922-26441193a6e9 · inbound

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models cites this paper.

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:48.897724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:43:48.897724Z digest=sha256:20bff6ede17e198cc40e148267852636c8c7579d7d0b91829c87ecba1e5b0a9d

Observation 96001828-996e-494c-92ac-e730bc7f142d · inbound

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting cites this paper.

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:49:33.106792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:49:33.106792Z digest=sha256:3696e6774cb4f9cd4dbfa72bbb0bccf7eff60033c3c05b1a3f4ccf846b26dd88

Observation 3569a863-06fd-4737-bdcf-5c53a2b1b5c7 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:28.951146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:6457e13a530cd089ed528ad49425c2de776b3a8f01ad0cfd927dc987a0b15e2b

Observation 9f4cb83e-1551-45f3-b03c-ff7da9755cbe · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:04.435702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:9e3314b99ffb2657efbc9fa6e790ed7c74690b8ba3c07ad7e7c5d95c65db8a31

Observation 35c5450e-e424-4d72-96eb-05b67664760f · inbound

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification cites this paper.

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:05:21.995701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:04:13.987667Z digest=sha256:e94b5bf27933bd92bcf644c581ec4109eaed9f7978b64f52fca0c920cfdad1c3

Observation 9dbf1532-b749-464d-87f9-222186cb1dd1 · inbound

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs cites this paper.

Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.498678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T08:34:14.310656Z digest=sha256:0bd767e99cde55e491379d6bd448b60f722ab6fe0ecd901c30da2ac99f6c1de3

Observation 3dc7a409-db2d-4c2c-bbd3-d78c2cdd19e2 · inbound

Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions cites this paper.

Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:30.228824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:31:40.195348Z digest=sha256:fa2ab9f1aa3a41426f91739b7d9f934e0c267341ebb36df18abf3e7c87a7335c

Observation 1d6b79df-6825-44b4-9e58-0f800a339e40 · inbound

Tracing Persona Vectors Through LLM Pretraining cites this paper.

Tracing Persona Vectors Through LLM Pretraining A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:29:27.871025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:28:17.086117Z digest=sha256:79dafb4d2fc57d5f60994d0d86bcbabf282e68e4544f886a0244d7f71a53d548

Observation c509664a-987b-47ec-b0eb-42feb21884fc · inbound

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models cites this paper.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.711820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:b82b5e4d0ad06bfe1cfe26a05f912fbec19541be62f46490b4dd3850b58ad6ae

Observation bcc42c2c-c0f5-4d3d-9cc5-a17265b858cd · inbound

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability cites this paper.

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:11:29.059463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:07:18.198225Z digest=sha256:478fdb475bcf30d4c89e529bbe03f88600af96e075db126381044ae735a5a478

Observation ac599741-e851-4fce-803a-7f96c139ecba · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.949014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:33a6b037412765d27a4db5b1955d2a388d94af5055f6e06177c5c4c37c1cefa9

Observation c6a21310-61a7-4cab-873d-3f68b75b92a5 · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.166430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:c3cfa660c3a8e56f90f80a58a17a50852089f9618ce57a32d29d536c09084446

Observation d45aad08-ea90-46c5-83e3-9014c53f1898 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.279218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:486452afaa2ec11b455df05136f38304e37027991630358f8e1e8cee281d144c

Observation 4b136081-fd06-4b5c-8f8e-060067cc56a6 · inbound

Tracking Representation Dynamics in Large Language Models with Persistent Homology cites this paper.

Tracking Representation Dynamics in Large Language Models with Persistent Homology A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.225595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:03:07.253265Z digest=sha256:017dffdbafddd44380491986845c234e01f12bbc0a44768714630fc6faf46188

Observation 739358e2-e78f-449e-9bf9-e17a8ffacc08 · inbound

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale cites this paper.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:4c699304a902eade3fc47ba6d8740b1a224fd9fbede6ae9c822bd617733298c1

Observation 365eb461-c285-4511-9f18-379cb93a83d8 · inbound

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates cites this paper.

Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:32:04.215688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:32:04.215688Z digest=sha256:e6561f3450efcdfae756191e42967846fa6088d9fb5edc538b17239df5a6defb