Pith. sign in

Paper Citation Record · LEDGER

What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2407.10264.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.10264 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:34:31.210142Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1189faa7-1708-40d2-8270-4626c3f42098 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.141571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:d605e93c3ce4daa914d4c59653eccde643b0c9749cb9162b60d993604e4e7bb9

Observation 2668cb19-a0eb-4783-b1ad-4cc1f96dd804 · inbound

PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning cites this paper.

PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T10:19:26.617429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:19:26.617429Z digest=sha256:a63f848af4e7acc9fdaa1f077e4008afad6faefa11279ea99c9f547ecf520a1d

Observation 2ded35e1-2345-404e-8218-711173e0cdef · inbound

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models cites this paper.

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:05.986317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:05.986317Z digest=sha256:90973ba73ac5da6f3568e4bfdf8f142b75b22bd98463930fc79752ce26f44185

Observation c093c635-2b6c-4879-8aa7-3b2d805c3683 · inbound

Obfuscated Activations Bypass LLM Latent-Space Defenses cites this paper.

Obfuscated Activations Bypass LLM Latent-Space Defenses What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.575679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.575679Z digest=sha256:365eba546ce8e6ff7871c179d1ef4b7da4e8d369dce4231995dd474c662dc1c7

Observation 6b9a9920-c739-4a89-8c69-f234c6012681 · inbound

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions cites this paper.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.274152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.274152Z digest=sha256:16c4d66b3d82b0cf73dfe023904dd1f8674f3cc84270f40e0d4374662b2536f0

Observation e7d42a64-65ba-4811-a66c-b4fb3ae22bbf · inbound

Layered Unlearning for Adversarial Relearning cites this paper.

Layered Unlearning for Adversarial Relearning What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:34:31.210142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:34:31.210142Z digest=sha256:13fce8c25468543d54356a66c53af48561926e7e1948ce76bd9633b98b74027b

Observation e7173816-d650-4d15-9788-34446761e211 · inbound

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? cites this paper.

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:29.382934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:29.382934Z digest=sha256:c466d683ebf91e1bb3d58e7bc63d33affd7a7c26b5ed7cd969de02c904654430

Observation 52b80434-8e73-415c-8e10-c27f4df731a7 · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:22:18.842470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:ca15e12fb646ac200266b5702a527e45844e056135ba155be8b1401113c01074

Observation 0310fc50-cfa6-42f9-8f08-c536caf3c83e · inbound

Efficient Safety Alignment of Language Models via Latent Personality Traits cites this paper.

Efficient Safety Alignment of Language Models via Latent Personality Traits What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:27:20.138378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T15:26:23.290009Z digest=sha256:7bfdb9cf6b0ccf91d5c269edf6001c4d82762600b5a9ca998703ffe7229b08d2

Observation 5c4aa29a-72e3-4562-89ae-0cbbc73cc1d3 · inbound

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning cites this paper.

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:29:55.413383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:29:55.413383Z digest=sha256:5796b6583577d93e498f49c90391a14d4de41a6c1e2c11f21c840262a45f20a0