Pith. sign in

Paper Citation Record · LEDGER

Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2402.05162.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.05162 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:10:21.164769Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cbea1368-0965-4239-a6f7-65a30a5b4428 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.149761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:05b1aeb7b60aba7a607ccaad9196f05506a5b9570c5d17072999b59ca5995860

Observation d19ef127-391d-4692-9095-2cebbeb04a3d · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.085448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:8f572561b62e9a43a290e1ab0b9bbc8b86e212e0d1baf9931c02cd532e24e452

Observation 83cee028-943d-4118-8b61-a51410ccfa8f · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.996646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:9e5951bd9cb5f19015b22e6edb616d6b73466784f76c35cf591e5e4b0c1ca2ef

Observation 0c2c6bea-7436-4b5f-9ab6-3ead28b1ec1b · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.164769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.164769Z digest=sha256:f4fa62cf6d27bd12ca0d07684aec0dca5d2dd30680d336e423516f26f810a6cd

Observation 4014d4f8-ebe4-4157-9e85-317ab0948343 · inbound

Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2 cites this paper.

Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2 Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:58:18.996512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T18:55:48.540435Z digest=sha256:3e293b76da529d5ea8ad9c45b475d2f75dbbc545c3a9f3c91276b4297bdcfb68

Observation 16c81941-8b7a-49bc-b0c1-c2975aaff959 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.571809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:4f2c35fc9716fc652430e746883c0d28373e021cecbb2077700e760c3f627a62

Observation 14bad936-292e-4a84-afaa-3747711aabec · inbound

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression cites this paper.

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:18:01.269418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T17:16:47.284564Z digest=sha256:8e0be29a124d0921b4602dc698f5ba5a14fdf449a38c2941d7cf71afa38b967e

Observation e1abcea7-6b34-4377-a29a-6c4741cb95eb · inbound

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints cites this paper.

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:01.737881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T16:04:25.851592Z digest=sha256:6a20d8a6008c6e8faba8563b7c0bb0421cf9ddf0a155899ffe784ff4b9ae70ac

Observation 39072647-bc4b-4d17-afb2-9e152a01511a · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:19.355933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:9c45255bfc06b3a828018d7fdc88d013e3001c19c9a5d8fee759d28cfbf4d7c9

Observation 7f3f0f3a-43c0-4b6b-8a7f-b6bd7a3dd9a8 · inbound

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models cites this paper.

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.385976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T01:41:44.898358Z digest=sha256:ce365eee3641a15a04ff3ac6d3881181192aef7a456a7c4497cc46417c8f2f64

Observation 535a40d2-b828-4c34-90dd-52b3f3e81075 · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.787885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:44:09.214779Z digest=sha256:4ea7fd516ceab6dccd7e38f5950c515361f3b64c873684a65b039492ddbfb5ea

Observation 49544792-b9a8-4f1d-bd8d-480e17de09ce · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:47.607612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T22:05:29.444682Z digest=sha256:c0637cf1b33189592da01b0d9c233e3ea9e333631b9d7cca5bd61babf669683f

Observation ab30518a-aa99-46d5-83eb-2fcf73a70a86 · inbound

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling cites this paper.

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:08:11.097483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:08:07.648295Z digest=sha256:d69a213c6e72bb6f71c6c8a31e26596a0b4df68f91cc76f74c7c7abffc4d1f92

Observation 2c8cc1d3-9a0c-4721-8b48-3d86c6f18b52 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.940745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:e68b74be9529ff751b577bde55517f314546a9b4c1f6233aec0e83ad108ea388

Observation 5b719db5-9a7a-4486-8ebb-3a322659aa89 · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:47.824267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:19af659d64e774bde0618301cfa86843a343d4cbe159dae0b066be575f3d0d90

Observation 9459fa0c-106d-4060-b66d-bd0e2592c9f4 · inbound

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets cites this paper.

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.485913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T14:44:57.205766Z digest=sha256:baa9dc933c024222b8a6ec7278f32cab5b8d26076b5e0547bad587386d69a883

Observation 4669ea2f-2c5e-4dd1-a35e-4804a3b3a840 · inbound

Faithfulness to Refusal: A Causal Audit of Neuron Selectors cites this paper.

Faithfulness to Refusal: A Causal Audit of Neuron Selectors Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-07T15:43:53.891301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T15:35:54.665268Z digest=sha256:bcd0d109089ede5c51d097c6550a6bc2b542ffa667263e502252392ff93eba70

Observation bcc8117c-adf3-4adc-a3af-4b28f7f06d9f · inbound

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge cites this paper.

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T06:50:14.933979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:50:14.933979Z digest=sha256:3ea0fe7036eb2cdb016b4bb396f16e790fa976af27fb062a68c2ad98f24c2040