Pith. sign in

Paper Citation Record · LEDGER

Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2402.05162.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.05162 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:09:30.214370Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cbea1368-0965-4239-a6f7-65a30a5b4428 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.149761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:3039bb0be7110d7f023538542fe374c99346f2da9e8ba13bc4252fc9008ebf44

Observation d19ef127-391d-4692-9095-2cebbeb04a3d · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.085448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:fe61202ea2c8ce52278d14bcf900a9f572ad0795b59795e54673047f28f2cf48

Observation 67c7ab49-d440-4829-ac92-4c419c2a1b82 · inbound

SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs? cites this paper.

SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs? Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T10:57:22.935082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:57:22.935082Z digest=sha256:a3ae76354fdef7118cb8549f28eb40361f0a61463f3f5a6a73345d94779168c5

Observation f43800fa-7f45-467f-9eed-79540da46314 · inbound

Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models cites this paper.

Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:18.443220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:33:18.443220Z digest=sha256:04d4fce39fbc219f460e612511b27140ae204f3bb1981ec26b3e6d8c345ac23b

Observation b3597628-016d-4927-98e2-0a3546b160f8 · inbound

RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response cites this paper.

RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:51:02.957947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:51:02.957947Z digest=sha256:f336813528efe262e7c12cee1e807e76478c84870fda5a997f06e204753ad123

Observation ec982510-da81-403d-b937-775830088e3b · inbound

Open Problems in Machine Unlearning for AI Safety cites this paper.

Open Problems in Machine Unlearning for AI Safety Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:10.723771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:10.723771Z digest=sha256:bcfadf21470c630ab31f3a49aa7e1de35d9976be6f4ebf99aead6b2a1925ba05

Observation 37ee5772-0a77-4982-9ca0-9adb1b4979c5 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.645718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.645718Z digest=sha256:7462808775b9df94204a3f7fef220eef6a6cfb4407f1d35cd7641a482531334d

Observation fdb2e949-cf83-4f26-a872-9c43d77b03bf · inbound

Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers cites this paper.

Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T21:55:56.884625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:55:56.884625Z digest=sha256:e4b9cb09a485733ad330dca24ebaeceab1dc09137780a652f8e0e96027b925b2

Observation 7fa0df6f-d9eb-48c7-a830-30ddec005595 · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.669881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.669881Z digest=sha256:8ab665609456dba10b97259f7ca2dc288bac14beaba4d97575bb937685d2051c

Observation 7200a08c-27e5-4891-a564-699913e1474f · inbound

Breaking Down Bias: On The Limits of Generalizable Pruning Strategies cites this paper.

Breaking Down Bias: On The Limits of Generalizable Pruning Strategies Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T11:42:07.679330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:42:07.679330Z digest=sha256:98b041d7fb5905ccab33667bd799ca52845328fa459326f0f988e5968a06c1d1

Observation 737f3945-a6cc-4a1f-b817-56b3c983d795 · inbound

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions cites this paper.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.378812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.378812Z digest=sha256:3d8576e6a30496230a364e34ce8e9e3463708f337de82de06b6a8e59a374b3a5

Observation 83cee028-943d-4118-8b61-a51410ccfa8f · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.996646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:2020f74151b3a16f87c75e1e86d541189ab0952084bf16f1b18772749ab6e095

Observation a05ecda7-04b6-42b0-a71c-0096387892ab · inbound

Mitigating Confounding in Speech-Based Dementia Detection through Weight Masking cites this paper.

Mitigating Confounding in Speech-Based Dementia Detection through Weight Masking Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:40.675703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:19:40.675703Z digest=sha256:a6a5c1707aabc92bb4fdaa6d3d915a4e4d0bfd2c7b2aacc5952e1037e36ed65e

Observation 3b74d362-e9c5-440d-a002-3e225ae1418b · inbound

SoK: Machine Unlearning for Large Language Models cites this paper.

SoK: Machine Unlearning for Large Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:00.309903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:59:00.309903Z digest=sha256:6cfe4fd13585612c2cad3677486679027cf75ae86027d38c641d210753ced900

Observation 1417b31e-acf9-4e0d-8e52-d63cb862f27c · inbound

Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs cites this paper.

Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:09:30.214370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:09:30.214370Z digest=sha256:772220f064756c3638a96a78b54a8bdf664e5c0b872844acad7681de677c563d

Observation e52b27ab-0ec8-40c8-8109-dde23d632199 · inbound

Linearly Decoding Refused Knowledge in Aligned Language Models cites this paper.

Linearly Decoding Refused Knowledge in Aligned Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.011420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.011420Z digest=sha256:54c758351be0012750e6851843c7cb8c3f7cb0df489eb5cd4eac263ca7b2da4c

Observation 1f295d79-792d-48c6-b9de-c9f683ddced6 · inbound

SDD: Self-Degraded Defense against Malicious Fine-tuning cites this paper.

SDD: Self-Degraded Defense against Malicious Fine-tuning Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.008458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.008458Z digest=sha256:d28c4ab7871dc35e400f21acce983513da37531520b7f73ed7ea83cacfc8a333

Observation 7567ac52-f1d8-4b55-91f0-4a4fd39dcca1 · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:36.334354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:36.334354Z digest=sha256:8e9a7379c7b763c9633fad4715549072917be8d28276f40ba94de1caaada1efb

Observation 83bd386f-8001-44d2-b4cf-dc27d23048cc · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.939155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.939155Z digest=sha256:f16ee2671db4a1c873f5a56e7749bbe13eaaa6edd2a0789a639023965909a9f0

Observation 0c2c6bea-7436-4b5f-9ab6-3ead28b1ec1b · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.164769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.164769Z digest=sha256:bad8eedfe7cfb7bacb1cef0d90c243e0f456d92a3a0368b1cd52fbd1cb5081b6

Observation 4014d4f8-ebe4-4157-9e85-317ab0948343 · inbound

Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2 cites this paper.

Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2 Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:58:18.996512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T18:55:48.540435Z digest=sha256:8783ee952d655e098607784dc83e15ca673d6ed0f17b35efbc187802267a16e4

Observation ff754a30-a656-4eb2-95b8-cea41a614f96 · inbound

Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment cites this paper.

Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:35.061007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:47:35.061007Z digest=sha256:3ce8daa42d5e6ae490441fcea8929cffb595e98a732514409442d9ae94a588db

Observation 16c81941-8b7a-49bc-b0c1-c2975aaff959 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.571809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:91284c3b2abed07b8c2ed40508a711b70faf004dc6ab0995a14178578c60b08b

Observation 14bad936-292e-4a84-afaa-3747711aabec · inbound

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression cites this paper.

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:18:01.269418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T17:16:47.284564Z digest=sha256:79eb8c9c5bc28d2a447d37580be02c862e7dfacbe6b73e89b4464934e5e8a422

Observation e1abcea7-6b34-4377-a29a-6c4741cb95eb · inbound

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints cites this paper.

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:01.737881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T16:04:25.851592Z digest=sha256:1dd398ccc84e12c672e0a56a75129771dec6c348a0cb84854b795be6fe8a0170

Observation 39072647-bc4b-4d17-afb2-9e152a01511a · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:19.355933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:13233cca7cc157be3836fd2afb941f8e0ca69d94e61358d52dbcc1a30e652d98

Observation 7f3f0f3a-43c0-4b6b-8a7f-b6bd7a3dd9a8 · inbound

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models cites this paper.

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.385976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T01:41:44.898358Z digest=sha256:252b3bc3c7871f3d5e240c2fffdb889268165fd798bbe4aaaeb010d38e097c65

Observation 535a40d2-b828-4c34-90dd-52b3f3e81075 · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:47:58.787885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T20:44:09.214779Z digest=sha256:cbe17c5ee385396dd90453740cb056b305632f444647f55a39be30b3022f61b2

Observation 49544792-b9a8-4f1d-bd8d-480e17de09ce · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:15:47.607612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:05:29.444682Z digest=sha256:786a86c531ef0ee49e27c7ed640b65187a0918a3d690be81667dc38b4855b24a

Observation ab30518a-aa99-46d5-83eb-2fcf73a70a86 · inbound

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling cites this paper.

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:08:11.097483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T10:08:07.648295Z digest=sha256:6df5767b25e37d4715f457bcb0cd458e2ee2bd8ac6a60e46a60b7e40c77c8f50

Observation 2c8cc1d3-9a0c-4721-8b48-3d86c6f18b52 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.940745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:06808194fcf23e97cef5fc51ac402e70eb55b150eaf0f297e6242d3a4f035c74

Observation 5b719db5-9a7a-4486-8ebb-3a322659aa89 · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:47.824267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:4045dd5e37e3f047a10f1cb026f97209cffc48256d148bbb81bb805491b72a07

Observation 9459fa0c-106d-4060-b66d-bd0e2592c9f4 · inbound

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets cites this paper.

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.485913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-03T14:44:57.205766Z digest=sha256:008297a5c54545cab9da54f71002adf1c6703513dd377b2c1156a14ee698ea2f

Observation 4669ea2f-2c5e-4dd1-a35e-4804a3b3a840 · inbound

Faithfulness to Refusal: A Causal Audit of Neuron Selectors cites this paper.

Faithfulness to Refusal: A Causal Audit of Neuron Selectors Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-07T15:43:53.891301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-07T15:35:54.665268Z digest=sha256:feb2ded513a7a17bd06230afc4bee089c61ecc568f946a4f5ce5427f5e863711

Observation bcc8117c-adf3-4adc-a3af-4b28f7f06d9f · inbound

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge cites this paper.

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T06:50:14.933979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:50:14.933979Z digest=sha256:d4c878b548bf27a9a0d1090151a8fa0e60d457e32eb483c95752756a9d710787

Observation b4088ce3-657f-4ac1-a4de-4894f73b47d5 · inbound

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning cites this paper.

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.773439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:46:10.773439Z digest=sha256:f4a0dbb9360e00e5d276c9a9b4c702abeda798dbb15c566b878d6df4178b692c

Observation df4dee33-b099-436b-821c-ee82e644ba1c · inbound

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs cites this paper.

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:46:15.991367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:46:15.991367Z digest=sha256:dfc0044f16cd951ab92ce2acad533736a7e0d635100527ec99d220e360e87be3