Pith. sign in

Paper Citation Record · LEDGER

SDD: Self-Degraded Defense against Malicious Fine-tuning

As of 16 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 3 inbound Pith citation observations for arXiv:2507.21182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21182 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:55:14.070225Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:46:08.836451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T04:45:00.946977Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19693b82-36a2-4966-b18f-52033639ce59 · outbound

This paper cites GPT-4 Technical Report.

SDD: Self-Degraded Defense against Malicious Fine-tuning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.725286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.725286Z digest=sha256:0c7d9830a40b968d04c10670e7e730e68f1d1c229187a5080ce7f504acf5f91e

Observation 5a7d25de-d0b7-4356-9b7a-ba51c792bef9 · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

SDD: Self-Degraded Defense against Malicious Fine-tuning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.736803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.736803Z digest=sha256:ee3149d980423c3fbc6184800074d2cc51c5842f5271d5516e3510d4ac33264c

Observation 0ac7078a-6ac4-46c3-9844-d444b27a95fb · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.198060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.743195Z digest=sha256:c043c5618f2dc37ade03c7cf54501a0fdeed502f2bfc30f965b5c94c70a603ff

Observation 7a6821dd-634d-4643-836b-c012c6f4e095 · outbound

This paper cites Invariant Risk Minimization.

SDD: Self-Degraded Defense against Malicious Fine-tuning Invariant Risk Minimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.748348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.748348Z digest=sha256:747e85ad2054e0547934718c449503c05dfd6ed0c59496a419e37fbfd3fe7914

Observation e0928c9a-0e0b-4a46-8acb-16bd7f5c0ed8 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

SDD: Self-Degraded Defense against Malicious Fine-tuning A General Language Assistant as a Laboratory for Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.753911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.753911Z digest=sha256:2e4e4fd2e6fa0477962fde3fbbf5310b563515fa3d57b843dcd87faa9395d336

Observation e079766f-74e4-41ad-9939-9923773acebd · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SDD: Self-Degraded Defense against Malicious Fine-tuning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.759174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.759174Z digest=sha256:8ebaccd295865c34f2ebd22aad58ccca1ae1173fc2533375493d8cfaf9e08ffc

Observation 9548ae66-4538-41e0-9c68-f900bab0d174 · outbound

This paper cites Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases.

SDD: Self-Degraded Defense against Malicious Fine-tuning Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.765034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.765034Z digest=sha256:490fbdfab5225ec6046646da5c1c2d39f5509c940d2d097a9a08b13441ac1da4

Observation f3d19aae-fa94-479d-b7b8-225520d0a681 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.770029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.770029Z digest=sha256:003d39d5702b875fb53924ff01e766983d18858dd23159fc081552eb135e3dc8

Observation 0560f7f3-a974-4fb3-858d-fe4b17f882a0 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.170421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.774585Z digest=sha256:7b36fdf705cec2f5a0dea1019a4f86c66d7712679d7cd4aa6a5e5e8d67df780f

Observation fea9462e-ecda-4490-b008-e3c95d2f2d9e · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

SDD: Self-Degraded Defense against Malicious Fine-tuning PaLM: Scaling Language Modeling with Pathways

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.779264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.779264Z digest=sha256:67493356b24fb157aecd6ac24b38a8a7ae149562ad6d36b143bae96ef96933b4

Observation df6cb45f-5197-4121-9d2f-8cecba6eaa46 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.784457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.784457Z digest=sha256:190614d7ad5e6e444db1be6c2eba95d5a4d4fb8abe6e1352548bd4aa590789a6

Observation 024a154c-afbc-4364-a7cd-85d5d7c6eb1b · outbound

This paper cites BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B.

SDD: Self-Degraded Defense against Malicious Fine-tuning BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.789814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.789814Z digest=sha256:f2336ab8c6d2196bfda47384a6a222ad924b69eaa519ccc31e6978730f8d0a93

Observation 96a3381c-ad50-460b-9fc0-ac9a05f6c113 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

SDD: Self-Degraded Defense against Malicious Fine-tuning ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.794700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.794700Z digest=sha256:67eb8825cab9add7647e09662f2552f37d515f4b5f14c050542c51d12f5b5744

Observation f34a0f7d-d202-4194-ac34-981ac8a18d5e · outbound

This paper cites Will releasing the weights of future large language models grant widespread access to pandemic agents?.

SDD: Self-Degraded Defense against Malicious Fine-tuning Will releasing the weights of future large language models grant widespread access to pandemic agents?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.799482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.799482Z digest=sha256:c14044495000988d755292373dd775c8ae5e80b43ccaade45a536e3dfe848ace

Observation 5cacd874-ab0d-458d-9ba2-2d0a5f4f251b · outbound

This paper cites Spear Phishing With Large Language Models.

SDD: Self-Degraded Defense against Malicious Fine-tuning Spear Phishing With Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.804188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.804188Z digest=sha256:1f848f3dd8f695275fd652278ca5a4a64de12f3a30d1635585f70b156e4b7f19

Observation 5f5cc971-ca58-4ecc-b5a0-41caa113420e · outbound

This paper cites Manning, Dan Jurafsky, and Chelsea Finn.

SDD: Self-Degraded Defense against Malicious Fine-tuning Manning, Dan Jurafsky, and Chelsea Finn

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:55:15.145099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.809512Z digest=sha256:4f95437a677cbeab0212379a4506e56ceaab0bcae1cec96f5919e1c4b12250d2

Observation 3cb7f07b-8f1b-4ce7-bc1d-aab912b3cfc7 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.127370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.814221Z digest=sha256:8fd43f0e3f3a1417282483c693916efa319fbd4131784d7b38e2dc3058f0347a

Observation 8bba5e72-74ee-4bf1-aed0-22d240228810 · outbound

This paper cites Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model.

SDD: Self-Degraded Defense against Malicious Fine-tuning Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.818599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.818599Z digest=sha256:198c28da70f39771664416dfa5f6a30b20b363aa4239119fe5f0b39d1f2d5c80

Observation 3bd5f0f0-a413-460d-9fb1-a00c2f80df74 · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

SDD: Self-Degraded Defense against Malicious Fine-tuning Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.824400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.824400Z digest=sha256:53c6a9d9bbba38f0cadadb01457171d00c63a7775e6c0a823118e8f15758f811

Observation c44ad53c-8969-49d4-b7f4-e87dee755ab3 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.109642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.829865Z digest=sha256:93e05690d2dba1afcecd5927bc489d6d12044c55a5ac9cc8a26e91b308efb859

Observation 49fba8c5-efad-467c-9814-a62acba02429 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

SDD: Self-Degraded Defense against Malicious Fine-tuning Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.834761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.834761Z digest=sha256:9b4b5ee643e1957945a3ba181438b89e11ee38692fc0890c897823572d47c87b

Observation 3a703f7f-b66b-4b07-9f4b-eb0a409e6e4e · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.839797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.839797Z digest=sha256:f3b64135c819a10a6f327b608248e3efc1f7828afa71549c6777aa8ddab6cf34

Observation 16691936-c84d-4cfd-b88f-ad1bdd547309 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.080480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.844486Z digest=sha256:aaac99efc2ab3136e72c73c52d398eac32f56c09405253e0ac1e462cafa25c7e

Observation 2d34fdb0-117e-4693-9828-ae8a0c129d4b · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

SDD: Self-Degraded Defense against Malicious Fine-tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.849490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.849490Z digest=sha256:730c2f4b8f818547851ab45482cc9fdfab0e7117e94f9fb954b825a20e8aec77

Observation c433e800-24dd-42bd-97de-7f0a43cdf9db · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.062431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.855311Z digest=sha256:a32b197052b48e749c02cf74eb5d1d66e39ef86cad753368c706698e18dc7abb

Observation ef7c91a6-7714-45f7-84c4-ec3ad290b026 · outbound

This paper cites Spurious Feature Diversification Improves Out-of-distribution Generalization.

SDD: Self-Degraded Defense against Malicious Fine-tuning Spurious Feature Diversification Improves Out-of-distribution Generalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.860473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.860473Z digest=sha256:02c1a44ac38544fd20741ae46f58c1017093ec8aa0c6eff002d5d1acc70132b2

Observation d753062d-7496-4b61-aca7-e85fe104860f · outbound

This paper cites Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation.

SDD: Self-Degraded Defense against Malicious Fine-tuning Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.865845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.865845Z digest=sha256:1cc4e2fcff354ff5ff282df68bad43719cc3a7f8a896b649dab9ca02e28e23ea

Observation 218c90f4-0d96-4391-a0fd-c8f4b35fb6dd · outbound

This paper cites Robustifying Safety-Aligned Large Language Models through Clean Data Curation.

SDD: Self-Degraded Defense against Malicious Fine-tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.870946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.870946Z digest=sha256:a9f3c428fccbe0e0e9d5b39adf3cb2f8d502f048ebeec76095b428db9bc6a06b

Observation 139c2e72-1d87-46fe-ba6a-a0cde68c243f · outbound

This paper cites BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine.

SDD: Self-Degraded Defense against Malicious Fine-tuning BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.876183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.876183Z digest=sha256:d1878c4b81779584df08ead0a10f815e855d2f4c660eb4e185780a2c9e605b15

Observation d0b3af2a-d5aa-4a63-8a1e-37716d4fdfa1 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

SDD: Self-Degraded Defense against Malicious Fine-tuning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.881021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.881021Z digest=sha256:e34c641e223eaef05b121edb7c2906b574d599d916fc5519987e24f74a370d6f

Observation cf9558ce-f833-4d55-9891-fdf2ad5b072f · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.045707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.886584Z digest=sha256:e63bc3d8f1b36546d61fc299eb9e3db0f3d1da6e759dfb01e452cdfcf8aaeeda

Observation e04cf625-f0a3-44c6-9259-b55b0ac7672d · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.891970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.891970Z digest=sha256:9c04ff01bd700efaef99306ca609f8ff3d165c7cad878fc077c0f2c66f2a2313

Observation f6596feb-7fb1-477a-b1ef-337b08a99336 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:15.014126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.897479Z digest=sha256:514c9378250c1cfbce783ce330fc68ea23890b384fe47973f2e5d17346b42bf2

Observation ac9dd152-c508-4102-a46c-86dd66dd27ba · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.993873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.902020Z digest=sha256:aafa9b95896e6bff38d5945880119da8d0dbb57cfeb2cc2d510efb8132a2e73d

Observation c54ba3fa-9e99-4458-8fb5-c8cae4f4c5c2 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.976518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.908358Z digest=sha256:ce0e27cc8840b5b6c68897a1d584aea0b6501f85feef37fa92b1f1674c386440

Observation af6d1605-f5d1-47e7-898a-228304c67d56 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.957139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.913713Z digest=sha256:68287d50dbf2d31f294c7d51d330f56c9c72f80b73305f0912692ca526fa59c5

Observation c28db790-7490-4122-ab14-0aea2495f2b1 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

SDD: Self-Degraded Defense against Malicious Fine-tuning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.918912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.918912Z digest=sha256:4570bde51400a63221ca0b059673d1773bc7097908eeed1949d5e91f73fc6bed

Observation 7c4a1bf2-204e-4ff9-adb8-2e6c9304e885 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

SDD: Self-Degraded Defense against Malicious Fine-tuning Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.928361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.928361Z digest=sha256:82ded32b88e616f6766f30f4783e8f194533ef9cfde442dec98e0b7ccbe03df1

Observation 72a95aad-b02b-4994-98fa-7f5dfb343766 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

SDD: Self-Degraded Defense against Malicious Fine-tuning Manning, Stefano Ermon, and Chelsea Finn

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:55:14.939942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.934314Z digest=sha256:fa1377a1475ef52378fc9f350210381512fa4eeaa821f8291224472d3722090e

Observation de66096f-486e-487d-aa45-cf8eb4c76568 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

SDD: Self-Degraded Defense against Malicious Fine-tuning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.939841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.939841Z digest=sha256:0baad0b3492323d492da47f7299e473f7811fe5af53facee93613b0091f79e68

Observation c19c6b36-8c91-4812-8500-d55804f2845e · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.923138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.945021Z digest=sha256:61f2e7d8af9c562a025eb64d57e4fba20e90b20689689b8a6b6f4c96a0b8ee56

Observation 4fa50c34-ca7b-4f84-8ade-5d6d844bfa87 · outbound

This paper cites The Risks of Invariant Risk Minimization.

SDD: Self-Degraded Defense against Malicious Fine-tuning The Risks of Invariant Risk Minimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.950811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.950811Z digest=sha256:a0bf08fee34ddc0179d85f3bebf5de8a24cb944d30bd3e7b5addaee03ac8446c

Observation 759cadca-3f04-40f2-9ad3-3128302714fe · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.907552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.956513Z digest=sha256:483ab9d9f714aacf3320fb59b87a0dfe22f65eb8a8987d4ea0b8d3b1dfaf016d

Observation a4dbacc2-2948-4795-a75c-1a05cf4278a3 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F.

SDD: Self-Degraded Defense against Malicious Fine-tuning Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:55:14.887588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.962549Z digest=sha256:4195552ab2d5dab80964ad5ce89f049c0451da5dbf75382c37bce31d61b35401

Observation 40c438cb-afbd-44c6-8952-a8430d44453d · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.863872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.967973Z digest=sha256:d4ec65f819abf9e30f9f669378a3e095b6456d2947bfca3b701936c10fdb3a54

Observation 08fc65b0-8348-4838-b18b-88f08afcf5f1 · outbound

This paper cites Hashimoto.

SDD: Self-Degraded Defense against Malicious Fine-tuning Hashimoto

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.972796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.972796Z digest=sha256:914f1e4bde2cc761884bf0a6d7c8a6d3ac29dadec77738d74fcd7041de27aac8

Observation afc5bc9c-b228-49f0-8183-97d65de57d47 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SDD: Self-Degraded Defense against Malicious Fine-tuning LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.983176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.983176Z digest=sha256:85df49f3db14c39d48d9cad9ba7192f12a404a957297642048924921d7c44b4f

Observation 76a7a51a-73fc-41d8-9c42-b5831e623813 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SDD: Self-Degraded Defense against Malicious Fine-tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.989079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.989079Z digest=sha256:0264538612c96c9d7786fbf70f933003d66089077481a1bce296b296b45f22b1

Observation 9a36cc09-9f00-459e-8dc5-5a2558606ac1 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.833702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:13.993870Z digest=sha256:580f257dd037cd0a8754246967fee9f21894b4738e92e9d3c019f2a528c60962

Observation cd38159c-9305-4696-89fe-48a08a622172 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.811946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:14.000306Z digest=sha256:c25231efb81516b1b2f30608ca38bb34567e34ad586492f3353045f165714487

Observation 1f295d79-792d-48c6-b9de-c9f683ddced6 · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

SDD: Self-Degraded Defense against Malicious Fine-tuning Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.008458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.008458Z digest=sha256:f7682787a7202419791365973c87206ce27f97b6bb916536c55700dd53576ff9

Observation 5358107b-11fe-4fd8-9e1c-3ef9ada4c501 · outbound

This paper cites Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M.

SDD: Self-Degraded Defense against Malicious Fine-tuning Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:55:14.786325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:14.013517Z digest=sha256:c0bc599ac35ce90a826b22a9bc8b7393c1e16d30d1bf78de3c8d8ea1a1c26298

Observation c5aa7db9-83ae-4f37-8887-021ba9e8eade · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

SDD: Self-Degraded Defense against Malicious Fine-tuning Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.018443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.018443Z digest=sha256:00543401ac56bedf98af84edb30626ea801202367ff023197b9bf08c404e7c30

Observation 9b9ab330-1cd2-4a80-b88a-dcb1d0c34424 · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.766705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:14.026348Z digest=sha256:4b612e0c55986d5576b68c1d875fcb182179ec28cb814c989ed18dbf4718c0ff

Observation b7964131-1802-4170-9105-c6f4ee58b793 · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

SDD: Self-Degraded Defense against Malicious Fine-tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.032498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.032498Z digest=sha256:fedfad4b97509d0266e7a59e1c0088cfc052e875d3a94a1736213b4ad190cdf9

Observation 6d69086d-697c-4e8c-9585-f6b5bbf74fa8 · outbound

This paper cites Self-Rewarding Language Models.

SDD: Self-Degraded Defense against Malicious Fine-tuning Self-Rewarding Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.037842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.037842Z digest=sha256:2f8dcf67044ff4e3a94674f276675633b877b4844dd9ebbf1d8874400c4429af

Observation 66e8f456-5755-460c-97bf-a83e487df1c3 · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

SDD: Self-Degraded Defense against Malicious Fine-tuning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.043471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.043471Z digest=sha256:afc91e3b912375c83233020e643070a8178d1a714fb5c1183b94423a8464be52

Observation b786a24d-dde6-4123-97a0-53bc3771afcf · outbound

This paper cites an unresolved cited work.

SDD: Self-Degraded Defense against Malicious Fine-tuning Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:55:14.736732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T17:55:14.048923Z digest=sha256:4649e7ca224fbc64092b8dfe3712b556ef936ba2210df58759d86f8ce1287647

Observation a0b91e01-9ded-4ea9-8f5f-030f3982a2c0 · outbound

This paper cites Making Harmful Behaviors Unlearnable for Large Language Models.

SDD: Self-Degraded Defense against Malicious Fine-tuning Making Harmful Behaviors Unlearnable for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.054517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.054517Z digest=sha256:484751d63325fd7d40d5f84b059fa12c8e0f3509d65bc22dbc8377a8cf59d8a3

Observation cea0e38b-c00d-4aa9-bf07-bfd4db651030 · outbound

This paper cites online" 'onlinestring :=.

SDD: Self-Degraded Defense against Malicious Fine-tuning online" 'onlinestring :=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.060124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.060124Z digest=sha256:d2b6251a6735f97821710acede7db3a0b19354715a5736c396453ac578be9d1b

Observation 91d9699d-f303-4c93-99f8-1c709b823b88 · outbound

This paper cites write newline.

SDD: Self-Degraded Defense against Malicious Fine-tuning write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.070225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.070225Z digest=sha256:70f5abf935d59207c5b30cb9870c478510e3682c6ebf8535a1164a1c682ad207

Pith citing papers

Observation 9621278c-7a10-4eb5-a00b-f6a485e84322 · inbound

Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps cites this paper.

Immunizing 3D Gaussian Generative Models Against Unauthorized Fine-Tuning via Attribute-Space Traps SDD: Self-Degraded Defense against Malicious Fine-tuning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:55:48.899871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:30:06.396482Z digest=sha256:21a0d35ea1fea85c987e7098181d262e85bbe0479a6be2e2d0dc2d30aa1286cb

Observation dcc4fd69-422b-40f8-8ead-9904507bdeae · inbound

GradShield: Alignment Preserving Finetuning cites this paper.

GradShield: Alignment Preserving Finetuning SDD: Self-Degraded Defense against Malicious Fine-tuning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:45:00.951155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-15T04:44:06.614390Z digest=sha256:f8d60215926d1a0ed97f46fbdf21bb97c894aab133e4a7424cdc39be35d7a4bf

Observation 988ed788-82f9-4433-bb8a-ba90cde03218 · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace SDD: Self-Degraded Defense against Malicious Fine-tuning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:08.836451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:08.836451Z digest=sha256:123e466aad565d1b91a204b8cfec60d15ad7cb03aeb28ce6b8c7ec08c4cc6c32