Pith. sign in

Paper Citation Record · LEDGER

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

As of 8 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2506.03850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03850 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:58:32.629027Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:01:25.549340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.110492Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85be4b79-7bc6-4a4b-8739-558370a5ff14 · outbound

This paper cites write newline.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.960888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.960888Z digest=sha256:742ba4567e9a646e6fc15bab9c09fdaae32a04744a1af24d8c0a8a96dbf8463c

Observation 9201a311-3786-4f57-8a1a-d464c3807d6d · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.049879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.049879Z digest=sha256:0c8710c38d03f3f941e5d416b6877aec41eeb119183b3d7782d9a4770eb03273

Observation 07745fbe-1cc1-4cbb-8cb0-60878b663edb · outbound

This paper cites D., Melenberg, B., and Rennen, G.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning D., Melenberg, B., and Rennen, G

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.806351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.166245Z digest=sha256:57aaea169260f0ba3c3435d2ac00a008118d69e126d25a7a66777a70bfe96d41

Observation 979dc913-5965-452a-ab2f-f847184f2920 · outbound

This paper cites Curriculum learning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Curriculum learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.617713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.274612Z digest=sha256:917916f4560814efc2bf7b06b07419865a2e58e3714ebb399dd6dea8a2d2f51e

Observation fc7b2970-884a-458e-87cd-2e018560ef7b · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.349795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.349795Z digest=sha256:332d82a8818ee7c94660837947a16ac6333b100310b60ffba10d00ab4e28858a

Observation 1b388d9c-4b5e-4e9b-a117-6132bfe03b9b · outbound

This paper cites Beyond factuality: A comprehensive evaluation of large language models as knowledge generators.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Beyond factuality: A comprehensive evaluation of large language models as knowledge generators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.468902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.468902Z digest=sha256:00452c1e40908190c1d5673a0caae3d6ac8fdc65037617cccffa366a4dd6e5b5

Observation 6b99841c-5412-4628-8e93-984729d3a5e3 · outbound

This paper cites W at ME : Towards lossless watermarking through lexical redundancy.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning W at ME : Towards lossless watermarking through lexical redundancy

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.552869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.552869Z digest=sha256:8f5567af1dc5e8e8e6eb85ff6a9773a733cc721a77ad2f21326438ad285cd823

Observation 8d33798b-9766-4add-9d45-ff2fc5d2eb2d · outbound

This paper cites Simple permutations can fool LL a MA : Permutation attack and defense for large language models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Simple permutations can fool LL a MA : Permutation attack and defense for large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.415083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.641840Z digest=sha256:10958bb3bafb3451757faac5f76e5ea394ee3e2fa85d8f7626de88990dd4de61

Observation c4a364ec-717d-4543-a634-bba1b2c0d6f6 · outbound

This paper cites PEARL : Towards permutation-resilient LLM s.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning PEARL : Towards permutation-resilient LLM s

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.239385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.780718Z digest=sha256:cd731b8c1e8c99c92656c71e424b8a5d1efd4f3ee1d35181f78296e4968abe05

Observation d4017aa4-c6e8-4a63-893e-37056372c498 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.922486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.922486Z digest=sha256:f9d185eff499a0455f2a8ea8655bfdfd0568fece125f9b13bf1c35683d21bf45

Observation ee26b7b9-f979-4f71-b552-d64461331b23 · outbound

This paper cites Statistics of robust optimization: A generalized empirical likelihood approach.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Statistics of robust optimization: A generalized empirical likelihood approach

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.054218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.035354Z digest=sha256:0321eb9182cec0050d8911179907c527daa0bc8d46a101bdd41155954d038574

Observation 5b6e0b8d-9112-4c43-abf2-31217a9a4b2e · outbound

This paper cites Sharpness-aware minimization for efficiently improving generalization.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Sharpness-aware minimization for efficiently improving generalization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.861219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.096459Z digest=sha256:6083963d79513980d9cad9c2d7cd6e127b768c0219a6fc45fcefcab73b877a36

Observation 0a7329e2-0b63-40cb-a845-8e5fd7ed970a · outbound

This paper cites A Survey of Uncertainty in Deep Neural Networks.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning A Survey of Uncertainty in Deep Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.226049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.226049Z digest=sha256:4db2d9fed166a2e5061c955c9f867fccf8212a4ec54b9ad6cdfb2a5848d130b1

Observation d79474ea-dfc3-4cb8-9fc8-8c788cb0b3da · outbound

This paper cites Safe lora: the silver lining of reducing safety risks when fine-tuning large language models, 2024.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Safe lora: the silver lining of reducing safety risks when fine-tuning large language models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.635342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.316371Z digest=sha256:e3f567279f2cbcb3b048f0bf4c2f026f9f4b190303789f19eecd82855c87226b

Observation 3039b9af-f8d2-4aeb-b693-0f6cb6283dd8 · outbound

This paper cites Does distributionally robust supervised learning give robust classifiers? In International Conference on Machine Learning (ICML), 2018.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Does distributionally robust supervised learning give robust classifiers? In International Conference on Machine Learning (ICML), 2018

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.437632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.426723Z digest=sha256:9f81d059f34bad6fa4f23dd2c9657798d5f5858fc810a8a2a320bde3e3987c3d

Observation 5ce7a588-a395-4725-802b-23fa97cb6001 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.556438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.556438Z digest=sha256:2c615c803738e002859c5a1959156ed03618bdd16f4300efbe126aa45c5ed872

Observation 432e1bfa-79a9-4200-a759-08918ef0de8c · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.726809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.726809Z digest=sha256:3da6b4042b57ed40e120631c15712b9f0c8dd8585c5c9b82af4716712cadeba0

Observation 68912f69-44e1-4a26-9e9d-4eccb783f9cf · outbound

This paper cites Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.862212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.862212Z digest=sha256:8af8b1257d643caf8e301282fb5c8673ce182fe98be358d4462254f44bf51635

Observation 2b67a653-2ba4-40ac-b187-9c56cd80cf82 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.014066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.014066Z digest=sha256:d991bd5447923f8570fc4058b12c952716443e195d9476260e46f6d26c4d047d

Observation 85e40ba6-bc56-47e8-9791-7ac500928d3a · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.125758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.125758Z digest=sha256:3bda4832ea84bcb8f525b24a044d4c7252007897e5ee5a164fbd1b7bd3e307ee

Observation 7b62c561-52f3-42b0-9a61-828f87d4bb22 · outbound

This paper cites A watermark for large language models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning A watermark for large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.258151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:30.240558Z digest=sha256:24fbf5fdfb14a25479605e5b5c276e562fbf9b22ae55645b2e93c823871b063e

Observation 9045197d-47fa-47bb-aca2-3d50cea61a3d · outbound

This paper cites and Zhou, E.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning and Zhou, E

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.043488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:30.373584Z digest=sha256:d0216a17e0f99d2500e3637792f8b739084628f01d3358f7e8c31356d34163bb

Observation 992acd13-9279-4488-99b5-01a2cc5cbbb7 · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.489568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.489568Z digest=sha256:eaebfc7494c4c852a9efb812cfc3f4641422f69c98467d8c012a8c30a79585d1

Observation 56a9ff7c-d06a-43d1-a381-a3109563f274 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods, 2022.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Truthfulqa: Measuring how models mimic human falsehoods, 2022

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.674810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.674810Z digest=sha256:d42f266d82487e0525ce2475e4b0279fc62c15de0bf2c8050df21a7a39b33c53

Observation 51fd9752-942b-42b2-90aa-7a1795e26c32 · outbound

This paper cites Decoupled Weight Decay Regularization.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Decoupled Weight Decay Regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.814159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.814159Z digest=sha256:d487e49f439a3943ce0fcb99f84351ea026126375b8a1564feebffde42477e11

Observation 4fed7e0f-c7a7-4eb8-808c-4691a40cde4a · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.941185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.941185Z digest=sha256:8857574368803e33fb8f33eaf28e2f5b20e85ac041affa31e75fb5923af684ac

Observation 93941efd-55ef-4822-9be0-74563ea52630 · outbound

This paper cites Virtual adversarial training: a regularization method for supervised and semi-supervised learning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Virtual adversarial training: a regularization method for supervised and semi-supervised learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.855444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.035740Z digest=sha256:79b2b641237a2fca1898e9eaaa222f4a6cbe2f5b6f5c5c1705cb5e20ba23809d

Observation 8aca17d5-d9ff-4695-9543-791e580d92ad · outbound

This paper cites Fine-tuning can cripple your foundation model; preserving features may be the solution.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Fine-tuning can cripple your foundation model; preserving features may be the solution

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.154739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.154739Z digest=sha256:755553fd72214b069583d9519568f88ed0784d0b08232c2ec20fbb8a98f5b4b3

Observation 0899e5f3-95e0-4dfc-9c9b-61cd7d5dd285 · outbound

This paper cites Robust stochastic approximation approach to stochastic programming.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Robust stochastic approximation approach to stochastic programming

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.235844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.235844Z digest=sha256:3449c6c3f86e7c367878b4738cd8eb784fd5b0c993bc2591a35ec5ef2e1525ba

Observation 31bf7f29-8d3d-4195-b990-02fa3aec3dd0 · outbound

This paper cites Distributionally robust language modeling.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Distributionally robust language modeling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.666994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.306031Z digest=sha256:f74499db92c375edf0484c6c91099eab277cabdbaa09a8e1c4f772760a63fdd3

Observation ed815a8c-94fe-46df-8523-01939a05d727 · outbound

This paper cites Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.384137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.384137Z digest=sha256:51eb65d0b16e26afd19973949af44f7b774050f3afe9ba089f359163788908e9

Observation e9fa4308-f270-4c84-a1b8-eccdf39237e8 · outbound

This paper cites Gradient starvation: A learning proclivity in neural networks.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Gradient starvation: A learning proclivity in neural networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.450064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.471377Z digest=sha256:eef0f7d2a42170c95ec17b8bd939bb5b920b3fc80baa93b8d05d2511e59837ad

Observation 9213de61-e392-432a-957a-01f18cc978a4 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.175737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.585683Z digest=sha256:f581f159ee14a5a2f1db1cd5066296aa5433d61c9bd303962c4fea7d22adf0c7

Observation e024dd50-b677-4397-9582-d93faf791fef · outbound

This paper cites Representation Noising: A Defence Mechanism Against Harmful Finetuning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Representation Noising: A Defence Mechanism Against Harmful Finetuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.656919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.656919Z digest=sha256:5b6eb9af665259515e92168631a994ee17e6265d2e6e79774660229852d47f9c

Observation 692cf9e7-b23c-4523-93e3-a926ffd1e9cc · outbound

This paper cites Immunization against harmful fine-tuning attacks.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Immunization against harmful fine-tuning attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.777027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.777027Z digest=sha256:7016a402fca68ec467e06f170b6cd496a86fbd9beab883c01d57fcc367f73220

Observation a6c52886-6a07-4020-a183-0c9c52698cd9 · outbound

This paper cites W., Hashimoto, T.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning W., Hashimoto, T

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.841598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.841598Z digest=sha256:e9d556d8ad47f698c71d3effd8d41c9afec8852b7a40b3a9b577691edcf840bb

Observation 6970cdf6-317a-4e6d-be18-0ac31cd90c05 · outbound

This paper cites D., Ng, A.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning D., Ng, A

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.918620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.918620Z digest=sha256:d83196fc85d98e09ef169a9aac334e29920163a8fc8d54cc93ad738c0ac2293a

Observation 03d73a50-2a49-4b4e-841c-df5b920aea78 · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.018033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.018033Z digest=sha256:a0917ae30deecace25a03f3fced55cc7d62c3a91fd0039838ae7088f794cd17d

Observation 836a6827-2136-4cc3-99cc-2cc4c5eabf0b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.116438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.116438Z digest=sha256:7790037f51dc667744feaefb85455c83565cee66d282530ca90feb2d0ade7a63

Observation 4f4a639c-f73a-4dc8-adfe-3056e664ee1f · outbound

This paper cites Qwen2 Technical Report.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.223942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.223942Z digest=sha256:926f3fc83f25f5892ae3cf28bb23b7acef1206a7c8cab76fe1fd2dcfe757218d

Observation 2e6a2e86-1dfb-49eb-bca1-21e2781a527e · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.316238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.316238Z digest=sha256:6b3d8cca0717c96ac6f2958d8bbfb4fe891395f0932977d8b3c070c5baf948dc

Observation 34a5e47a-548e-4a19-b784-3991868611b4 · outbound

This paper cites A safety realignment framework via subspace-oriented model fusion for large language models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning A safety realignment framework via subspace-oriented model fusion for large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.433283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.433283Z digest=sha256:a4ccda4c0592e67a93672074e8293d1ad7a826128e85c143b07bcc4a9fef2ad1

Observation 382b390f-a3a8-4031-8574-31eb66cb9f27 · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.538591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.538591Z digest=sha256:ef1a9c68df54ebd62df7da91e32678b91d7dee03f6f44c86c53da8b9e0dd8ce6

Observation c09f958c-6138-4837-b436-8b3cca594162 · outbound

This paper cites Character-level convolutional networks for text classification.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Character-level convolutional networks for text classification

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.629027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.629027Z digest=sha256:7fd9a1808040f4c536be11ab24ae201fe07bc3652bd02f460c2f5ba747283965

Pith citing papers

Observation 7f15e16c-f8f7-4280-8880-d2d3f12f7f79 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.112067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:bf6472db6d0535950994c6115ad0c5f7999f5f32ad98df905a984a6fc77761a7

Observation 485d1a7f-70e7-45c9-800c-5887115546fb · inbound

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection cites this paper.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.667357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:49:56.311711Z digest=sha256:aae6c5c7849d35b6f8c438fd84a7e72b5114ec4e661e7e8cfa93297c84ea86b5