Pith. sign in

Paper Citation Record · LEDGER

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

As of 15 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2506.03850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03850 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:58:32.629027Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:01:25.549340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.110492Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85be4b79-7bc6-4a4b-8739-558370a5ff14 · outbound

This paper cites write newline.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.960888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.960888Z digest=sha256:e9296bef1e1db17178573d86a6b21d8ea73b29b8b7649915d90e83f7a93f6986

Observation 9201a311-3786-4f57-8a1a-d464c3807d6d · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.049879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.049879Z digest=sha256:c8703fc8ab272343d231f6d08bd11346165d43fcf001b739d910a2713485946d

Observation 07745fbe-1cc1-4cbb-8cb0-60878b663edb · outbound

This paper cites D., Melenberg, B., and Rennen, G.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning D., Melenberg, B., and Rennen, G

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.806351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.166245Z digest=sha256:6bebdf6ae801dbf95fb920f22ceb7b43c09e3aadc5b98f472ccd2f6be70cf8c9

Observation 979dc913-5965-452a-ab2f-f847184f2920 · outbound

This paper cites Curriculum learning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Curriculum learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.617713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.274612Z digest=sha256:df2816b11573d0f1bca4432d19e50f14df4311ed4f1d0d51dc67479b869a7793

Observation fc7b2970-884a-458e-87cd-2e018560ef7b · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.349795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.349795Z digest=sha256:0d222f176bb06a9ccf62fe30eb1f9816d4a92d0f0d794a6d6804f797aa8f3541

Observation 1b388d9c-4b5e-4e9b-a117-6132bfe03b9b · outbound

This paper cites Beyond factuality: A comprehensive evaluation of large language models as knowledge generators.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Beyond factuality: A comprehensive evaluation of large language models as knowledge generators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.468902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.468902Z digest=sha256:b6ee748408ab0446e48cc7f2ce74eb508a296f8e12d03a99e690ee939dd346e0

Observation 6b99841c-5412-4628-8e93-984729d3a5e3 · outbound

This paper cites W at ME : Towards lossless watermarking through lexical redundancy.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning W at ME : Towards lossless watermarking through lexical redundancy

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.552869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.552869Z digest=sha256:872d206f0b19b20cda00cb6cd7e613d7e4ab26682ce275b191ae9fe6dcece17c

Observation 8d33798b-9766-4add-9d45-ff2fc5d2eb2d · outbound

This paper cites Simple permutations can fool LL a MA : Permutation attack and defense for large language models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Simple permutations can fool LL a MA : Permutation attack and defense for large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.415083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.641840Z digest=sha256:8d46786aa7d4153789189f1408c4e44272eddef89290f063b3ce50af822b5fc7

Observation c4a364ec-717d-4543-a634-bba1b2c0d6f6 · outbound

This paper cites PEARL : Towards permutation-resilient LLM s.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning PEARL : Towards permutation-resilient LLM s

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.239385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:28.780718Z digest=sha256:49f0ae447bcdd85621ad28114b812e16e7af5c59aefaa78b4a77055bc32eaa95

Observation d4017aa4-c6e8-4a63-893e-37056372c498 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:28.922486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:28.922486Z digest=sha256:8a787e11f24fd2f695a3b630d67bcfbe4ffa82941f4d5a68a8ce684d4ceec8f1

Observation ee26b7b9-f979-4f71-b552-d64461331b23 · outbound

This paper cites Statistics of robust optimization: A generalized empirical likelihood approach.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Statistics of robust optimization: A generalized empirical likelihood approach

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:35.054218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.035354Z digest=sha256:3e036a19f4dc1d80049951ba43087599d2c506e06350a99c1b5ccfa9b78a0edc

Observation 5b6e0b8d-9112-4c43-abf2-31217a9a4b2e · outbound

This paper cites Sharpness-aware minimization for efficiently improving generalization.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Sharpness-aware minimization for efficiently improving generalization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.861219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.096459Z digest=sha256:48ddd6a006eb5e4375bf9dcc9ccd9686803c71827da41594d03221d5a1f99066

Observation 0a7329e2-0b63-40cb-a845-8e5fd7ed970a · outbound

This paper cites A Survey of Uncertainty in Deep Neural Networks.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning A Survey of Uncertainty in Deep Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.226049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.226049Z digest=sha256:d147fbe6ee6b1731fd9189ee65d6f1d7d64381629ec9b9b41f103896eb09a263

Observation d79474ea-dfc3-4cb8-9fc8-8c788cb0b3da · outbound

This paper cites Safe lora: the silver lining of reducing safety risks when fine-tuning large language models, 2024.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Safe lora: the silver lining of reducing safety risks when fine-tuning large language models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.635342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.316371Z digest=sha256:16524d3ffdc8381b5b79231b19c6e4e34a614bcce64b94e36483b2e52686305b

Observation 3039b9af-f8d2-4aeb-b693-0f6cb6283dd8 · outbound

This paper cites Does distributionally robust supervised learning give robust classifiers? In International Conference on Machine Learning (ICML), 2018.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Does distributionally robust supervised learning give robust classifiers? In International Conference on Machine Learning (ICML), 2018

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.437632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:29.426723Z digest=sha256:765daea33879dc59fd2ab42d3b84ecc45a09d3d1f99202765fc9112a79260005

Observation 5ce7a588-a395-4725-802b-23fa97cb6001 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.556438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.556438Z digest=sha256:a36bde27860f951177c9d4d6ce966e5fd8875505f68312e326eb3f2763ff6c2f

Observation 432e1bfa-79a9-4200-a759-08918ef0de8c · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.726809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.726809Z digest=sha256:bb405eef8fcd0bd54df46783b5e2eed2dca38fab789aa5448a104e740d698dba

Observation 68912f69-44e1-4a26-9e9d-4eccb783f9cf · outbound

This paper cites Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:29.862212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:29.862212Z digest=sha256:7ac344914d2c30ec095ccb17cadc835102c668b0efb66e3b7165589b6ce580dd

Observation 2b67a653-2ba4-40ac-b187-9c56cd80cf82 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.014066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.014066Z digest=sha256:e1bc99b2d793ba2da56854b06ad2b2fcf1217d00dc72461d51d1ca95ddedae02

Observation 85e40ba6-bc56-47e8-9791-7ac500928d3a · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.125758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.125758Z digest=sha256:f4fb68c5edade2d278d2e69456ad60eb1c5c78afe81f3392c7a54f4e3e048052

Observation 7b62c561-52f3-42b0-9a61-828f87d4bb22 · outbound

This paper cites A watermark for large language models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning A watermark for large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.258151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:30.240558Z digest=sha256:bdbcb64f2a9ac3eb3c4fc0109fc289889f05dbf9157f7e872e03fb4b89be4fea

Observation 9045197d-47fa-47bb-aca2-3d50cea61a3d · outbound

This paper cites and Zhou, E.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning and Zhou, E

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:34.043488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:30.373584Z digest=sha256:11f971edba175048db4313e255731b7072ec0bd5f40128a8289ccef4a0ebd283

Observation 992acd13-9279-4488-99b5-01a2cc5cbbb7 · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.489568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.489568Z digest=sha256:c5aaa56542ecc3f53aa438cf9551e345c8c68f30eac0de77feea5a13e8a137c7

Observation 56a9ff7c-d06a-43d1-a381-a3109563f274 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods, 2022.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Truthfulqa: Measuring how models mimic human falsehoods, 2022

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.674810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.674810Z digest=sha256:78035d7f4f099b8a9098cbcff97526499d4ae5e974d14de9d301f30149e5009e

Observation 51fd9752-942b-42b2-90aa-7a1795e26c32 · outbound

This paper cites Decoupled Weight Decay Regularization.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Decoupled Weight Decay Regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.814159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.814159Z digest=sha256:00e240db8e024144d5776e1c0c84a9d33f5de7279f8a22fc86902f6479a32581

Observation 4fed7e0f-c7a7-4eb8-808c-4691a40cde4a · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.941185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:30.941185Z digest=sha256:e8a0429508c359e9db3db665a20777382c60e29cfcd41ab0780f3a2b5b4700ff

Observation 93941efd-55ef-4822-9be0-74563ea52630 · outbound

This paper cites Virtual adversarial training: a regularization method for supervised and semi-supervised learning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Virtual adversarial training: a regularization method for supervised and semi-supervised learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.855444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.035740Z digest=sha256:baee063b342642eebdec61b338d7ba3293db122480e5b38ab109fff697ffc80b

Observation 8aca17d5-d9ff-4695-9543-791e580d92ad · outbound

This paper cites Fine-tuning can cripple your foundation model; preserving features may be the solution.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Fine-tuning can cripple your foundation model; preserving features may be the solution

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.154739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.154739Z digest=sha256:be0d76775dd8c902c03d050db76781e7f73b98e771a71528e4c1ad94c6dd9a4a

Observation 0899e5f3-95e0-4dfc-9c9b-61cd7d5dd285 · outbound

This paper cites Robust stochastic approximation approach to stochastic programming.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Robust stochastic approximation approach to stochastic programming

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.235844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.235844Z digest=sha256:d15e895dfcf256c22f5e3096ded38d2ba46ef93ce2a64be037dcb18f54011b43

Observation 31bf7f29-8d3d-4195-b990-02fa3aec3dd0 · outbound

This paper cites Distributionally robust language modeling.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Distributionally robust language modeling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.666994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.306031Z digest=sha256:d2597aa561a0853a243ce09075a655e0cfc4111957010dd1393a64c80db1a115

Observation ed815a8c-94fe-46df-8523-01939a05d727 · outbound

This paper cites Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.384137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.384137Z digest=sha256:867696912401d8bd0a7444f82f593f15d64220bb0decfa7e5b8a538721a66a7c

Observation e9fa4308-f270-4c84-a1b8-eccdf39237e8 · outbound

This paper cites Gradient starvation: A learning proclivity in neural networks.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Gradient starvation: A learning proclivity in neural networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.450064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.471377Z digest=sha256:fd17651b8a734412daf859767aab96a107bf06f573fb2aaa8c169454a6fe0618

Observation 9213de61-e392-432a-957a-01f18cc978a4 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:58:33.175737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:58:31.585683Z digest=sha256:a10f7c7ddb1c037bdc0bded8698a6b6fedce7b8d31c841e9c4ff7a00acc0c000

Observation e024dd50-b677-4397-9582-d93faf791fef · outbound

This paper cites Representation Noising: A Defence Mechanism Against Harmful Finetuning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Representation Noising: A Defence Mechanism Against Harmful Finetuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.656919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.656919Z digest=sha256:35e3609b075a106c1852444bbb8e85e878c5d2a23109cb4656fc3c2b5f0529b2

Observation 692cf9e7-b23c-4523-93e3-a926ffd1e9cc · outbound

This paper cites Immunization against harmful fine-tuning attacks.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Immunization against harmful fine-tuning attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.777027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.777027Z digest=sha256:05e83578fd4bf4d8414d88dc28a715b7a157c24b66f4e6bf557c58c6713fb3da

Observation a6c52886-6a07-4020-a183-0c9c52698cd9 · outbound

This paper cites W., Hashimoto, T.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning W., Hashimoto, T

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.841598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.841598Z digest=sha256:a9b59dad1731a095d8c82e35f973bb13abb1ce336d1cd538bedfae6c5c0741f7

Observation 6970cdf6-317a-4e6d-be18-0ac31cd90c05 · outbound

This paper cites D., Ng, A.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning D., Ng, A

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:31.918620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:31.918620Z digest=sha256:2fc1dc71605bd44b8c624357962993e61edf882051c2e574c82a8635a8a063ee

Observation 03d73a50-2a49-4b4e-841c-df5b920aea78 · outbound

This paper cites an unresolved cited work.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.018033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.018033Z digest=sha256:c226193a376917cb3994363a234ceca0f98af7801296f851f2c7116841f1accb

Observation 836a6827-2136-4cc3-99cc-2cc4c5eabf0b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.116438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.116438Z digest=sha256:05158c3a1ebcb3c30e15fbc340b14e3f0fa948994691d9f8b76c0d93c78546ec

Observation 4f4a639c-f73a-4dc8-adfe-3056e664ee1f · outbound

This paper cites Qwen2 Technical Report.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Qwen2 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.223942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.223942Z digest=sha256:976a81acd3e465b1ade291bcc2e95f99275e96ec1db00430cf6a16628f7fe345

Observation 2e6a2e86-1dfb-49eb-bca1-21e2781a527e · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.316238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.316238Z digest=sha256:8227417a39b57bbe58e1798bcda7df0e541d1276ab91ec305ab4021be7153cd7

Observation 34a5e47a-548e-4a19-b784-3991868611b4 · outbound

This paper cites A safety realignment framework via subspace-oriented model fusion for large language models.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning A safety realignment framework via subspace-oriented model fusion for large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.433283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.433283Z digest=sha256:8ce599763b7d30d832b0ec165c54a1aa0d24c0294a20bede120e7ed44d4d265a

Observation 382b390f-a3a8-4031-8574-31eb66cb9f27 · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.538591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.538591Z digest=sha256:3f8fa51636917626fd35972c2390597df55bf40d25b686d55b42d7d67d6494b2

Observation c09f958c-6138-4837-b436-8b3cca594162 · outbound

This paper cites Character-level convolutional networks for text classification.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Character-level convolutional networks for text classification

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.629027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.629027Z digest=sha256:2d426147ffd5c54dd811f6711a084f7ca7873af1a6009816b7ba4eeb75d2106f

Pith citing papers

Observation 7f15e16c-f8f7-4280-8880-d2d3f12f7f79 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.112067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:2f31098d4f52073000d4c4f8b04d2b6c64327ed006f5902fae157d1915ec4461

Observation 485d1a7f-70e7-45c9-800c-5887115546fb · inbound

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection cites this paper.

SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.667357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:49:56.311711Z digest=sha256:85ca3a2e0ccffb14a2c6cbf71189823889c3d9bf56b4e8a691e58d9a4e9ef186